2B - "AI Infrastructure & Evidence Systems"
Tracks
Track 2
Case studies on AI use in public health
Educating our workforce for AI
Enabling infrastructure for AI (sharing data, federated learning, etc.)
Equity and Ethics (privacy, accessibility, information bias, confidentiality and security)
| Wednesday, November 11, 2026 |
| 1:30 PM - 3:00 PM |
Speaker
Dr Lavender Otieno
Research Fellow
Flinders University
Digitising Information for Practice in Public Health (SMART-PH)
Abstract
Background and Aim
COVID 19 revealed major gaps in public health systems, particularly the lack of real time, integrated insights to guide prevention and response. SMART-PH was established in SA to address this by developing a statewide public health data analytics platform with embedded machine learning capability. By integrating multi-agency data, the platform will generate timely intelligence to strengthen decision making.
Methods and Analysis
SMART PH is being designed and delivered through a co-design, co-development, co-evaluation and implementation framework involving public health practitioners, data scientists, statisticians, epidemiologists, clinicians, local councils, Aboriginal organisations, primary health networks, non government organisations and communities.
• Phase 1: Co design - Engaged 20 participants across nine organisations through interviews, focus groups and a large co design workshop. Stakeholders identified priority public health areas, workflow requirements, implementation enablers and barriers, and desired platform features.
• Phase 2: Data Lake development - Establishing a statewide Public Health Data Lake by introducing new public health data assets into the existing Digital Analytics Platform, guided by a comprehensive data audit.
• Phase 3: Machine learning tools - Developing, testing and evaluating public health decision support tools using functional, analytical and interactive AI techniques.
• Phase 4: Implementation - Co developing implementation strategies, a SMART PH Governance Framework and a National Public Health AI Community of Practice to support capability building and national scalability.
Outcomes
Phase 1 findings highlight diverse organisational data needs, including access to additional datasets, real time data, automated reporting and analytics capability building. Subsequent phases will deliver a statewide data lake with AI driven tools and sustainable implementation strategies.
Conclusion and Future Action
SMART-PH aims to future proof SA's public health system to address climate related health impacts, rising chronic disease, preventable injury and emerging infectious diseases, advancing proactive, technology enabled prevention across the life course.
COVID 19 revealed major gaps in public health systems, particularly the lack of real time, integrated insights to guide prevention and response. SMART-PH was established in SA to address this by developing a statewide public health data analytics platform with embedded machine learning capability. By integrating multi-agency data, the platform will generate timely intelligence to strengthen decision making.
Methods and Analysis
SMART PH is being designed and delivered through a co-design, co-development, co-evaluation and implementation framework involving public health practitioners, data scientists, statisticians, epidemiologists, clinicians, local councils, Aboriginal organisations, primary health networks, non government organisations and communities.
• Phase 1: Co design - Engaged 20 participants across nine organisations through interviews, focus groups and a large co design workshop. Stakeholders identified priority public health areas, workflow requirements, implementation enablers and barriers, and desired platform features.
• Phase 2: Data Lake development - Establishing a statewide Public Health Data Lake by introducing new public health data assets into the existing Digital Analytics Platform, guided by a comprehensive data audit.
• Phase 3: Machine learning tools - Developing, testing and evaluating public health decision support tools using functional, analytical and interactive AI techniques.
• Phase 4: Implementation - Co developing implementation strategies, a SMART PH Governance Framework and a National Public Health AI Community of Practice to support capability building and national scalability.
Outcomes
Phase 1 findings highlight diverse organisational data needs, including access to additional datasets, real time data, automated reporting and analytics capability building. Subsequent phases will deliver a statewide data lake with AI driven tools and sustainable implementation strategies.
Conclusion and Future Action
SMART-PH aims to future proof SA's public health system to address climate related health impacts, rising chronic disease, preventable injury and emerging infectious diseases, advancing proactive, technology enabled prevention across the life course.
Biography
Dr Otieno is a Kenyan Australian early career researcher (ECR) and Co-Deputy Discipline Lead of Trauma and Injury at the College of Medicine and Public Health at Flinders University. Positioned at the interface of research and health policy, her career is dedicated to bridging the gap between scientific evidence and real-world public health practice for translatable change. As a Research Fellow and project manager, Dr Otieno oversees SMART-PH (Digitising Information for Practice in Public Health), a NRCI MRFF aimed at developing the first public health data late, with integrated artificial intelligence for insight delivery. This innovative initiative seeks to unify public health data from diverse agencies, enhancing data interoperability and decision-making capabilities for public health practice.
Ms Yuan Gao
Phd Student
Adelaide Health Technology Assessment
Reporting Artificial Intelligence Use in Health Technology Assessment: The ExplAIn Framework
Abstract
Background and Aim: Recent advances in artificial intelligence (AI) have been reshaping workflows in the evidence submitted, evaluated and appraised by Health Technology Assessment (HTA) and regulatory agencies internationally. This scoping review maps how AI is reported in current HTA and regulatory tasks.
Methods and Analysis: Following a published protocol (10.17605/OSF.IO/BSCZW), we conducted literature searches across seven databases (2020-2025), two HTA research websites, and HTA/regulatory websites across 22 countries. Papers were included if they concerned tasks related to the HTA or regulatory evaluation of technologies and involved AI. Papers evaluating AI health technologies were excluded. Data were extracted using a pre-defined table.
Outcomes: 65 peer-reviewed papers and 22 HTA/regulatory documents met the inclusion criteria. There was high heterogeneity in the types of AI used and HTA/regulatory tasks addressed. Reporting of AI model development generally lacked transparency. After excluding commercial products, 10 studies (20%) did not describe the model optimisation process, and 32 (65%) did not provide access to the source code. Furthermore, the rationale for using AI in specific HTA/regulatory tasks was often difficult to assess, as only around half of the included studies reported appropriate comparators for benchmarking performance and provided performance metrics (n = 33, 51%). Performance evaluation was also highly heterogeneous, with the greatest consistency observed in free-text processing tasks and the highest variability in comparative effectiveness assessments. Detailed HTA and regulatory guidance often remains limited, as the AI use cases and technology performance are still evolving. To address these challenges, we propose the ExplAIn framework, a novel reporting framework designed to improve transparency in the use of AI within HTA and regulatory contexts.
Conclusion and Future actions: AI reporting standards for HTA and regulatory tasks should be improved. Given AI's adaptability across use contexts, its performance should be evaluated empirically within the intended use setting.
Methods and Analysis: Following a published protocol (10.17605/OSF.IO/BSCZW), we conducted literature searches across seven databases (2020-2025), two HTA research websites, and HTA/regulatory websites across 22 countries. Papers were included if they concerned tasks related to the HTA or regulatory evaluation of technologies and involved AI. Papers evaluating AI health technologies were excluded. Data were extracted using a pre-defined table.
Outcomes: 65 peer-reviewed papers and 22 HTA/regulatory documents met the inclusion criteria. There was high heterogeneity in the types of AI used and HTA/regulatory tasks addressed. Reporting of AI model development generally lacked transparency. After excluding commercial products, 10 studies (20%) did not describe the model optimisation process, and 32 (65%) did not provide access to the source code. Furthermore, the rationale for using AI in specific HTA/regulatory tasks was often difficult to assess, as only around half of the included studies reported appropriate comparators for benchmarking performance and provided performance metrics (n = 33, 51%). Performance evaluation was also highly heterogeneous, with the greatest consistency observed in free-text processing tasks and the highest variability in comparative effectiveness assessments. Detailed HTA and regulatory guidance often remains limited, as the AI use cases and technology performance are still evolving. To address these challenges, we propose the ExplAIn framework, a novel reporting framework designed to improve transparency in the use of AI within HTA and regulatory contexts.
Conclusion and Future actions: AI reporting standards for HTA and regulatory tasks should be improved. Given AI's adaptability across use contexts, its performance should be evaluated empirically within the intended use setting.
Biography
Yuan Gao is a PhD candidate in Health Technology Assessment (HTA) at the University of Adelaide and a Health Data Scientist at Flinders University. With a background in public health and biostatistics, her previous research spans public decision-making, healthcare reimbursement, cancer epidemiology, and the application of machine learning and statistical modelling to clinical trial data. Her current research focuses on the evaluation and transparent reporting of artificial intelligence (AI) in HTA and regulatory decision-making, with the aim of supporting more rigorous, trustworthy, and evidence-informed use of AI in HTA.
A/prof JODIE AVERY
RRI Research Leadership Co-lead - Chronic Reproductive Conditions
Robinson Research Institute, Adelaide University
AI-Enabled Structured Data Extraction: Revolutionizing Endometriosis Research and Public Health Policy
Abstract
Background and Aim Endometriosis affects 1 in 7 women and individuals assigned female at birth, yet diagnosis remains plagued by a 6.4-year average delay. Addressing this public health crisis requires large-scale clinical research to address diagnostic delay, including data driven projects like IMAGENDO® which combines ultrasounds and magnetics resonance images to diagnoses endometriosis. However, a critical bottleneck exists: the labour-intensive extraction of structured data from unstructured ultrasound reports. This manual process limits research scalability and reproducibility, hindering the development of evidence-based policies. Our aim was to evaluate whether Large Language Models (LLMs) could automate this extraction, creating the enabling infrastructure needed for high-quality, large-scale public health datasets.
Methods and Analysis We conducted a case study evaluating three locally hosted LLMs (Llama3-8b, Mistral-7b, and GPT-oss:20b) to transform 49 narrative endometriosis transvaginal ultrasound (eTVUS) reports into structured data. To address conference sub-themes of Ethics and Privacy, LLMs were hosted on local servers to safeguard sensitive health information. Performance was benchmarked against a blinded human extractor and validated by an expert sonographer to interrogate potential information bias and ensure data integrity.
Outcomes While human accuracy (98.4%) surpassed the LLMs (78.89%–86.02%), we identified a complementary "hybrid" potential. LLMs outperformed humans in extracting numeric fields, whereas humans were superior in semantic and categorical fields. This finding demonstrates how AI can maximize benefits by reducing manual workload and highlighting inconsistencies, while human oversight minimizes the harm of semantic errors.
Conclusion and Future actions Future public health practice should adopt a hybrid human–AI workflow to enable scalable, high-quality imaging research. We must prioritise educating our workforce to manage these tools and strengthen cross-sector collaboration to build robust governance frameworks. By automating data extraction, we can accelerate the research necessary to improve diagnostic pathways, ultimately reducing the profound physical and economic burdens of endometriosis.
Methods and Analysis We conducted a case study evaluating three locally hosted LLMs (Llama3-8b, Mistral-7b, and GPT-oss:20b) to transform 49 narrative endometriosis transvaginal ultrasound (eTVUS) reports into structured data. To address conference sub-themes of Ethics and Privacy, LLMs were hosted on local servers to safeguard sensitive health information. Performance was benchmarked against a blinded human extractor and validated by an expert sonographer to interrogate potential information bias and ensure data integrity.
Outcomes While human accuracy (98.4%) surpassed the LLMs (78.89%–86.02%), we identified a complementary "hybrid" potential. LLMs outperformed humans in extracting numeric fields, whereas humans were superior in semantic and categorical fields. This finding demonstrates how AI can maximize benefits by reducing manual workload and highlighting inconsistencies, while human oversight minimizes the harm of semantic errors.
Conclusion and Future actions Future public health practice should adopt a hybrid human–AI workflow to enable scalable, high-quality imaging research. We must prioritise educating our workforce to manage these tools and strengthen cross-sector collaboration to build robust governance frameworks. By automating data extraction, we can accelerate the research necessary to improve diagnostic pathways, ultimately reducing the profound physical and economic burdens of endometriosis.
Biography
A/Prof Jodie Avery is an epidemiologist and public health researcher with a background in medical radiations, psychology, and social sciences. Her extensive experience in developing research policy partnerships and conducting evaluations is evidenced by over 100 publications and more than 25 commissioned government reports across diverse areas. Her work has influenced policy changes through Adelaide University, SA Health, SAHMRI, and in her role as Vice President of the SA Branch of the Public Health Association of Australia (PHAA). An expert in systematic review methodology, she teaches this alongside public health, medical, and psychological subjects at the tertiary level.
She is the Research Leadership Co-Lead for Chronic Reproductive Conditions at the Robinson Research Institute and Program Director of the IMAGENDO Study within the Endometriosis Group at the Institute. Over more than 25 years, she has conducted research and advocacy in women’s chronic reproductive health.
Ms Tina La
Policy And Program Officer
The Australian Multicultural Health Collaborative
Missing from the Data, Missing from the Algorithm
Abstract
Artificial intelligence in public health is only as equitable as the data it learns from. In Australia, that data systematically excludes culturally and linguistically diverse (CALD) communities. Over 23% of Australians currently speak a language other than English at home, yet national health datasets remain extremely unrepresentative of multicultural populations. This abstract identifies the data gaps that must be first addressed to then build the AI infrastructures capable of delivering equitable public health outcomes for all Australians.
FECCA conducted a structured policy analysis of data collection practices across administrative datasets, population surveys, and health and social research in Australia. This included a qualitative assessment of how cultural and linguistic diversity is measured across Commonwealth and state/territory systems, benchmarking Australian practice against international standards. Analysis was informed by consultations with key stakeholders including the ABS, AIHW, and academics from ANU and the University of Sydney, resulting in the published paper 'If We Don't Count It... It Doesn't Count!'.
The analysis found systemic failure across three domains that are now critical AI infrastructure concerns: inconsistent application of ABS Standards for Statistics on Cultural and Language Diversity; active exclusion of CALD populations through English-proficiency criteria in research that AI models are trained on; and no mandated CALD data reporting across government agencies. Findings highlight new concerns around how the existing datasets being used to build and validate current AI health tools will carry the same gaps identified years ago.
CALD data reform must be treated as a non-negotiable condition before AI deployment across Australian health systems. Our recommendations begin with first mandating cultural and language diversity variables across health datasets as core enabling infrastructure for AI. Ensuring CALD communities are visible to the AI algorithms making decisions about them will also support the mitigation of adverse health impacts from newly procured AI tools.
FECCA conducted a structured policy analysis of data collection practices across administrative datasets, population surveys, and health and social research in Australia. This included a qualitative assessment of how cultural and linguistic diversity is measured across Commonwealth and state/territory systems, benchmarking Australian practice against international standards. Analysis was informed by consultations with key stakeholders including the ABS, AIHW, and academics from ANU and the University of Sydney, resulting in the published paper 'If We Don't Count It... It Doesn't Count!'.
The analysis found systemic failure across three domains that are now critical AI infrastructure concerns: inconsistent application of ABS Standards for Statistics on Cultural and Language Diversity; active exclusion of CALD populations through English-proficiency criteria in research that AI models are trained on; and no mandated CALD data reporting across government agencies. Findings highlight new concerns around how the existing datasets being used to build and validate current AI health tools will carry the same gaps identified years ago.
CALD data reform must be treated as a non-negotiable condition before AI deployment across Australian health systems. Our recommendations begin with first mandating cultural and language diversity variables across health datasets as core enabling infrastructure for AI. Ensuring CALD communities are visible to the AI algorithms making decisions about them will also support the mitigation of adverse health impacts from newly procured AI tools.
Biography
Tina brings an evidence-based analytical approach to her role in advocacy and project delivery of initiatives targeting multicultural health outcomes. With an Honours in Psychology from the ANU, her research degree has equipped her with experience in communicating the cultural, social, and psychological factors impacting health outcomes through written work. During this time, her work bridged empirical insight with a commitment to amplifying marginalized voices, translating research into meaningful social change. From here, Tina has collaborated on papers addressing different marginalised populations, most recently publishing a paper on the importance of diversity initiatives to improve education outcomes. She is committed to improving wellbeing and equity for diverse and underrepresented groups through leveraging her research-informed experience alongside community-led collaboration.
Dr Sashika Harasgama
Public Health Registrar
Western Public Health Unit
Adoption ahead of evidence? Governing AI evidence-synthesis tools in public health
Abstract
Problem
Public health agencies are adopting AI to speed up evidence synthesis - the foundation of evidence-based policy and practice - faster than the evidence needed to govern that adoption. Decision-makers face a crowded, fast-changing tool market with little guidance on which tools are reliable, what they cost, who they exclude, and what harms they carry.
What we did
We conducted a systematic scoping review (Joanna Briggs Institute; PRISMA-ScR) of 222 studies, mapping 65 evaluated AI tools across the evidence-synthesis pathway and the risks attached to each, complemented by a conceptual appraisal of the underlying approaches, from traditional machine learning to large language models (LLMs).
Results
Adoption is outpacing evidence. Studies of generative AI chatbots underpinned by LLMs - conceptually the highest-risk approach (hallucination, opacity, non-determinism) - rose steeply in 2024, yet no studies compared them with established machine-learning tools, and only 4.1% measured the workload or time savings that justify adoption. Almost all tools sat behind paywalls or subscriptions, raising equity and access concerns. Most were built for biomedical research and optimised for randomised controlled trials, not the multidisciplinary, heterogeneous evidence base typical of public health. The sustainability of LLMs was essentially unstudied.
Lessons
Organisations should not adopt AI evidence-synthesis tools on accuracy claims alone. Procurement and governance should require feasibility and comparative-effectiveness evidence, account for cost and access equity, apply responsible-use standards (e.g. RAISE [1]) with a researcher kept in the loop, and consider sustainability.
References:
1. Thomas J, Flemyng E, Noel-Storr A, Moy W, Marshall IJ, Hajji R, et al. Responsible use of AI in evidence SynthEsis (RAISE): recommendations for practice. In: Open Science Framework [https://osf.io/], Washington DC: Center for Open Science (accessed 22 September 2025). https://doi.org/10.17605/OSF.IO/FWAUD
Public health agencies are adopting AI to speed up evidence synthesis - the foundation of evidence-based policy and practice - faster than the evidence needed to govern that adoption. Decision-makers face a crowded, fast-changing tool market with little guidance on which tools are reliable, what they cost, who they exclude, and what harms they carry.
What we did
We conducted a systematic scoping review (Joanna Briggs Institute; PRISMA-ScR) of 222 studies, mapping 65 evaluated AI tools across the evidence-synthesis pathway and the risks attached to each, complemented by a conceptual appraisal of the underlying approaches, from traditional machine learning to large language models (LLMs).
Results
Adoption is outpacing evidence. Studies of generative AI chatbots underpinned by LLMs - conceptually the highest-risk approach (hallucination, opacity, non-determinism) - rose steeply in 2024, yet no studies compared them with established machine-learning tools, and only 4.1% measured the workload or time savings that justify adoption. Almost all tools sat behind paywalls or subscriptions, raising equity and access concerns. Most were built for biomedical research and optimised for randomised controlled trials, not the multidisciplinary, heterogeneous evidence base typical of public health. The sustainability of LLMs was essentially unstudied.
Lessons
Organisations should not adopt AI evidence-synthesis tools on accuracy claims alone. Procurement and governance should require feasibility and comparative-effectiveness evidence, account for cost and access equity, apply responsible-use standards (e.g. RAISE [1]) with a researcher kept in the loop, and consider sustainability.
References:
1. Thomas J, Flemyng E, Noel-Storr A, Moy W, Marshall IJ, Hajji R, et al. Responsible use of AI in evidence SynthEsis (RAISE): recommendations for practice. In: Open Science Framework [https://osf.io/], Washington DC: Center for Open Science (accessed 22 September 2025). https://doi.org/10.17605/OSF.IO/FWAUD
Biography
Dr Sashika Harasgama is a public health registrar with the Western Public Health Unit in Melbourne, training with the Australasian Faculty of Public Health Medicine. Her work spans communicable disease control and population health, with increasing interest in innovative digital health technologies for public health. She holds a Master of Public Health from the London School of Hygiene and Tropical Medicine and a medical degree from Griffith University.
Before returning to Australia, Sashika worked at the Health Equity Evidence Centre at Queen Mary University of London, where she co-led a Health Foundation commissioned project on the role of AI tools in rapid evidence synthesis. Her team also developed living evidence maps using machine learning to address health inequalities. Her research interests sits at the intersection of health equity, digital health, and the responsible use of artificial intelligence in public health practice.
Dr Navoda Liyana Pathirana
Postdoctoral Research Fellow
Deakin University
SCANNER – deep learning system to automatically detect and classify health-harming marketing
Abstract
Background and aim:
The expansion of digital media has transformed marketing, enabling artificial intelligence (AI)–driven, highly targeted advertising. This increasingly includes promotion of health-harming commodities such as gambling, alcohol, tobacco, vaping, and unhealthy foods, often directed at children and youth. Regulation lags behind these developments due to challenges in monitoring personalised digital marketing. We developed SCANNER, a deep-learning system that automatically detects and classifies marketing of health-harming commodities in video and image data.
Methods and Analysis:
Using a supervised deep-learning approach, SCANNER was trained to detect over 600 brands: 145 gambling, 141 alcohol, 320 food, and 27 breastmilk substitute brands. The system employs YOLOv11 (You Only Look Once, version 11) for logo detection and Paddle Optical Character Recognition (OCR) for brand name identification across gambling, alcohol, and food marketing. For tobacco and vape detection, SCANNER integrates the Gemma Vision language model with YOLO and OCR to capture both branded and non-branded promotional content. To demonstrate SCANNER’s application in digital marketing evaluation, SCANNER was used to quantify digital marketing exposure in screen recordings from 290 Australian children aged 8–25 years.
Outcomes:
SCANNER demonstrated high detection accuracy: 100% for tobacco/vape, 89.8% for gambling, 98.6% for alcohol, and 88.5% for food marketing across structured and unstructured digital environments. SCANNER efficiently analysed over 200 hours of screen recordings, providing the first comprehensive assessment of digital marketing exposure to health-harming commodities among Australian children and youth.
Conclusion and Future actions:
SCANNER enables rapid, reliable quantification of digital marketing exposure, offering a scalable tool for monitoring harmful marketing practices and supporting policy evaluation and regulatory enforcement.
The expansion of digital media has transformed marketing, enabling artificial intelligence (AI)–driven, highly targeted advertising. This increasingly includes promotion of health-harming commodities such as gambling, alcohol, tobacco, vaping, and unhealthy foods, often directed at children and youth. Regulation lags behind these developments due to challenges in monitoring personalised digital marketing. We developed SCANNER, a deep-learning system that automatically detects and classifies marketing of health-harming commodities in video and image data.
Methods and Analysis:
Using a supervised deep-learning approach, SCANNER was trained to detect over 600 brands: 145 gambling, 141 alcohol, 320 food, and 27 breastmilk substitute brands. The system employs YOLOv11 (You Only Look Once, version 11) for logo detection and Paddle Optical Character Recognition (OCR) for brand name identification across gambling, alcohol, and food marketing. For tobacco and vape detection, SCANNER integrates the Gemma Vision language model with YOLO and OCR to capture both branded and non-branded promotional content. To demonstrate SCANNER’s application in digital marketing evaluation, SCANNER was used to quantify digital marketing exposure in screen recordings from 290 Australian children aged 8–25 years.
Outcomes:
SCANNER demonstrated high detection accuracy: 100% for tobacco/vape, 89.8% for gambling, 98.6% for alcohol, and 88.5% for food marketing across structured and unstructured digital environments. SCANNER efficiently analysed over 200 hours of screen recordings, providing the first comprehensive assessment of digital marketing exposure to health-harming commodities among Australian children and youth.
Conclusion and Future actions:
SCANNER enables rapid, reliable quantification of digital marketing exposure, offering a scalable tool for monitoring harmful marketing practices and supporting policy evaluation and regulatory enforcement.
Biography
Dr Navoda Liyana Pathirana is a Postdoctoral Research Fellow at Deakin Centre for Global Preventive Health and Nutrition, within the Deakin Institute for Health Transformations. She is a multidisciplinary researcher working at the intersection of artificial intelligence, public health, and sustainable food systems. Her research focuses on developing data‑driven, AI‑assisted policy evaluation frameworks to address harmful digital marketing, unhealthy food environments, and the commercial determinants of health. As part of her fellowship, she leads ECHO (Evaluating the Co‑benefits of food policies on Health, environmental and socio‑economic Outcomes), which applies environmental‑economic modelling and programming tools to quantify the multiple impacts of food and public health policies. Dr Pathirana has led and contributed to internationally funded projects with partners including UNICEF, WHO, VicHealth, and government agencies, and has published in BMJ, Nature Food, and Public Health Nutrition. Her work directly informs public health regulation in Australia and South Asia.
Dr Sashika Harasgama
Public Health Registrar
Western Public Health Unit
What AI approaches are safe for public health evidence synthesis?
Abstract
Problem
The public health evidence base differs from the clinical trial literature: multidisciplinary and heterogeneous, often qualitative or drawn from non-randomised and grey sources, and highly context dependent. As organisations increasingly use AI to synthesise evidence rapidly, the most accessible tools - general-purpose chatbots such as ChatGPT - may be the least suited, with implications for the reliability of reviews informing public health policy and practice.
What we did
We conducted a conceptual appraisal of AI approaches used in evidence synthesis, from traditional machine learning to LLMs, examining which suit the demands of public health evidence and where their limitations lie, corroborated by a scoping review of 222 studies evaluating 65 distinct tools.
Results
General-purpose chatbots generate text through prediction rather than grounded retrieval, and are prone to hallucination, fabricated citations, opaque reasoning, and variable outputs from identical inputs - at odds with the reproducibility and transparency evidence synthesis demands. They appeared least reliable for precisely this kind of evidence. Purpose-built tools may be more appropriate: machine-learning screening tools are well validated for defined tasks such as priority screening; while retrieval-augmented LLM tools, such as Elicit and Consensus, ground outputs in source documents and may reduce fabrication. However, most were optimised for biomedical trials, none directly compared with general-purpose chatbots, and few addressed public health evidence types.
Lessons
General-purpose chatbots are poorly suited to synthesising public health evidence, and their use warrants caution. Purpose-built tools offering grounded outputs and reproducible methods may be more appropriate, though they remain unvalidated across the multidisciplinary evidence on which public health decision-making depends. Until such tools are validated, their outputs should inform rather than replace expert synthesis.
The public health evidence base differs from the clinical trial literature: multidisciplinary and heterogeneous, often qualitative or drawn from non-randomised and grey sources, and highly context dependent. As organisations increasingly use AI to synthesise evidence rapidly, the most accessible tools - general-purpose chatbots such as ChatGPT - may be the least suited, with implications for the reliability of reviews informing public health policy and practice.
What we did
We conducted a conceptual appraisal of AI approaches used in evidence synthesis, from traditional machine learning to LLMs, examining which suit the demands of public health evidence and where their limitations lie, corroborated by a scoping review of 222 studies evaluating 65 distinct tools.
Results
General-purpose chatbots generate text through prediction rather than grounded retrieval, and are prone to hallucination, fabricated citations, opaque reasoning, and variable outputs from identical inputs - at odds with the reproducibility and transparency evidence synthesis demands. They appeared least reliable for precisely this kind of evidence. Purpose-built tools may be more appropriate: machine-learning screening tools are well validated for defined tasks such as priority screening; while retrieval-augmented LLM tools, such as Elicit and Consensus, ground outputs in source documents and may reduce fabrication. However, most were optimised for biomedical trials, none directly compared with general-purpose chatbots, and few addressed public health evidence types.
Lessons
General-purpose chatbots are poorly suited to synthesising public health evidence, and their use warrants caution. Purpose-built tools offering grounded outputs and reproducible methods may be more appropriate, though they remain unvalidated across the multidisciplinary evidence on which public health decision-making depends. Until such tools are validated, their outputs should inform rather than replace expert synthesis.
Biography
Dr Sashika Harasgama is a public health registrar with the Western Public Health Unit in Melbourne, training with the Australasian Faculty of Public Health Medicine. Her work spans communicable disease control and population health, with increasing interest in innovative digital health technologies for public health. She holds a Master of Public Health from the London School of Hygiene and Tropical Medicine and a medical degree from Griffith University.
Before returning to Australia, Sashika worked at the Health Equity Evidence Centre at Queen Mary University of London, where she co-led a Health Foundation commissioned project on the role of AI tools in rapid evidence synthesis. Her team also developed living evidence maps using machine learning to highlight strategies that address health inequalities. Her research interests sits at the intersection of health equity, digital health, and the responsible use of artificial intelligence in public health practice.
Dr Benjamin Scalley
Medical Director
WA Health
Taming the beast: generating reliable answers to public health questions using AI.
Abstract
Background and aim: Public health units receive a broad range of questions from stakeholders and need to be across a large amount of information and guidelines. Searching guidelines and synthesising an answer can be time consuming and difficult. Using a large language model (LLM) to find an answer could be of benefit but also introduces risks. LLMs can produce a response that is not appropriately tailored to the context of the unit or completely hallucinate an answer. It is also important that the source information is free from bias and is evidence based.
Methods and Analysis: A retrieval augment generation (RAG) model was produced on relevant Australian guidelines and evidence. The system that was developed uses two parallel paths. Firstly semantic RAG which involves embedding the question, cosine-scoring against indexed guideline chunks, LLM reranking, followed by an LLM generated answer from the highest ranked chunks. Secondly text to SQL over structured tables that returns appropriate tabular information to be included in the LLM provided answer.
Outcome: A proof of concept has been produced and information regarding the outcome of its use within a public health unit within Western Australia will be presented. Example questions and answers, including its strengths, weaknesses, and approaches to improve the quality of answers will be presented.
Conclusion and Future action: The use of RAG increases the relevance and accuracy of the answers given. The highly referenced answer allows quick investigation of sources. Future broader use of this approach within public health units would be of benefit.
Methods and Analysis: A retrieval augment generation (RAG) model was produced on relevant Australian guidelines and evidence. The system that was developed uses two parallel paths. Firstly semantic RAG which involves embedding the question, cosine-scoring against indexed guideline chunks, LLM reranking, followed by an LLM generated answer from the highest ranked chunks. Secondly text to SQL over structured tables that returns appropriate tabular information to be included in the LLM provided answer.
Outcome: A proof of concept has been produced and information regarding the outcome of its use within a public health unit within Western Australia will be presented. Example questions and answers, including its strengths, weaknesses, and approaches to improve the quality of answers will be presented.
Conclusion and Future action: The use of RAG increases the relevance and accuracy of the answers given. The highly referenced answer allows quick investigation of sources. Future broader use of this approach within public health units would be of benefit.
Biography
Dr Scalley is a public health physician with recent experience across the breadth of public health. He is currently the medical director of the Boorloo (Perth) Public Health Unit, the largest public health unit in Australia. Through this role he has expanded and modernised the unit with an emphasis on data science, automation, and technical innovation. During COVID-19 he was the head of the Public Health Operations section of the WA Health response. Through this role he lead the establishment of this new area, created a scalable structure and implemented surge processes that resulted in highly successful outcomes for contact tracing and other public health functions in WA. He was also previously the Director of Environmental Health in NSW where he responded to the emergence of PFAS as an important contaminant, drinking water quality issues, and to climate and heat related incidents.
Mrs Marketa Reeves
Program Director, Australian Child And Youth Wellbeing Atlas
University Of Western Australia
AI-Enabled Infrastructure for Public Health Insight: The Child and Youth Wellbeing Atlas
Abstract
Background and Aim
Public health decision-making is increasingly constrained not by a lack of data, but by fragmented access, limited analytical capacity, and barriers to translating complex evidence into actionable insights. The Australian Child and Youth Wellbeing Atlas (the Atlas) addresses this by integrating multi-domain, place-based data to support evidence-informed policy and practice. However, as demand grows, enabling infrastructure is needed to improve discoverability, usability, and analytical depth. This presentation will demonstrate how Generative AI (GenAI) can enhance public health data infrastructure through improved navigation, insight generation, and scalable access to integrated datasets.
Methods and Analysis
We draw on the Atlas as a national data platform (www.australianchildatlas.com.au), alongside an in-progress GenAI integration that focuses on two key capabilities: (1) AI-supported filtering and navigation to facilitate access to complex, multi-domain indicators, and (2) AI-enabled analysis and insight generation to identify patterns and trends across geographies and populations. The development adopts a low-risk approach by limiting GenAI use to clearly defined tasks and analytical contexts, thereby reducing the likelihood of hallucinated outputs. Additional safeguards are provided through a controlled model platform with oversight and governance mechanisms to support responsible and trustworthy AI use.
Outcomes
The Atlas has already demonstrated impact through real-world use cases, including informing place-based advocacy and planning, supporting service design, and enabling cross-sector collaboration. Integrating GenAI is expected to amplify these impacts by reducing technical barriers, enabling rapid synthesis of evidence, and supporting more equitable access to insights. It can also help prioritise interventions by identifying actions most likely to improve outcomes, while keeping human oversight.
Conclusion and Future Actions
AI-enabled infrastructure offers significant opportunity to maximise the benefits of public health data. Future work will further expand GenAI, strengthen governance, and evaluate impacts on decision-making and equity. This work highlights the need for trusted, scalable and accessible infrastructure to support responsible AI adoption in public health.
Public health decision-making is increasingly constrained not by a lack of data, but by fragmented access, limited analytical capacity, and barriers to translating complex evidence into actionable insights. The Australian Child and Youth Wellbeing Atlas (the Atlas) addresses this by integrating multi-domain, place-based data to support evidence-informed policy and practice. However, as demand grows, enabling infrastructure is needed to improve discoverability, usability, and analytical depth. This presentation will demonstrate how Generative AI (GenAI) can enhance public health data infrastructure through improved navigation, insight generation, and scalable access to integrated datasets.
Methods and Analysis
We draw on the Atlas as a national data platform (www.australianchildatlas.com.au), alongside an in-progress GenAI integration that focuses on two key capabilities: (1) AI-supported filtering and navigation to facilitate access to complex, multi-domain indicators, and (2) AI-enabled analysis and insight generation to identify patterns and trends across geographies and populations. The development adopts a low-risk approach by limiting GenAI use to clearly defined tasks and analytical contexts, thereby reducing the likelihood of hallucinated outputs. Additional safeguards are provided through a controlled model platform with oversight and governance mechanisms to support responsible and trustworthy AI use.
Outcomes
The Atlas has already demonstrated impact through real-world use cases, including informing place-based advocacy and planning, supporting service design, and enabling cross-sector collaboration. Integrating GenAI is expected to amplify these impacts by reducing technical barriers, enabling rapid synthesis of evidence, and supporting more equitable access to insights. It can also help prioritise interventions by identifying actions most likely to improve outcomes, while keeping human oversight.
Conclusion and Future Actions
AI-enabled infrastructure offers significant opportunity to maximise the benefits of public health data. Future work will further expand GenAI, strengthen governance, and evaluate impacts on decision-making and equity. This work highlights the need for trusted, scalable and accessible infrastructure to support responsible AI adoption in public health.
Biography
Marketa Reeves (M.A.) is a Program Director and policy specialist based at the School of Population and Public Health at the University of Western Australia. With over 20 years’ experience across government, academia, and cultural institutions, her work focuses on the use of data to inform public health policy and improve outcomes for children and young people. She currently leads the Australian Child and Youth Wellbeing Atlas, a nationally significant data platform supporting evidence-informed decision-making. Marketa has led major research and engagement initiatives, including large-scale consultations and survey programs, and brings expertise in translating complex data into actionable insights. Her skills span quantitative and qualitative research, stakeholder engagement, and strategic planning. She is committed to advancing equitable, evidence-based approaches in public health and policy through accessible and impactful data systems.
Ms Edwina Mead
Phd Candidate
Monash University
Harnessing LLMs safely: reducing evidence synthesis bottlenecks in maternity care
Abstract
Background and Aim
Evidence-informed, equitable public health policy requires rapid synthesis of global literature. However, exponential growth in scientific publishing means that traditional methods of evidence synthesis are growing increasingly difficult, delaying translation into preventative policy. This is already apparent in maternal and child health, where a comprehensive review comparing long-term outcomes following Caesarean section or vaginal birth is impractical due to large volumes of evidence. This presentation outlines a project harnessing AI to safely accelerate article screening for a large-scale scoping review while minimising automation biases.
Methods and Analysis
Combining software engineering techniques with public health research, we developed a hybrid pipeline that decouples LLM data extraction from study inclusion decisions. This method eliminates extensive prompt engineering, instead using a simple prompt to extract 12 structured variables using Google Gemini. This was followed by a seven-stage rule-based pipeline that mirrors familiar, transparent text-based database searching to maintain human control. The workflow processed 51,281 abstracts and was validated against a large-scale manual audit.
Outcomes
The pipeline reduced the manual screening burden by 96% with an API cost of US$458.73. It achieved an estimated sensitivity of 89.9% (95% CI: 87.3–92.1%), specificity of 96.4% (95% CI: 96.2–96.5%), and a very low false-negative rate of just 0.18% (95% CI: 0.09–0.35%).
Conclusion and Future Actions
Our approach offers a potential way forward for developing living systematic reviews in rapidly evolving clinical domains. The method’s low cost, processing speed, and avoidance of opaque decision-making by the LLM provides an accessible framework that public health researchers can trust and control. Shortening the research-to-policy window improves the likelihood of health guidelines remaining robust, current, and community-trusted.
Evidence-informed, equitable public health policy requires rapid synthesis of global literature. However, exponential growth in scientific publishing means that traditional methods of evidence synthesis are growing increasingly difficult, delaying translation into preventative policy. This is already apparent in maternal and child health, where a comprehensive review comparing long-term outcomes following Caesarean section or vaginal birth is impractical due to large volumes of evidence. This presentation outlines a project harnessing AI to safely accelerate article screening for a large-scale scoping review while minimising automation biases.
Methods and Analysis
Combining software engineering techniques with public health research, we developed a hybrid pipeline that decouples LLM data extraction from study inclusion decisions. This method eliminates extensive prompt engineering, instead using a simple prompt to extract 12 structured variables using Google Gemini. This was followed by a seven-stage rule-based pipeline that mirrors familiar, transparent text-based database searching to maintain human control. The workflow processed 51,281 abstracts and was validated against a large-scale manual audit.
Outcomes
The pipeline reduced the manual screening burden by 96% with an API cost of US$458.73. It achieved an estimated sensitivity of 89.9% (95% CI: 87.3–92.1%), specificity of 96.4% (95% CI: 96.2–96.5%), and a very low false-negative rate of just 0.18% (95% CI: 0.09–0.35%).
Conclusion and Future Actions
Our approach offers a potential way forward for developing living systematic reviews in rapidly evolving clinical domains. The method’s low cost, processing speed, and avoidance of opaque decision-making by the LLM provides an accessible framework that public health researchers can trust and control. Shortening the research-to-policy window improves the likelihood of health guidelines remaining robust, current, and community-trusted.
Biography
Edwina Mead BEng (Hons), MPH, MGH
Edwina is a PhD candidate at the Monash Centre for Health Research and Implementation (MCHRI). Her research focuses on evaluating the long-term health and economic implications of current maternity care practices, aiming to ensure that guidelines are built on robust, timely, and equitable evidence. Drawing on a former career in software engineering, Edwina uses programmatic modelling and innovative data science workflows at a scale rarely accessible through traditional research methods.
A/Prof Son Nghiem
Principal Research Fellow
The University Of Queensland
Modern Evidence Synthesis with AI-Supported Systematic Reviews
Abstract
Introductions
Systematic reviews are the cornerstone of evidence-based public health, yet conventional methods are slow and labour-intensive, often taking months to years. This delay creates a persistent gap between emerging evidence and policy, slowing public health responses when timeliness matters most. We developed the ReviewFilter (https://sonnghiem.shinyapps.io/ReviewFilter/), an AI-supported application that streamlines the systematic review workflow while keeping researchers in control, aiming to close this evidence-to-policy gap.
Methods
ReviewFilter integrates the full review pipeline within a single R Shiny application: it generates structured search strings from a plain-language question using PICO synonyms (population, intervention, comparator, outcomes), queries bibliographic databases, screens records against inclusion and exclusion criteria, and supports data extraction and meta-analysis. Users connect their preferred large language model (e.g., Claude, OpenAI, Gemini, or a locally run model via Ollama) through their own API keys, the local option allowing sensitive data to remain on-device. Unlike existing tools (Rayyan, Covidence, Elicit) that highlight eligibility cues for manual decisions, ReviewFilter integrates search and filter records directly, with human oversight and override retained at every step.
Results
In applied reviews, the tool reduced screening time from months to days while achieving approximately 90% precision against human reviewer decisions. A single plain-language request. For example, a simple request of "primary studies on applications of AI in public health interventions", generated a PICO-based search returning 9,995 records from PubMed. The app merges the direct database and saved records and filters them to a manageable set. The app is publicly available and adheres to open-science principles (FAIR data, open-access outputs).
Discussion
ReviewFilter shows that AI can accelerate evidence synthesis without surrendering rigour or human judgement. By embedding transparency, human control, and open standards, it illustrates how AI can be harnessed responsibly, maximising benefits while minimising harms. Future work will expand database coverage and formal validation.
Systematic reviews are the cornerstone of evidence-based public health, yet conventional methods are slow and labour-intensive, often taking months to years. This delay creates a persistent gap between emerging evidence and policy, slowing public health responses when timeliness matters most. We developed the ReviewFilter (https://sonnghiem.shinyapps.io/ReviewFilter/), an AI-supported application that streamlines the systematic review workflow while keeping researchers in control, aiming to close this evidence-to-policy gap.
Methods
ReviewFilter integrates the full review pipeline within a single R Shiny application: it generates structured search strings from a plain-language question using PICO synonyms (population, intervention, comparator, outcomes), queries bibliographic databases, screens records against inclusion and exclusion criteria, and supports data extraction and meta-analysis. Users connect their preferred large language model (e.g., Claude, OpenAI, Gemini, or a locally run model via Ollama) through their own API keys, the local option allowing sensitive data to remain on-device. Unlike existing tools (Rayyan, Covidence, Elicit) that highlight eligibility cues for manual decisions, ReviewFilter integrates search and filter records directly, with human oversight and override retained at every step.
Results
In applied reviews, the tool reduced screening time from months to days while achieving approximately 90% precision against human reviewer decisions. A single plain-language request. For example, a simple request of "primary studies on applications of AI in public health interventions", generated a PICO-based search returning 9,995 records from PubMed. The app merges the direct database and saved records and filters them to a manageable set. The app is publicly available and adheres to open-science principles (FAIR data, open-access outputs).
Discussion
ReviewFilter shows that AI can accelerate evidence synthesis without surrendering rigour or human judgement. By embedding transparency, human control, and open standards, it illustrates how AI can be harnessed responsibly, maximising benefits while minimising harms. Future work will expand database coverage and formal validation.
Biography
Son Nghiem is the Head of the Health Economics Research and Modeling Unit at the School of Public Health, University of Queensland. He has more than 18 years of experience in health economics, applied econometrics, and health services research, with expertise spanning cost-effectiveness analysis, economic evaluation, and disease progression modelling. His research addresses critical questions at the intersection of health economics, health services, and public health, including patient safety and hospital-acquired complications, cardiovascular disease, climate-sensitive infectious diseases, and health technology assessment. He has an outstanding publication record of more than 150 peer-reviewed articles, over 4,000 citations, and an H-index of 38. He has been awarded 18 competitive grants totaling more than $20M, including 10 as CI-A. CI Nghiem has extensive experience in health economic evaluations for infectious disease surveillance.