Figures
Abstract
Objective
Primary health care (PHC) is the cornerstone of health systems worldwide, yet faces growing pressures from aging populations, workforce shortages, and constrained resources. While artificial intelligence (AI) and machine learning (ML) are increasingly used for clinical tasks such as diagnostics and decision support, their potential to address organizational and system-level processes remains underexplored. This scoping review maps AI and ML applications targeting non-clinical, organizational, and system-level functions in PHC, examining techniques employed, application purposes, data sources, and implementation maturity, and identifying evidence gaps to inform future research.
Methods
This scoping review will be conducted using the Arksey and O’Malley framework and reported according to PRISMA-ScR. We will search Ovid MEDLINE, EBSCO CINAHL, Ovid Embase, Cochrane Library (Wiley), Ovid PsycINFO, and IEEE Xplore, as well as grey literature sources, from January 2010 to December 2025, with a pre-submission update. Two independent reviewers will screen titles, abstracts, and full texts. Eligible studies include empirical research describing, evaluating, implementing, or developing AI and ML applications at the meso-level (organizational, e.g., care scheduling, population stratification) and macro-level (system-wide, e.g., workforce planning, funding allocation) of PHC. Studies on micro-level clinical applications, non-empirical research, and secondary literature will be excluded. Applications will be classified using a two-layer taxonomy (AI technique family and core algorithm) and mapped to PHC building blocks adapted from WHO operational frameworks. Synthesis will examine technique-building block intersections, application purposes, data sources, implementation maturity, and evidence gaps. The protocol is registered in Open Science Framework (osf.io/wzj5x).
Conclusions
This review will provide a timely synthesis of AI and ML applications at organizational and system levels of PHC, identifying which building blocks have been addressed, which remain underexplored, and where the field stands in terms of implementation readiness. The findings will inform a targeted research agenda and provide actionable evidence for decision-makers seeking to leverage AI and ML for PHC strengthening.
Citation: Galvez-Hernandez P, Audet L-A, Shakeri Z, Wodchis WP (2026) Artificial intelligence and machine learning for organizational and system-level functions in primary health care: A scoping review protocol. PLoS One 21(9): e0329426. https://doi.org/10.1371/journal.pone.0329426
Editor: Mohammad Saadati, Khoy University of Medical Sciences, IRAN, ISLAMIC REPUBLIC OF
Received: July 16, 2025; Accepted: August 28, 2026; Published: September 25, 2026
Copyright: © 2026 Galvez-Hernandez et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All relevant data are within the manuscript and its Supporting information files.
Funding: PGH’s and LA’s time was supported by a Postdoctoral Fellowship at the University of Toronto, Institute of Health Policy, Management and Evaluation, funded by the Canadian Institutes of Health Research (CIHR) (grant #513679) through the Optimizing Teams for Interprofessional Care in Primary Health Care (OPTIC-PHC) project, awarded under the CIHR Project Grant program to WW. PGH also received support from the Artificial Intelligence for Public Health (AI4PH) Research Training Platform. The funders had no role in the study design, data collection and analysis, decision to publish, or preparation of the manuscript. No additional external funding was received for this study.
Competing interests: The authors have declared that no competing interests exist.
Background
Applications of Artificial Intelligence (AI) and Machine Learning (ML) in healthcare have grown exponentially over the past decade [1]. AI can be defined as a broad field focused on systems that exhibit intelligent behavior by deriving knowledge from data [2]. Machine learning (ML), commonly considered a subfield of AI, uses algorithms to learn associations with predictive power from data examples through statistical models [2]. These tools enable the analysis of large, complex data to uncover patterns and predict outcomes, offering promising advancements in how healthcare is delivered.
Most studies on AI and ML in healthcare have focused on their clinical applications, including diagnostic tools, chatbots, and predictive analytics embedded in healthcare IT infrastructure for patient care delivery [3,4]. Additionally, there is increasing interest in population-based applications in public health, often referred to as precision population health management, which uses advanced analytics (e.g., AI, machine learning, and predictive modeling) to stratify populations and tailor interventions based on clinical, behavioral, and social data, aiming to deliver targeted, equitable care at scale [5]. Numerous empirical data have shown that AI and ML can assist clinicians to make better decisions, early diagnosis and patient monitoring [6,7]. These benefits have also been synthesized in systematic reviews, which highlighted the advantages of AI and ML for patients, families, and healthcare professionals in various healthcare contexts [8,9].
From a systems thinking perspective, effective health care delivery is not only a function of clinical decision-making but is fundamentally shaped by the structural and organizational processes that support it, including governance, financing, workforce deployment, and service planning [10]. These meso-level (organizational and team-level) and macro-level (regional, provincial, and national regulatory and funding structures) dimensions constitute the infrastructure upon which clinical care operates [11]. While AI and ML have demonstrated transformative potential for clinical tasks, the capacity of these technologies to address the organizational and system-level challenges that constrain health care delivery has received comparatively limited attention in the literature.
Within health systems, Primary Health Care (PHC), a cornerstone of efficient health systems and universal health coverage, faces persistent challenges in sustaining high-quality service delivery [12]. Examples of structural and process-related challenges include resource allocation favouring specialized care; difficulties in planning and funding PHC strategies that involve multisectoral collaboration; adapting care models to local contexts; inadequate workforce planning amid shortages; inefficiencies in payment and purchasing systems; and deploying integrated monitoring and evaluation frameworks [12,13]. These challenges are compounded by an aging population with complex multimorbidity and growing social needs, placing additional strain on PHC systems [14]. While economic and workforce resource constraints present a major burden, existing infrastructures could be strengthened through more effective resource allocation, streamlined care pathways, needs-based service and workforce planning, and better integration supported by timely data [15,16]. AI and ML might have the potential to support these improvements; however, their role in addressing these challenges remains unclear and has not been comprehensively synthesized in existing reviews.
Some empirical research has explored the application of AI and ML beyond clinical encounters in primary care, including population-level risk stratification [17], patient phenotyping and panel size optimization to improve provider workload distribution [18], and appointment scheduling [19]. However, a comprehensive understanding of AI and ML applications across the broader organizational and system-level dimensions of primary care remains absent, as existing reviews have focused predominantly on micro-level applications involving individual patients, providers, and their clinical interactions. For example, Abbasgholizadeh et al. mapped AI applications in community-based PHC and found that diagnosis, risk stratification, disease surveillance, and clinical decision support were the dominant use cases, relying primarily on machine learning and natural language processing [8]. These existing syntheses confirm the dominance of micro-level clinical applications in the PHC AI/ML literature and underscore the need of reviews addressing the organizational and system-level domains targeted by this protocol. While AI techniques are increasingly being applied to structural challenges such as strategic purchasing and funding allocation in health systems [20], evidence of such applications within the PHC context specifically remains limited.
Addressing this gap is essential for understanding how AI and ML may support core PHC system functions, including governance, resource allocation, workforce planning, and intersectoral coordination. This scoping review aims to: (i) map and characterize current AI and ML applications in organizational and system-level dimensions of PHC; (ii) identify the type of data sources and methods used to develop and apply AI/ML models; (iii) examine reported outcomes, strengths, and limitations; and (iv) identify evidence gaps and future directions.
Methods
This scoping review follows the five stages proposed by Arksey and O’Malley: (i) identifying the research questions; (ii) identifying relevant studies; (iii) study selection; (iv) charting the data; and (v) collating, summarizing and reporting the results [21]. A scoping review was selected because it is designed to map the breadth of a heterogeneous evidence base, identify key concepts, and clarify the boundaries of an emerging research area [21]. Results will be reported according to the PRISMA Extension for Scoping Reviews (PRISMA-ScR) checklist [22]. The analytical approach integrates a classification taxonomy for AI/ML techniques that captures the type of methods employed and a set of PHC building blocks adapted from WHO operational and measurement frameworks [12,15], which will enable a systematic cross-mapping of techniques to PHC building blocks, which constitutes the primary analytical output of this review. This protocol has been registered in Open Science Framework (OSF) (osf.io/wzj5x). In addition, a preprint version was posted on medRxiv through the journal’s facilitated preprint posting service during the manuscript submission process. Ethics approval was not required for this study because human subjects are not involved.
Scoping review stages
Stage 1: Defining the research questions.
The review will be guided by the following research questions: (i) What are the current applications of AI and ML at the organizational and system-levels of PHC as reported in the literature? (ii) What types of data sources (e.g., electronic health records, administrative databases, national registries, survey data) have been used to develop and apply AI and ML models in this context? (iii) What are the strengths and limitations of these applications as reported by study authors? (iv) What gaps exist in the current evidence base in terms of underused AI and ML techniques across PHC building blocks, and implementation maturity?
Stage 2: Identifying relevant studies.
Studies will be searched in Ovid MEDLINE, EBSCO CINAHL, Ovid Embase, Cochrane Library (Wiley), Ovid PsycINFO, and IEEE Xplore. A grey literature search will be conducted in OpenGrey, Google Scholar, and ProQuest Dissertations & Theses Global. Reference lists of included articles and relevant literature reviews will also be searched to identify additional studies [21]. The search will be restricted to English-language publications from January 2010 to December 2025, with a final search update conducted prior to manuscript submission, to ensure the review reflects the most current evidence available at the time of publication. The start date was chosen to capture the period of rapid expansion of AI and ML in healthcare, marked by the widespread adoption of deep learning techniques in the early 2010s [23].
A search strategy was iteratively developed following guidance from the Joanna Briggs Institute [24] and the Cochrane Handbook for Systematic Reviews of Interventions [25]. The strategy was informed by preliminary searches, comparison with search strategies from relevant published systematic and scoping reviews on AI/ML in health care, and term analysis of five relevant articles identified through an initial limited search in MEDLINE (Ovid) using the terms “artificial intelligence,” “machine learning,” and “primary health care.” Titles, abstracts, keywords, and index terms from these articles were examined using the Yale MeSH Term Analyzer [26] to identify additional controlled vocabulary and free-text terms. The resulting strategy combined keywords and index terms for two core concepts, AI/ML and primary health care, using Boolean operators (Table 1). It was then reviewed for sensitivity during consultation with an expert librarian at the University of Toronto and adapted to the syntax and controlled vocabulary of each database. The five hand-searched articles were used as a known-item check to confirm retrieval of relevant studies. The term “data mining” was considered but not included as a separate term because of substantial overlap with machine learning.
Stage 3: Study selection.
All references will be uploaded to Rayyan software for deduplication and screening. Two independent reviewers (PGH, LA) will conduct title, abstract, and full text screening using a detailed protocol outlining the eligibility criteria. To validate the screening process, both reviewers will first simultaneously screen a sample of 25 articles as a training exercise [21]. A calibration exercise will then be conducted, with both reviewers independently screening 20% of articles at each stage, targeting an inter-rater reliability above 80% [27]. Discrepancies will be resolved through periodic meetings to discuss decision rationales, address inconsistencies, and reach consensus, with the involvement of a third reviewer as tie-breaker when necessary.
The eligibility criteria are defined a priori, drawing from the conceptual definitions in Table 2. Inclusion criteria are: (i) published and unpublished primary studies that (ii) describe, evaluate, implement, or develop AI/ML applications (iii) at the organizational and system-levels of PHC systems; for example, clinic management, workforce planning, primary care policy analysis, or resource allocation across services. Unpublished primary studies include dissertations, theses, conference papers and proceedings, and reports. AI methods in this review primarily refer to computational techniques that enable learning of data-driven decision-making; knowledge-driven AI (e.g., expert systems) is also included where it appears in the context of PHC systems. All categories and techniques of AI and ML are eligible (e.g., supervised, unsupervised, semi-supervised, reinforcement learning, deep learning, natural language processing).
Exclusion criteria are: (i) non-empirical publications (e.g., commentaries, editorials); (ii) secondary studies such as literature reviews and meta-analyses; (iii) studies reporting micro-level AI and ML applications in PHC (e.g., diagnostic tools, patient-facing digital technologies); (iv) studies conducted in settings other than PHC, such as clinical, public health, or epidemiological research focused solely on the efficacy of technologies for individual patients or on measuring population health profiles.
Stage 4: Charting the data.
The data charting process will involve a training phase, a reliability check, and full extraction. A pilot test will be conducted on 10% of included studies to refine the charting form prior to full extraction. One reviewer (PGH) will chart the data, which will be independently verified by a second reviewer (LA). Discrepancies will be resolved through discussion or by a third reviewer if necessary.
The following information will be charted for each included study across four domains: (i) Study characteristics: title, author(s), year of publication, country, source, study design, study objectives and funding source. (ii) AI/ML application: AI technique family (Traditional ML, deep learning, NLP, or other AI/Hybrid); core algorithm(s) (the specific models used for the main analytical task, e.g., random forest, logistic regression); ML learning approach (supervised, unsupervised, semi-supervised, or reinforcement); application purpose (prediction, classification, clustering/segmentation, optimization, monitoring/surveillance, decision support, or other); and primary data type (structured EHR/claims, unstructured text, administrative data, survey data, images, or mixed). (iii) Primary system context: PHC system level (meso, macro, or both); type of setting (e.g., primary care centers, community health centers); target population; and PHC building block, as defined in Table 3. (iv) Outcomes and appraisal: implementation stage (development/validation only, pilot tested, or fully deployed); main findings; and strengths and limitations of the AI/ML application as reported by the study authors.
Stage 5: Collating, summarizing, and reporting the results.
Extracted data will be coded and categorized according to the classification framework defined in Table 2 and mapped to the PHC building blocks defined in Table 3. To organize the mapping of AI/ML applications to non-clinical PHC functions, we defined seven organizational and system-level building blocks (Table 3), adapted from two WHO frameworks: the Operational Framework for Primary Health Care [12] and the Primary Health Care Measurement Framework [15]. These frameworks identify structural and process-level levers for strengthening PHC systems, including workforce deployment, financing, service planning, quality improvement, digital health infrastructure, and community engagement. While the original frameworks encompass a broader set of PHC dimensions, we consolidated those pertaining to organizational and system-level functions into seven building blocks that capture the non-clinical domains to which AI/ML could be applied. Each study will be assigned one primary building block based on the stated purpose of the AI/ML application, guided by the question: “What PHC function does this application directly improve or enable?” Downstream or indirect effects will not be used for assignment.
Using this framework, extracted data will be synthesized in three components: First, a descriptive overview will present temporal trends, geographic distribution, publication venues, PHC settings, and system levels. Second, the landscape analysis will cross-tabulate AI technique families against PHC building blocks, presented as a heat map in which cell size or color intensity reflects the number of studies and empty cells identify research gaps. Application purposes will be mapped to building blocks with illustrative examples from included studies. Third, a methodological quality snapshot will summarize across all included studies the proportion reporting a validation strategy, performance metrics, external validation, and use of publicly available data or code. Findings will be presented narratively, organized by PHC building block, and supported by figures and summary tables. Author-reported strengths and limitations will be compiled, categorized by building block, and analyzed thematically. A dedicated gap analysis section will identify underexplored PHC building block-technique intersections and inform a targeted research agenda. The review process will be documented using the PRISMA-ScR flow diagram and the PRISMA-P checklist (S1 File). Consistent with the aims of a scoping review of mapping the breadth and nature of a research area rather than evaluate intervention effectiveness, a formal risk of bias assessment will not be conducted [21].
Discussion
This scoping review protocol outlines a structured approach to mapping AI and ML applications at the organizational and system levels of PHC. The review has completed Stage 3 (study selection), and completion of the full scoping review is expected by July 2026. By synthesizing evidence from published and unpublished literature across six databases and grey literature, the review will characterize how AI and ML have been applied to non-clinical PHC functions, including workforce planning, resource allocation, service demand estimation, quality monitoring, and community engagement, and identify where evidence is accumulating and where significant gaps remain. This focus extends beyond the predominantly micro-level clinical applications explored in previous reviews [8,19], addressing a largely uncharted area in health services and policy research.
The review is anticipated to make three contributions. First, the cross-tabulation of AI/ML technique families against PHC building blocks will produce the first structured map of technique-area intersections at the organizational and system levels of PHC, identifying both populated and empty cells that can directly inform a targeted research agenda. Second, by classifying implementation stage and methodological approaches, the review will characterize the field’s readiness for real-world deployment, which has direct relevance for health system decision-makers weighing AI/ML investments. Third, the building blocks taxonomy and classification framework developed for this review may serve as reusable tools for future syntheses in this domain.
Some limitations are anticipated in this scoping review. The restriction to English-language publications may exclude relevant research from regions where PHC literature is published in local languages, particularly in Latin America, East Asia, and Francophone Africa, where primary care systems may be organized differently and where AI/ML applications to system-level challenges could yield distinct insights. Additionally, the use of specific “primary health care” terminology may limit the capture of studies from contexts where equivalent first-level care services use different nomenclature. Additionally, although the search strategy was developed iteratively with librarian consultation and informed by published AI/ML search strategies, the absence of a librarian or information specialist as a co-author may have limited the comprehensiveness and reporting quality of the search methods, as librarian involvement has been associated with higher-quality search reporting in knowledge synthesis studies [32]. Finally, this scoping review will prioritize breadth of mapping over depth of outcome evaluation to identify what exists and where gaps lie, but will not assess the effectiveness of individual AI/ML applications.
Despite these limitations, this review will provide a comprehensive and timely synthesis of AI and ML applications at understudied levels of PHC. The findings will contribute to health services and policy research by identifying evidence gaps, characterizing methodological patterns, and informing recommendations to strengthen the organizational and structural capacity of PHC systems.
Acknowledgments
The authors would like to thank the Faculty Liaison & Instruction Librarian at University of Toronto for their input when designing and validating the search strategy.
References
- 1.
Panesar A. Machine learning and AI for healthcare. Coventry, UK: Apress; 2019. 10.1007/978-1-4842-3799-1
- 2. Panch T, Szolovits P, Atun R. Artificial intelligence, machine learning and health systems. J Glob Health. 2018;8(2):020303. pmid:30405904
- 3. Abdulazeem H, Whitelaw S, Schauberger G, Klug SJ. A systematic review of clinical health conditions predicted by machine learning diagnostic and prognostic models trained or validated using real-world primary health care data. PLoS One. 2023;18(9):e0274276. pmid:37682909
- 4. Bory C, Schmutte T, Davidson L, Plant R. Predictive modeling of service discontinuation in transitional age youth with recent behavioral health service use. Health Serv Res. 2022;57(1):152–8. pmid:34396526
- 5. Han A, Isaacson A, Muennig P. The promise of big data for precision population health management in the US. Public Health. 2020;185:110–6. pmid:32615477
- 6. Altini N, Rossini M, Turkevi-Nagy S, Pesce F, Pontrelli P, Prencipe B, et al. Performance and limitations of a supervised deep learning approach for the histopathological Oxford Classification of glomeruli with IgA nephropathy. Comput Methods Programs Biomed. 2023;242:107814. pmid:37722311
- 7. Bachtiger P, Petri CF, Scott FE, Ri Park S, Kelshiker MA, Sahemey HK, et al. Point-of-care screening for heart failure with reduced ejection fraction using artificial intelligence during ECG-enabled stethoscope examination in London, UK: a prospective, observational, multicentre study. Lancet Digit Health. 2022;4(2):e117–25. pmid:34998740
- 8. Abbasgholizadeh Rahimi S, Légaré F, Sharma G, Archambault P, Zomahoun HTV, Chandavong S, et al. Application of artificial intelligence in community-based primary health care: systematic scoping review and critical appraisal. J Med Internet Res. 2021;23(9):e29839. pmid:34477556
- 9. Bhatt P, Liu J, Gong Y, Wang J, Guo Y. Emerging artificial intelligence-empowered mHealth: scoping review. JMIR Mhealth Uhealth. 2022;10(6):e35053. pmid:35679107
- 10.
World Health Organization. Monitoring the building blocks of health systems: a handbook of indicators and their measurement strategies. Geneva: World Health Organization; 2010.
- 11. Sheikh K, Gilson L, Agyepong IA, Hanson K, Ssengooba F, Bennett S. Building the field of health policy and systems research: framing the questions. PLoS Med. 2011;8(8):e1001073. pmid:21857809
- 12.
World Health Organization, United Nations Children’s Fund. Operational framework for primary health care: transforming vision into action. Geneva: World Health Organization; 2020.
- 13. Galvez-Hernandez P, Shankardass K, Puts M, Tourangeau A, Gonzalez-de Paz L, Gonzalez-Viana A, et al. Mobilizing community health assets through intersectoral collaboration for social connection: associations with social support and well-being in a nationwide population-based study in Catalonia. PLoS One. 2025;20(3):e0320317. pmid:40138367
- 14. Organisation for Economic Co-operation and Development. Health at a glance 2023: OECD indicators. Paris: OECD Publishing; 2023.
- 15.
World Health Organization, United Nations Children’s Fund. Primary health care measurement framework and indicators: monitoring health systems through a primary health care lens. Geneva: World Health Organization; 2022.
- 16. Galvez-Hernandez P, Gonzalez-Viana A, Gonzalez-de Paz L, Shankardass K, Muntaner C. Generating contextual variables from web-based data for health research: tutorial on web scraping, text mining, and spatial overlay analysis. JMIR Public Health Surveill. 2024;10:e50379. pmid:38190245
- 17. Orlando LA, Wu RR, Myers RA, Neuner J, McCarty C, Haller IV, et al. At the intersection of precision medicine and population health: an implementation-effectiveness study of family health history based systematic risk assessment in primary care. BMC Health Serv Res. 2020;20(1):1015. pmid:33160339
- 18. Rajkomar A, Yim JWL, Grumbach K, Parekh A. Weighting primary care patient panel size: a novel electronic health record-derived measure using machine learning. JMIR Med Inform. 2016;4(4):e29. pmid:27742603
- 19. Sørensen NL, Bemman B, Jensen MB, Moeslund TB, Thomsen JL. Machine learning in general practice: scoping review of administrative task support and automation. BMC Prim Care. 2023;24(1):14. pmid:36641467
- 20. Ramezani M, Takian A, Bakhtiari A, Rabiee HR, Fazaeli AA, Sazgarnejad S. The application of artificial intelligence in health financing: a scoping review. Cost Eff Resour Alloc. 2023;21(1):83. pmid:37932778
- 21. Arksey H, O’Malley L. Scoping studies: towards a methodological framework. Int J Soc Res Methodol. 2005;8(1):19–32.
- 22. Tricco AC, Lillie E, Zarin W, O’Brien KK, Colquhoun H, Levac D, et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): checklist and explanation. Ann Intern Med. 2018;169(7):467–73. pmid:30178033
- 23. Jiang F, Jiang Y, Zhi H, Dong Y, Li H, Ma S, et al. Artificial intelligence in healthcare: past, present and future. Stroke Vasc Neurol. 2017;2(4):230–43. pmid:29507784
- 24. Peters MDJ, Godfrey CM, Khalil H, McInerney P, Parker D, Soares CB. Guidance for conducting systematic scoping reviews. Int J Evid Based Healthc. 2015;13(3):141–6. pmid:26134548
- 25.
Lefebvre C, Glanville J, Briscoe S, Featherstone R, Littlewood A, Metzendorf MI, et al. Cochrane Handbook for Systematic Reviews of Interventions version 6.5.1 Cochrane, 2025. Available from cochrane.org/handbook.
- 26. Yale University. Yale mesh analyzer. Cushing/Whitney Medical Library; 2015. Available from: https://mesh.med.yale.edu/
- 27. Belur J, Tompson L, Thornton A, Simon M. Interrater reliability in systematic review methodology: exploring variation in coder decision-making. Sociol Methods Res. 2018;50(2):837–65.
- 28.
Santosh KC, Gaur L. Artificial intelligence and machine learning in public healthcare: opportunities and societal impact. Cham: Springer Nature; 2022.
- 29. Doupe P, Faghmous J, Basu S. Machine learning for health services researchers. Value Health. 2019;22(7):808–15. pmid:31277828
- 30. Muldoon LK, Hogg WE, Levitt M. Primary care (PC) and primary health care (PHC): what is the difference? Can J Public Health. 2006;97(5):409–11.
- 31. Sawatzky R, Kwon JY, Barclay R, Chauhan C, Frank L, van den Hout WB, et al. Implications of response shift for micro-, meso-, and macro-level healthcare decision-making using results of patient-reported outcome measures. Qual Life Res. 2021;30(12):3343–57.
- 32. Pawliuk C, Cheng S, Zheng A, Siden HH. Librarian involvement in systematic reviews was associated with higher quality of reported search methods: a cross-sectional survey. J Clin Epidemiol. 2024;166:111237. pmid:38072177