Figures
Abstract
Background
Clinical referral letters direct patients into acute, urgent, and specialist care pathways, but their triage is typically manual, heterogeneous, and resource intensive, contributing to waiting-list pressure and potential patient harm. Artificial intelligence (AI) methods, including natural language processing, machine learning, and large language models, offer potential to automate or augment referral triage, yet the performance, safety, and implementation characteristics of AI applied to referral documentation have not been systematically synthesised. A February 2026 PROSPERO search identified no completed or ongoing systematic reviews addressing this question.
Objectives
To evaluate the prioritisation performance of AI models used to triage clinical referral documentation for acute and specialist care pathways, and to map how this emerging field defines and evaluates AI-assisted referral triage, including model types, reference standards, validation practices, safety outcomes, and implementation-related outcomes.
Methods
This protocol is reported in accordance with PRISMA-P. Systematic searches will be undertaken in MEDLINE, EMBASE, Web of Science, Scopus, CINAHL, and the Cochrane Central Register of Controlled Trials, covering January 2016 to the 5 May 2026, with no language restrictions. Eligible studies will include diagnostic accuracy, model development and validation, and comparative studies that apply an AI model to referral documentation and report at least one triage-performance outcome. Two reviewers will independently screen, extract data, and assess risk of bias using PROBAST-AI as appropriate. A structured narrative synthesis and evidence map using the SWiM reporting guideline will be the primary synthesis output, stratified by AI model class and clinical pathway. Meta-analysis will be treated as conditional and exploratory, and will be undertaken only where studies report comparable outcomes using a similar model class, clinical setting, reference standard, outcome definition, and threshold structure.
Expected outcomes and significance
This review will produce the first systematic synthesis and evidence map of AI-assisted referral triage across acute and specialist care pathways, characterising how the field defines and evaluates this task as well as what is currently known about prioritisation performance. Findings will inform clinicians, service planners, and policymakers considering AI adoption in referral workflows, and will identify methodological priorities for future research.
Systematic review registration: PROSPERO CRD420251244654.
Citation: Foley J, Hession E, Umana E, Boland F (2026) The use of artificial intelligence models for interpretation and triage of clinical referral letters to acute and specialist care pathways: Protocol for a systematic review. PLoS One 21(8): e0355476. https://doi.org/10.1371/journal.pone.0355476
Editor: Mergan Naidoo, University of Kwazulu-Natal, SOUTH AFRICA
Received: May 11, 2026; Accepted: July 22, 2026; Published: August 6, 2026
Copyright: © 2026 Foley et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: No datasets were generated or analysed during the current study. All relevant data from this study will be made available upon study completion.
Funding: The author(s) received no specific funding for this work.
Competing interests: The authors have declared that no competing interests exist.
Introduction
The problem
Referral letters and clinical documentation are central to directing patients into acute healthcare. Across health systems, patients are referred from primary care, emergency settings, community services, or other clinical teams into acute, elective, urgent, or tertiary-level services. These referrals rely on clinical documentation, including referral letters, clinical notes, clinical summaries, or electronic referral forms, which all vary considerably in quality, structure, and completeness [1–3]. The triage of these referrals determines urgency, waiting times, and access to care, and is a major contributor to service delays across acute and specialist pathways [4–6].
The triage and prioritisation of referrals is frequently slow, inconsistent, and resource intensive [7]. In many health systems, this process is performed manually by clinicians or experienced nursing staff, who must integrate heterogeneous and often incomplete information to assign clinical priority [8–10]. This can be challenging, time-consuming, and cause frustration to the individuals triaging the referrals, and subsequent delays to the patient for whom the referral was sent. Under and over-triage of referral letters has been reported, and the subjectivity of the triage disposition often depends on the quality of the information in the letter, and the experience of the individual triaging [10].
As populations increase, the demand on health systems increases concurrently, and the volume of referrals that must be triaged becomes increasingly difficult with these challenges [7]. Waiting lists in healthcare are a focus of intent for health governing bodies internationally, and current triage systems to reduce these lists have no adopted universal framework, nor are they built to manage a rapidly rising referral burden. The triage and prioritisation of referral documentation across acute and specialist care represents an analogous challenge: a process that is slow, inconsistent, and reliant on manual clinical review of heterogenous documentation.
AI as a potential solution
Artificial intelligence (AI), in particular machine learning (ML), deep learning (DL), and natural language processing (NLP) have significant and rapidly expanding potential to support clinical decision-making. NLP systems can analyse free-text clinical documentation to extract clinically relevant features and classify patient urgency, if the system has been trained to a high standard [11,12]. ML models have demonstrated the ability to predict triage acuity, disposition, and clinical deterioration from structured patient data and early clinical observations in emergency settings while large language models (LLMs) can generate clinical reasoning based on the acquisition of extensive clinical knowledge, which represents a further development for AI-assisted triage from unstructured referral text [12–14]. Collectively, these approaches could offer a scalable, consistent, efficient and potentially safer adjunct or alternative to purely manual referral triage, with the capacity to reduce delays and improve prioritisation accuracy at scale.
Evidence gap
Despite this promise, the performance, safety, and clinical implementation characteristics of AI models for the triage of referral documentation have not been systematically synthesised. A search of PROSPERO conducted in February 2026 identified no completed or ongoing systematic reviews specifically evaluating AI models for interpretation and triage of clinical referral letters across acute and specialist care pathways. Existing reviews have addressed adjacent topics including AI for triage using physiological data and structured EHR inputs, AI for diagnostic support in specific clinical domains, and AI for the interpretation of diagnostic imaging and clinical notes in specific specialties but none has evaluated the specific intersection of free-text or semi-structured referral documentation and AI-assisted triage across the range of acute and specialist care pathways [15,16]. This gap limits the ability of clinicians, health service planners, and researchers to make evidence-informed decisions about the adoption of AI for referral triage.
Objectives
Primary objective.
To evaluate the prioritisation performance of AI models used to triage clinical referral documentation, including referral letters, clinical notes and electronic referral forms for acute and specialist care pathways, and to map how this emerging field defines and evaluates AI-assisted referral triage.
Prioritisation performance will be assessed as agreement with each study’s defined reference standard and interpreted in the context of the study’s clinical pathway, reference standard and level of urgency, rather than a transferable construct.
Secondary objectives.
- To map the evidence base for AI-assisted referral triage, including the types of referral documentation processed, the clinical pathways addressed, the reference standards adopted, and the reporting of calibration, internal validation, and external validation.
- To identify and describe the types of AI models (including machine learning, deep learning, and natural language processing) developed for referral triage, including their data sources, methods of development and validation, and the clinical areas in which they have been applied.
- To assess safety-related outcomes, with particular attention to under-triage rates, over-triage rates, and any reported adverse clinical outcomes associated with AI-based referral triage.
- To examine implementation-related outcomes, including efficiency, feasibility, interpretability, and acceptability of AI models used in referral triage contexts.
- To identify methodological gaps in the current evidence base relevant to the real-world adoption of AI for referral triage, and to highlight priorities for future research.
Methods
This protocol is reported in accordance with the Preferred Reporting Items for Systematic Review and Meta-Analysis Protocols (PRISMA-P) guidelines [17]. The completed systematic review will be conducted and reported in accordance with the PRISMA guidelines [18–20]. A comprehensive systematic search will be undertaken across relevant databases. All identified records will be imported into a reference management system, and duplicates will be removed.
Eligibility criteria
All studies that examine the use of an AI tool for triage of referral documentation will be considered against the inclusion and exclusion criteria outlined in Table 1.
Exclusion criteria.
Studies will be excluded if they,
- Do not involve referrals to acute or specialist care pathways (for example, administrative, social, or non-clinical referrals).
- Do not include a documentation component to a referral, or if the AI model was purely used for administrative purposes.
- Do not include a triage or prioritisation standard.
- Case reports, case series, editorials, opinion pieces, narrative reviews and conference abstracts without sufficient methodological detail (unless full text is available).
Information sources and search strategy
An electronic search strategy will be performed using MEDLINE, EMBASE, Web of Science, Scopus, CINAHL, and the Cochrane Central Register of Controlled Trials. The search strategy will be broad, and focus on four conceptual blocks, using Artificial Intelligence, Clinical Referral or Documentation Triage/Prioritisation and Acute/Specialist Care as key MeSH terms, exploded where available. There will be no language restrictions, papers not in English will be included and translated where necessary with translation processes documented in the review. To enhance completeness, we will perform backward and forward citation chaining of all included studies and relevant reviews. To capture emerging evidence in this rapidly developing field, we will additionally search Europe PMC for relevant preprints; these will be screened against the same eligibility criteria, flagged as non-peer-reviewed, and reported separately, and will not be pooled with peer-reviewed studies in any quantitative synthesis. The review will include studies published from January 2016–5 May 2026 (the date of the search execution), reflecting the emergence of modern LLM and deep learning methods in clinical applications and is consistent with the ten-year limit specified in the PROSPERO registration. An example search strategy is shown in Supplementary Material, Table 1 (S1 File).
Study records
Data management.
Search results will be imported into Covidence, where duplicate records will be identified and removed. Screening of titles, abstracts, and full texts will be conducted using Covidence. Data extraction will be completed using a piloted standardised extraction form in Covidence or Microsoft Excel. All data will be stored on a secure, access-restricted institutional drive with regular backups.
Selection process.
Two authors will independently screen the results of the search strategy, first by title and abstract, and then by examination of the full articles according to the inclusion criteria. Any unresolved discrepancy between these two authors will be resolved by the third author. A PRISMA 2020 compliant flow diagram will document the full screening process, from total records identified through each database, through deduplication, screening, full-text review, and final inclusion. Reasons for exclusion at the full-text stage will be reported. Where multiple publications report on the same model, dataset or cohort, the most methodologically rigorous manuscript will be designated the primary record, with additional reports cross-referenced and treated as supplementary to avoid double counting. Successive versions of the same model (e.g., GPT 3.5 versus GPT 4.0 evaluations) will be documented and reported separately in the synthesis.
Data extraction and data items
Two authors will extract data independently using a standardised extraction form developed within Covidence or Microsoft Excel. The form will be piloted on a random subset of three to five studies before formal extraction commences and refined as necessary. Discrepancies will be resolved through discussion, with a third reviewer consulted where consensus cannot be reached. All extraction decisions will be documented.
Missing data will be documented and reported as ‘not reported’ where applicable. Authors will be contacted for clarification where key data is missing. Analyses will be conducted using available data, and the potential impact of missing data on the findings will be considered in the interpretation of results.
The data items to be extracted, where possible, are presented in Table 2.
Risk of bias in individual studies
The risk of bias will be assessed independently by two reviewers. The following validated tools will be applied according to study design.
- AI Prediction Model Development and/or Validation: PROBAST-AI where applicable [21,22]
- Randomised Controlled Trials: Cochrane Risk of Bias 2 (ROB 2) [23]
- Non-randomised comparative studies: ROBINS-I (Risk of Bias in Non-Randomised Studies of Interventions) [24]
- Diagnostic accuracy studies evaluating an AI model against a clinical reference standard without explicit model development: QUADAS-2 ([25]
Where a study combines elements of multiple designs (for example, model development with concurrent prospective validation), the most appropriate tool will be selected, and the rationale documented. Disagreements in risk of bias judgements will be resolved by discussion; third-reviewer adjudication will be sought where consensus cannot be reached.
Data synthesis
Given the anticipated clinical, methodological, and statistical heterogeneity of included studies, results will be synthesised primarily using a structured narrative approach as the primary method, supplemented by quantitative synthesis where feasible. Metrics such as accuracy, sensitivity, specificity, and AUC may not be directly comparable across studies even when reported in a similar form, because their meaning depends on the reference standard, and the clinical consequence of over- and under-triage in each pathway. Accordingly, terms such as “performance” and “agreement” will be interpreted in relation to each study’s own reference standard and pathway, not as generalisable transferable constructs.
The review therefore has two complimentary synthesis aims. The first is a structured evidence map of the field, describing the types of AI models utilised, the types of referral documentation processed, the clinical pathway, and the reference standards assessed. The second is a synthesis of prioritisation performance, agreement with reference standards, and safety and implementation outcomes, if meaningful comparison can be achieved.
For both aims, a narrative synthesis will be conducted using the SWiM (Synthesis Without Meta-Analysis) reporting guideline, with structured summaries of studies tabulated to enable comparison of the primary and secondary objectives [26]. The synthesis will be stratified by (i) AI model class (traditional machine-learning classifiers, NLP pipelines, deep learning models, and LLM-based systems), reflecting their differing validation requirements and failure modes, and (ii) clinical pathway (emergency/acute versus outpatient/elective), reflecting differences in the clinical meaning of under- and over-triage. Stratification by model class will extend beyond performance to how each class has been evaluated, including validation designs, reference standards, and reported failure modes. Meta-analysis will be treated as a conditional and exploratory output and will be performed only where a minimum of three studies report comparable outcomes using a similar AI model class, clinical setting, reference standard, outcome definition, and threshold structure. Quantitative analyses will be conducted using StataNow 19.5 SE [27]. Where diagnostic accuracy meta-analysis is appropriate, bivariate random-effects or hierarchical summary receiver operating characteristic (HSROC) models will be used, depending on the availability of sensitivity and specificity estimates and the degree of threshold variation across studies.
Where meta-analysis is not feasible, results will be described narratively with reference to the direction and magnitude of effects across studies.
Reporting bias
Reporting bias will be assessed by considering potential publication bias and selective outcome reporting. Where sufficient studies are available (typically ≥10 and sufficiently homogeneous), funnel plots may be generated to explore potential small-study effects. Statistical tests for funnel plot asymmetry (e.g. Egger’s test) will be considered where appropriate. In addition, where study protocols or trial registry entries are available, reported outcomes will be compared against prespecified outcomes to assess selective outcome reporting. However, it is anticipated that there will be too few studies to explore funnel plots, and availability of study protocols may be limited in AI-based studies.
Certainty of evidence
GRADE will not be applied as the review focuses on AI model development and performance rather than intervention effects on patient outcomes. Instead, methodological quality and risk of bias will be assessed as outlined above.
Study status and timeline
This review has been prospectively registered with PROSPERO (CRD420251244654; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251244654). Database searches were completed on 5 May 2026, and title/abstract screening is ongoing at the time of submission. No full-text screening, data extraction, risk-of-bias assessment, or synthesis has yet been undertaken. We anticipate the following timeline: (A) title/abstract and full-text screening will be completed by August 2026; (B) data extraction and risk-of-bias assessment will be completed by November 2026; and (C) data synthesis and final results are expected by February 2027. Any substantive amendments to the protocol after registration will be recorded in the PROSPERO record and reported transparently in the final review.
Protocol amendments
Any amendments to this protocol after the date of PROSPERO registration and prior to the completion of the review will be documented in a dated protocol-amendment log, with each amendment recorded as follows: (1) date of amendment; (2) section of protocol affected; (3) nature of the amendment; (4) rationale for the amendment; and (5) anticipated impact of the amendment on the review findings.
Substantive amendments affecting the review question, eligibility criteria, information sources, search strategy, outcomes, risk-of-bias assessment, or synthesis approach will also be reflected in an updated PROSPERO registration, with the version history maintained transparently in the registry. Minor amendments, such as corrections to typographical errors or clarification of pre-specified procedures that do not alter the review’s methodology, will be recorded in the internal amendment log only.
All deviations from the protocol encountered during the conduct of the review will be reported in the final review manuscript, in a dedicated subsection of the Methods, in accordance with PRISMA 2020 [20].
Discussion
This review will provide the first systematic synthesis of evidence on the use of AI models for the interpretation and triage of clinical referral documentation across acute and specialist care pathways. A central contribution of the review will be to map how this emerging field defines and evaluates AI-assisted referral triage including the reference standards used, the urgency-category structures adopted, and the reporting of calibration and external validation so that subsequent performance comparisons can be interpreted in their proper clinical and methodological context. The results will be of value to health services where referral volumes are high, and triage is performed manually by nursing or medical staff. If AI models are found to demonstrate high sensitivity and specificity for urgent case identification, then policymakers and service planners may consider incorporating AI-assisted triage into referral workflows, particularly in conjunction with a sequential or hybrid human-AI assessment. This has the potential to reduce delays to clinical review, improve triage consistency, and support safer prioritisation of high-acuity referrals.
The planned synthesis will also allow comparison across AI model types and input modalities, including NLP approaches applied to free-text letters, structured eReferral data, and hybrid inputs. If specific model architectures or input types are found to have comparable or superior performance, this may provide a basis for prioritising their development and evaluation in settings where more resource-intensive approaches are currently unavailable or impractical.
A key strength of this review is its broad scope across care pathways and settings, and its application of PROBAST-AI, the methodological standard specifically developed for AI prediction model studies. The inclusion of studies from multiple countries and service configurations will increase the generalisability of findings.
The main anticipated limitation is the absence of a standardised definition of triage across the included literature, with studies likely applying a range of triage scales and institution-specific thresholds. This will require caution in interpreting pooled accuracy estimates and may limit comparability across settings. More fundamentally, given that triage is a context-dependent clinical prioritisation judgement rather than a single stable diagnostic target, superficially similar metrics may carry different meaning across studies, allowing for the proposed meta-analysis to be treated as exploratory rather than a central expected output. A second limitation is the expected predominance of internal validation studies: where external or prospective validation is absent; findings should be interpreted conservatively with respect to generalisability. Both limitations will be addressed explicitly in risk of bias assessments and will inform recommendations for future research.
References
- 1. O’Cathain A, Knowles E, Maheswaran R, Pearson T, Turner J, Hirst E, et al. A system-wide approach to explaining variation in potentially avoidable emergency admissions: national ecological study. BMJ Qual Saf. 2014;23(1):47–55. pmid:23904507
- 2. Soto CM, Kleinman KP, Simon SR. Quality and correlates of medical record documentation in the ambulatory care setting. BMC Health Serv Res. 2002;2(1):22. pmid:12473161
- 3. Demsash AW, Kassie SY, Dubale AT, Chereka AA, Ngusie HS, Hunde MK, et al. Health professionals’ routine practice documentation and its associated factors in a resource-limited setting: a cross-sectional study. BMJ Health Care Inform. 2023;30(1):e100699. pmid:36796855
- 4. Fogarty C, Cronin P. Waiting for healthcare: a concept analysis. J Adv Nurs. 2008;61(4):463–71. pmid:18234043
- 5. Murray M, Berwick DM. Advanced access: reducing waiting and delays in primary care. JAMA. 2003 Feb 26;289(8):1035.
- 6. Naiker U, FitzGerald G, Dulhunty JM, Rosemann M. Time to wait: a systematic review of strategies that affect out-patient waiting times. Aust Health Rev. 2018;42(3):286–93. pmid:28355525
- 7. Pierce A, Teeling SP, McNamara M, O’Daly B, Daly A. Using lean six sigma in a private hospital setting to reduce trauma orthopedic patient waiting times and associated administrative and consultant caseload. Healthcare (Basel). 2023;11(19):2626. pmid:37830663
- 8. Rathnayake D, Clarke M. The effectiveness of different patient referral systems to shorten waiting times for elective surgeries: systematic review. BMC Health Serv Res. 2021;21(1):155. pmid:33596882
- 9. McVeigh TP, Donnelly D, Al Shehhi M, Jones EA, Murray A, Wedderburn S, et al. Towards establishing consistency in triage in a tertiary specialty. Eur J Hum Genet. 2019;27(4):547–55. pmid:30622329
- 10. Khou V, Ly A, Moore L, Markoulli M, Kalloniatis M, Yapp M, et al. Review of referrals reveal the impact of referral content on the triage and management of ophthalmology wait lists. BMJ Open. 2021;11(9):e047246. pmid:34493511
- 11. Kreimeyer K, Foster M, Pandey A, Arya N, Halford G, Jones SF, et al. Natural language processing systems for capturing and standardizing unstructured clinical information: a systematic review. J Biomed Inform. 2017;73:14–29. pmid:28729030
- 12. Levin S, Toerper M, Hamrock E, Hinson JS, Barnes S, Gardner H, et al. Machine-learning-based electronic triage more accurately differentiates patients with respect to clinical outcomes compared with the emergency severity index. Ann Emerg Med. 2018;71(5):565–74.
- 13. Ahmed A, Nureldaim A, Dawoud E, Mustafa M, Ali A, Abdelgadir A, et al. Clinical impact of artificial intelligence-based triage systems in emergency departments: a systematic review. Cureus. 2025.
- 14. Rego A, Arango-Ibanez JP, Taylor RA, Smith ME, Jones DD, Pelletier J, et al. Artificial intelligence in emergency medicine: a narrative review. Am J Emerg Med. 2026;102:155–65. pmid:41616395
- 15. Almagharbeh WT, Alharrasi M, Khan Rony MK, Kabir S, Alrazeeni DM, Akter F. The applicability of artificial intelligence in managing emergency patients: an umbrella review. Int Emerg Nurs. 2025;83:101710. pmid:41242110
- 16. Kaboudi N, Firouzbakht S, Eftekhar MS, Fayazbakhsh F, Joharivarnoosfaderani N, Ghaderi S, et al. Diagnostic performance of ChatGPT to perform emergency department triage: a systematic review and meta-analysis. Emergency Medicine. 2024.
- 17. Moher D, Shamseer L, Clarke M, Ghersi D, Liberati A, Petticrew M, et al. Preferred reporting items for systematic review and meta-analysis protocols (PRISMA-P) 2015 statement. Syst Rev. 2015;4(1):1. pmid:25554246
- 18. Shamseer L, Moher D, Clarke M, Ghersi D, Liberati A, Petticrew M, et al. Preferred reporting items for systematic review and meta-analysis protocols (PRISMA-P) 2015: elaboration and explanation. BMJ. 2015;:g7647–g7647.
- 19. McInnes MDF, Moher D, Thombs BD, McGrath TA, Bossuyt PM, PRISMA-DTA Group, et al. Preferred reporting items for a systematic review and meta-analysis of diagnostic test accuracy studies: the PRISMA-DTA statement. JAMA. 2018 Jan 23;319(4):388.
- 20. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. pmid:33782057
- 21. Collins GS, Dhiman P, Andaur Navarro CL, Ma J, Hooft L, Reitsma JB, et al. Protocol for development of a reporting guideline (TRIPOD-AI) and risk of bias tool (PROBAST-AI) for diagnostic and prognostic prediction model studies based on artificial intelligence. BMJ Open. 2021;11(7):e048008. pmid:34244270
- 22. Wolff RF, Moons KGM, Riley RD, Whiting PF, Westwood M, Collins GS, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51–8. pmid:30596875
- 23. Sterne JAC, Savović J, Page MJ, Elbers RG, Blencowe NS, Boutron I, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;l4898.
- 24. Sterne JA, Hernán MA, Reeves BC, Savović J, Berkman ND, Viswanathan M, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ. 2016;355:i4919. pmid:27733354
- 25. Whiting PF, Rutjes AWS, Westwood ME, Mallett S, Deeks JJ, Reitsma JB, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155(8):529–36. pmid:22007046
- 26. Campbell M, McKenzie JE, Sowden A, Katikireddi SV, Brennan SE, Ellis S, et al. Synthesis without meta-analysis (SWiM) in systematic reviews: reporting guideline. BMJ. 2020;368:l6890. pmid:31948937
- 27.
Stata Statistical Software: Release 19.5. College Station, TX: StataCorp LLC; 2025.