Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

Artificial intelligence in spine care: A scoping review of diagnostic applications

  • Victoria A. Bensel ,

    Roles Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Writing – original draft

    vabensel@gmail.com

    Affiliations Department of Biomedical Informatics and Data Science, Yale School of Medicine, New Haven, Connecticut, United States of America, VA Connecticut Healthcare System, West Haven, Connecticut, United States of America

  • Anne Habeck,

    Roles Conceptualization, Formal analysis, Writing – review & editing

    Affiliation Bristol, Connecticut, United States of America

  • Marcda Hilaire Brunot,

    Roles Conceptualization, Writing – review & editing

    Affiliation Private Practice, Jacksonville, Florida, United States of America

  • Eleni-James Becton,

    Roles Writing – review & editing

    Affiliation Department of Clinical Research, The University of Jamestown, Jamestown, North Dakota, United States of America

  • Monika Ray,

    Roles Investigation, Project administration, Supervision, Writing – review & editing

    Affiliations Department of Internal Medicine, School of Medicine, University of California Davis, Sacramento, California, United States of America, Center for Healthcare Policy and Research, University of California Davis, Sacramento, California, United States of America

  • Alexandria L. Brackett,

    Roles Data curation, Methodology, Resources, Software, Writing – review & editing

    Affiliation Harvey Cushing/John Hay Whitney Medical Library at Yale University, New Haven, Connecticut, United States of America

  • Anthony J. Lisi

    Roles Conceptualization, Methodology, Writing – review & editing

    Affiliations Department of Biomedical Informatics and Data Science, Yale School of Medicine, New Haven, Connecticut, United States of America, VA Connecticut Healthcare System, West Haven, Connecticut, United States of America

Abstract

Background

Artificial intelligence (AI) is increasingly used to enhance diagnostic accuracy, automate image interpretation, and support clinical decision-making. In the field of spine care, applications include MRI and CT-based detection of lumbar disc degeneration, spinal stenosis, vertebral fractures, and axial spondyloarthritis, as well as emerging symptom-based and multimodal diagnostic tools. However, evidence remains dispersed across modalities and conditions, and the quality and clinical readiness of AI systems vary. This scoping review maps current AI applications for diagnosing spinal disorders and identifies gaps for future research and clinical translation.

Methods

This review followed Joanna Briggs Institute (JBI) and PRISMA-ScR guidelines. Ovid MEDLINE, AMED, Embase, Cochrane CENTRAL, Web of Science, and Scopus were searched from January 2019 to December 2024. Eligible studies were mapped according to AI methodology, diagnostic target, data source, and validation approach, and were required to involve human participants, include sufficient methodological detail, and published in English peer-reviewed journals. No geographic restrictions were applied. Data was extracted on study design, AI methodology, diagnostic target, validation approach, and usability. Methodological quality was assessed using a 19-point scoring system covering study design, reporting clarity, data validation, and feature selection.

Results

Forty-six studies met the inclusion criteria, conducted primarily in Asia and Europe, with two studies from North America and one from South America. Most investigations were retrospective, imaging-based deep learning models applied to MRI or CT for detecting disc herniation, lumbar spinal stenosis, modic changes, vertebral fractures, and sacroiliitis. Several studies used prospective designs or external validation. Diagnostic performance was generally high across imaging models, with many studies describing accuracy that approached or matched clinician benchmarks, particularly in sacroiliitis classification, disc disease detection, and stenosis grading. Methodological scores ranged from 7.5 to 17.5 out of 19, with recurrent weaknesses in handling missing data, feature selection, and data element validation.

Conclusion

This review maps a growing body of literature on AI applications for diagnosing spinal disorders, with studies most frequently reporting favorable performance for MRI- and CT-based detection of degenerative and inflammatory conditions. Evidence remains preliminary and heterogeneous.

Introduction

Spinal disorders are among the leading causes of disability worldwide, contributing substantially to pain, reduced quality of life, and healthcare expenditures [13]. Conditions such as low back pain, spinal stenosis, spondylolisthesis, ankylosing spondylitis, and vertebral fractures represent a diverse spectrum of pathologies that often present with overlapping symptoms [4]. Accurate diagnosis is critical, as treatment pathways vary widely depending on the underlying etiology, disease severity, and patient comorbidities [5]. However, diagnostic evaluation in spine care remains complex and often inconsistent [6].

Current diagnostic approaches rely heavily on clinical examination and imaging modalities such as plain radiographs, magnetic resonance imaging (MRI), and computed tomography (CT) [7,8]. While these tools are indispensable, they are also limited by subjective interpretation, inter-observer variability, and frequent discrepancies between radiographic findings and clinical symptoms [911]. For example, asymptomatic degenerative changes are common on MRI, which complicates the differentiation between incidental findings and clinically meaningful pathology [12,13]. This variation in diagnostic accuracy among practitioners can ultimately lead to misdiagnosis, delayed treatment, and unnecessary interventions.

Additionally, the growing demand for spine-related imaging has further highlighted inefficiencies in traditional diagnostic pathways. Radiological services face increasing workload pressures, and primary care clinicians may lack the specialized expertise required to interpret complex spinal findings [1416]. Consequently, there is a pressing need for early and precise diagnosis to guide appropriate management and avoid overtreatment, particularly given the global rise in musculoskeletal disability [17,18].

Artificial intelligence (AI) offers potential solutions to these challenges by enhancing diagnostic precision, standardizing interpretation, and integrating multimodal data sources [1921]. Machine learning algorithms and deep learning models can be applied to spinal imaging to automatically classify pathologies, quantify structural changes, and detect subtle abnormalities [22,23]. Natural language processing tools are increasingly being used to extract diagnostic insights from electronic health records and radiology reports [24,25]. Predictive models trained on large datasets also show promise for risk stratification and early disease detection [26,27]. Despite this potential, the current evidence base is fragmented, with studies varying widely in methodology, diagnostic targets, and reporting standards [2830].

Given these challenges and opportunities, this scoping review aims to provide an overview of AI-based diagnostic applications in spine care, summarize key findings, and highlight areas where further validation and clinical integration are required.

Methods

This scoping review was conducted in accordance with the Joanna Briggs Institute (JBI) methodology for scoping reviews and follows the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) guidelines.

Search strategy

A comprehensive search strategy was developed in collaboration with a medical librarian (AB) to identify relevant studies on the application of AI in spine diagnostics. The search covered literature published between January 1, 2019, and December 31, 2024. This time window was selected to capture the period of rapid expansion in clinical AI research following the widespread adoption of deep learning methods in medical imaging, which accelerated substantially from 2019 onward [31]. Electronic databases searched included Ovid MEDLINE, AMED, Embase, Cochrane CENTRAL, Web of Science, and Scopus. Search terms combined AI-related concepts (e.g., “machine learning,” “deep learning,” “neural networks,” “natural language processing”) with spine care and diagnostic terms (e.g., “diagnosis,” “classification,” “detection,” “prediction,” “imaging”). Controlled vocabulary was also used when applicable. The only filter applied was to limit the publication years. All identified citations were deduplicated using the Yale University Harvey Cushing/John Hay Whitney Medical Library Reference Deduplicator tool prior to importation and then screening in Covidence. See S1 File for the full search string.

Source evidence selection

After de-duplication, three independent reviewers (VB, MB, AH) screened titles and abstracts for eligibility. Full-text articles were retrieved for studies that met the inclusion criteria or when eligibility was uncertain. Discrepancies were resolved by discussion among primary reviewers (VB, MB, AH). When consensus could not be reached, a third reviewer served as a tiebreaker to make the final inclusion decision. Third reviewer adjudication was needed occasionally throughout the screening process. Cohen’s kappa was not calculated as reviewers were not assigned in fixed pairs; however, an overall agreement rate of >85% was achieved across screening and extraction stages. The study selection process was documented in a PRISMA flow diagram.

Inclusion/exclusion

Studies were eligible if they examined AI applications for diagnostic purposes in spine care (Table 1). Eligible studies involved human participants and compared AI-based diagnostic approaches either to non-AI methods (e.g., clinician interpretation) or to accepted diagnostic benchmarks (e.g., radiology standards, pathology-confirmed findings). Both imaging-based and non-imaging diagnostic applications were included. The unifying criterion across all included studies was the application of an AI model to a diagnostic task in spine care, regardless of input data modality. Exclusion criteria were studies focused exclusively on administrative, financial, or non-clinical applications of AI; animal or cadaveric studies; and reports lacking a clinical decision-making component.

thumbnail
Table 1. Eligibility criteria and their rationale.

https://doi.org/10.1371/journal.pone.0352200.t001

Data extraction

A sample data extraction was performed independently by two reviewers (AH, VB) using a structured extraction template, achieving an > 85% level of agreement, with remaining studies divided between the two extractors. Extracted data included study characteristics (author, year, country, design, and population), AI model type, input data source (e.g., MRI, CT, X-ray, clinical notes), comparator (e.g., radiologist or gold standard), diagnostic task (e.g., detection, classification, grading), performance metrics (e.g., sensitivity, specificity, area under the curve (AUC)), validation approach, and key findings related to diagnostic accuracy, clinical utility, and implementation considerations.

Data analysis

Extracted data were synthesized narratively and summarized in tabular form. Studies were categorized based on diagnostic tasks, AI technology type, and clinical application area. Comparisons between AI and conventional diagnostic methods were emphasized, focusing on diagnostic performance, interpretability, workflow integration, and research gaps were identified. The included studies demonstrated substantial heterogeneity in model architecture, input data sources, outcome definitions, and validation approaches. Quantitative synthesis was therefore not appropriate.

Quality assessment

While formal quality appraisal is not required in scoping reviews, we conducted a structured quality assessment to contextualize the methodological rigor of included studies. We developed a custom checklist adapted from the APPRAISE-AI tool [32], and the TRIPOD-AI extension [33]. This tool evaluates key elements of AI-based clinical research, including data representativeness, transparency, bias mitigation, model performance, and reporting practices (Table 2).

thumbnail
Table 2. Structured quality appraisal domains and criteria adapted from APPRAISE-AI and TRIPOD-AI.

https://doi.org/10.1371/journal.pone.0352200.t002

Each item was scored on a 0–1 scale (0 = not met, 0.5 = partially met, 1 = fully met), with results summarized to guide interpretation of study quality. A score of 0.5 was assigned when a criterion was addressed but incompletely, such as when a study acknowledged missing data without describing how it was handled. Borderline cases were discussed between reviewers, and the 0.5 designation was applied consistently across studies. To ensure consistency in scoring, two independent reviewers assessed all studies, and a third reviewer was available when discrepancies occurred.

Results

A total of 1,485 manuscripts were identified through searches across six databases. After removing 31 duplicates, 1,454 studies were screened by title and abstract. Ultimately, 46 studies met all criteria and were included in the final review (Fig 1).

thumbnail
Fig 1. PRISMA flow diagram of study selection for AI diagnosis applications in spine care.

https://doi.org/10.1371/journal.pone.0352200.g001

Forty-six studies met inclusion criteria (Table 3), published between 2019 and 2025 and conducted predominantly in Asia (n = 24), followed by Europe (n = 14), North America (n = 3), South America (n = 1), and the Middle East (n = 2), with two multinational cohorts. The majority of studies were retrospective and imaging-based diagnostic evaluations using deep learning architectures applied to lumbar MRI or CT for conditions such as disc herniation, spinal stenosis, modic changes, vertebral fractures, and axial spondyloarthritis [3438]. Several prospective studies were also identified [3942], along with cross-sectional or diagnostic accuracy designs [43,44]. Across studies, artificial intelligence was used primarily for detection, segmentation, classification, and grading of spinal pathology on imaging.

thumbnail
Table 3. Overview of articles included in the scoping review (N = 46).

https://doi.org/10.1371/journal.pone.0352200.t003

AI modalities and diagnostic purposes

Deep learning models constituted the majority of approaches (Table 3), with convolutional neural networks (CNNs) frequently applied to automated detection and classification tasks [34,4555]. Several studies integrated segmentation networks such as U-Net [41,45,5658] or in combination with specialized architectures like Mask-RCNN and EfficientDet [57,59]. Traditional machine-learning approaches, including random forest and support vector machines, were also used for feature-based prediction [6062], and one study evaluated a generative large language model (LLM) (ChatGPT-3.5) for diagnostic recommendations [63].

Overall, the included studies fell into four broad methodological clusters. The largest comprised deep learning models applied to structural imaging for detection, classification, or grading of pathology. A second cluster used traditional machine learning approaches, including random forest and support vector machines, applied to either imaging-derived features or structured clinical inputs. A third, smaller cluster examined multimodal models that integrated imaging findings with clinical or laboratory variables to improve diagnostic classification, particularly for axial spondyloarthritis. A fourth cluster consisted of a single study evaluating a generative large language model for clinical consultation support. These clusters differed substantially in their input data, model architecture, outcome definitions, and validation strategies, which precluded direct cross-study comparison and informed the decision to conduct a scoping rather than quantitative synthesis.

Models addressed varied diagnostic applications, including sacroiliitis detection [37,39,45,48,49,57,58], lumbar disc disease [35,5355,6468], spinal stenosis [36,40,50,52], modic changes [34,46,66,69], and vertebral compression fractures [46,53]. Two studies focused on differentiating acute versus chronic fractures [46,53], while others addressed symptom-based diagnostic classification [70] or combined clinical-imaging models for axial spondyloarthritis [37,38,62].

Application trends

Imaging-based AI models were frequently reported to show high performance in identifying structural pathology across lumbar and sacroiliac joint disorders. For example, CNN-based systems were described as accurately identifying sacroiliac joint inflammation and structural lesions consistent with axial spondyloarthritis [37,43,45,49,57,71] and distinguishing calcified disc herniations and degenerative lumbar changes [46,66,68,72,73]. Studies evaluating lumbar spinal stenosis and disc herniation reported successful classification and grading [36,40,50,52,54,55,64]. Prognostic imaging models predicted progression of modic changes and other degenerative features [35].

Models based on structured clinical input were reported to perform less consistently [44,70]. though studies examining multimodal approaches combining clinical and imaging features described improvements in diagnostic classification in axial spondyloarthritis [37,38,62]. One LLM-based diagnostic assistant demonstrated potential for clinical consultation support but was not benchmarked against gold-standard criteria [63].

Diagnostic performance

Across imaging-based systems, studies reported generally high accuracy, with AUCs frequently ≥0.90 in internal testing. Examples include sacroiliitis and BME detection/classification [37,38,45,58,60,74], vertebral compression fracture classification [46,53], multi-feature lumbar grading [54,55,64,72], and lumbar stenosis severity [36]. Several studies reported strong accuracy for disc-related pathology and modic changes [34,35,4649,51,66,67,69,71,73]. Reconstruction/acceleration work showed preserved lesion detectability with shorter scans [78,79] and improved image quality [77]. One clinical decision support system reported AUC 0.84 for triage of disc herniation [45], while a symptom-driven model matched clinician diagnosis 72% of the time (Table 4) [70].

thumbnail
Table 4. Summary table of AI-based diagnostic methods for spinal conditions: outcomes, validation, and usability (N = 46).

https://doi.org/10.1371/journal.pone.0352200.t004

thumbnail
Fig 2. Approach to rating quality and reporting of A-based diagnosis approaches across included studies (N = 46).

https://doi.org/10.1371/journal.pone.0352200.g002

Inter- and intra-rater reliability

Reliability was variably reported. Where provided, agreement between AI outputs and reference readers or between human readers improved or remained high [36,39,42,5356,74,77,78]. Several studies did not quantify AI-specific reliability despite using expert labels [35,37,38,40,41,43,49,51,52,68].

Errors and misclassification

Common error patterns included false positives from artifacts or degenerative changes and false negatives for subtle or small lesions [59,74]. Grading confusion at class boundaries was recurrent for lumbar disc categories and early sacroiliitis grades [37,54,58], with misclassification more frequent for intermediate grades or external datasets [53,54,64]. Dataset size and class imbalance affected performance [50,54,64], and moderate AUCs were observed for radiomics models in active sacroiliitis [71]. A prospective LLM evaluation showed substantial diagnostic and management errors across several spinal conditions [63].

Validation approach

The majority of included studies used retrospective designs in which algorithms were developed and tested on previously collected imaging datasets. A smaller subset employed prospective designs [3943,57,63,70,75,79] or included external validation cohorts [45,46,5355,58,59,64]. External validation was performed in 53% of included studies, often with performance attenuation relative to internal splits; across sacroiliitis/BME, LDH/LCCS/LNRC, fracture, and multi-feature grading [45,46,5355,58,59,64]. Cross-validation frameworks included five-fold and ten-fold schemes [36,57,69], and one study reported internal plus prospective validation [46]. Many reports remained single center with internal splits only [35,37,38,41,42,48,49,72].

Impact on clinicians and workflow integration

Several papers suggested potential to match or exceed clinician performance or to reduce workload and variability, though most did not quantify downstream clinical impact [36,38,45,46,54,55,62,64]. Time savings or efficiency gains were reported with AI-assisted interpretation or accelerated acquisition [52,78,79]. Nevertheless, almost all systems remained research-only without routine clinical deployment [35,37,39,4144,4951,53,5660,6466,68,70,72,74,77].

Usability, interpretability, and implementation considerations

Most of the included studies are proof-of-concept or research-only, with no studies reporting routine clinical deployment. Nearly all systems were evaluated within research pipelines, and authors consistently noted further validation would be required prior to clinical integration.

Interpretability methods included feature importance (Gini) and probability outputs [35,60] saliency/heat-map techniques (Grad-CAM or integrated gradients) [45,55,57,74], and transparent geometric metrics for stenosis [36]. Several models remained “black-box” CNNs with limited explainability [46,50,64]. Some facilitators that were noted included improved image quality and reliability [77,78], potential picture archiving sand communication systems (PACS) integration or clinical net benefit [36,38], open-source release [57], and regulatory clearance for the underlying reconstruction tool [79]. Barriers included single-center training, lack of external validation, protocol heterogeneity, class imbalance, limited demographics, and the need for standardized acquisition and labeling [35,38,57,59,64,67,74].

Quality assessment ratings

Across the 46 studies evaluated, reporting quality was generally strong for criteria related to privacy and consent compliance, financial and ethical disclosures, target variable definition, model justification, and baseline model comparisons, which were consistently well documented. Most studies also provided clear analytical package descriptions and demonstrated reasonable transparency and data documentation. However, notable deficiencies were observed in the handling of missing data and class imbalances, data element validation, model specification, and feature selection methods, each of which were addressed fully in fewer than half of the studies. Moderate reporting quality was evident for domains such as data availability, bias identification and mitigation, and data splitting procedures. These findings are summarized in Fig 2, which depicts the number of studies receiving full or partial credit for each reporting criterion.

Discussion

This scoping review provides an overview of recent applications of artificial intelligence for the diagnosis of spinal disorders. Across 46 included studies, the majority reported on imaging-based models, most often applied to MRI, CT, or radiographs. A smaller number incorporated clinical data or combined clinical and imaging inputs. The emphasis on imaging aligns with broader trends in medical AI research, where computer vision techniques dominate because of the relative availability of structured image datasets [31,80].

Many studies reported high diagnostic performance, with authors describing accuracies that approached or matched clinician benchmarks, across various applications including, but not limited to, lumbar disc herniation, spinal stenosis, vertebral fractures, scoliosis, and inflammatory disorders (i.e., axial spondyloarthritis). However, despite encouraging performance metrics, important limitations temper these findings. Most studies were retrospective, single-center analyses with small-to-moderate sample sizes. Few studies included external validation, and prospective testing in clinical workflows was rare. The lack of external testing reduces confidence in generalizability and is consistent with limitations reported in other AI reviews in spine care [30,81].

The diagnostic targets were heterogeneous, reflecting both the versatility of AI methods and the fragmented state of the literature. Some studies focused on binary classification of pathology, while others attempted severity grading or prognostic prediction. This variation limited direct comparison across studies. A small subset examined the integration of multimodal data, combining imaging with clinical or demographic variables, suggesting that broader data inputs may enhance diagnostic utility. At the same time, several reports lacked transparency in data sources, preprocessing, and feature selection, making reproducibility difficult to assess. These gaps align with concerns raised in prior reviews, which emphasized the need for standardized outcome measures and consistent reporting frameworks [82,83].

An emerging pattern in the literature was that AI was most frequently studied in roles augmenting rather than replacing human interpretation. Several studies reported reduced interobserver variability or improved efficiency when AI tools were applied as decision-support systems alongside clinician review. However, the path from research prototype to clinical deployment involves challenges that were largely unaddressed in the included literature, including regulatory approval requirements, integration with existing clinical infrastructure such as PACS systems, interpretability standards for clinician-facing tools, and the need for prospective evidence in diverse patient populations [84]. Addressing these implementation challenges will be necessary before the diagnostic potential suggested by this literature can be realized in routine care.

Overall, diagnostic AI in spine care shows potential to support earlier detection, improve efficiency, and standardize interpretation across diverse conditions. However, evidence remains preliminary. Several specific evidence gaps identified in this review warrant further investigation. First, the predominance of imaging-based retrospective studies from Asia and Europe means that prospective, multi-center studies in diverse geographic and clinical settings are needed to establish generalizability. Second, the near-absence of non-imaging and multimodal diagnostic tools in the literature represents an underdeveloped area where primary research is warranted, particularly given the potential of combined clinical-imaging approaches demonstrated in axial spondyloarthritis studies. Third, no included studies evaluated real-world clinical deployment or measured patient outcomes associated with AI-assisted diagnosis, representing a critical gap that future prospective studies should address. Fourth, the heterogeneity in outcome definitions, validation approaches, and reporting standards across the current literature suggest that a methodological consensus or reporting framework for AI diagnostic studies in spine care would be a valuable contribution.

Our results provide a snapshot of a rapidly evolving domain. The advancement of AI in healthcare, paralleling its expansion in other sectors, has precipitated substantial concerns regarding data privacy, ethical considerations, and regulatory obstacles. Emerging contentious issues encompass data ownership, model stewardship, data and model bias, as well as model transparency and interpretability [8589]. These concerns are particularly pertinent to patient privacy, data drift, the potential for model manipulation, biases in models that may unjustly affect marginalized populations, and disparities in access to high-quality care. To address these challenges, guidelines aimed at mitigating risks associated with AI systems have been promulgated by the US Federal Trade Commission [90], the European Union [91], China [92], and various industry and professional stakeholders [93]. As AI becomes increasingly integrated into healthcare delivery and operational processes, discussions regarding the ethical and effective utilization of AI are expected to intensify.

Limitations

The search was restricted to studies published in English between January 1st, 2019, and December 31st, 2024, which may have led to the exclusion of relevant work published in other languages or outside this time frame. No formal risk of bias assessment was performed, consistent with the scoping review methodology, although a structured quality appraisal was applied to provide context. Several common sources of bias were identified narratively across the included literature. The predominance of retrospective, single-center designs introduces selection bias and limits the representativeness of training datasets seen in included studies. Small sample sizes were frequently noted, particularly in studies evaluating non-imaging or multimodal models, which increases the risk of overfitting and reduces confidence in reported performance metrics. The reliance on internal validation splits in a substantial number of the studies included in this review.

Underrepresentation of certain populations, settings, and geographic regions is also a notable limitation of the current evidence base. A large portion of included studies were conducted in Asia and Europe, with limited representation from North America, South America, the Middle East, and Africa. No studies were identified from low- or middle-income country settings, which raises questions about the generalizability of findings to healthcare systems with different imaging infrastructure, patient demographics, and resource availability. Within studies, demographic reporting was frequently incomplete, with many studies not reporting patient age distributions, sex, race, or comorbidity profiles. This limits the ability to assess whether AI models perform equitably across patient subgroups, which is a recognized concern in medical AI research.

There was substantial heterogeneity across included studies in terms of AI model design, data inputs, diagnostic targets, and often without standardized thresholds or reporting conventions. This degree of methodological variation meant that performance metrics were not directly comparable across studies, and pooled quantitative analysis was not appropriate.

Most included studies were retrospective and often single center, which may affect generalizability. Details regarding data sources, preprocessing, and model development were not consistently reported, making reproducibility difficult to assess. In addition, many of the identified AI applications were evaluated in narrowly defined patient groups. The absence of multi-center or prospective validation methods within the included studies indicates that clinical applicability is uncertain.

Conclusions

We present a comprehensive overview of the literature on the application of AI in the diagnosis of spinal conditions. Studies frequently reported AI system performance comparable to or supportive of human interpretation. However, methodological variability, reporting transparency limitations, and reliance on retrospective single-center datasets were common. Based on the evidence mapped in this review, future studies should prioritize prospective and multi-center designs, adopt standardized reporting frameworks such as TRIPOD-AI [33], utilize diverse patient populations with complete demographic reporting, and evaluate AI tools within clinical workflows rather than isolated research pipelines. Lastly, investment in the infrastructure needed to support prospective AI validation, including standardized imaging protocols, data sharing frameworks, and regulatory pathways for AI-based diagnostic tools, will be critical to translating research findings into safe and equitable clinical practice.

References

  1. 1. Wu A, March L, Zheng X, Huang J, Wang X, Zhao J, et al. Global low back pain prevalence and years lived with disability from 1990 to 2017: estimates from the Global Burden of Disease Study 2017. Ann Transl Med. 2020;8(6):299. pmid:32355743
  2. 2. Zhang C, Zi S, Chen Q, Zhang S. The burden, trends, and projections of low back pain attributable to high body mass index globally: an analysis of the global burden of disease study from 1990 to 2021 and projections to 2050. Front Med (Lausanne). 2024;11:1469298. pmid:39507709
  3. 3. Chang D, Lui A, Matsoyan A, Safaee MM, Aryan H, Ames C. Comparative review of the socioeconomic burden of lower back pain in the United States and globally. Neurospine. 2024;21(2):487–501. pmid:38955526
  4. 4. Huntoon E, Huntoon M. Differential diagnosis of low back pain. Semin Pain Med. 2004;2(3):138–44.
  5. 5. Alharbi TAF, Rababa M, Alsuwayl H, Alsubail A, Alenizi WS. Diagnostic challenges and patient safety: the critical role of accuracy - a systematic review. J Multidiscip Healthc. 2025;18:3051–64. pmid:40470160
  6. 6. Mathieu J, Pasquier M, Descarreaux M, Marchand A-A. Diagnosis value of patient evaluation components applicable in primary care settings for the diagnosis of low back pain: a scoping review of systematic Reviews. J Clin Med. 2023;12(10):3581. pmid:37240687
  7. 7. Ruiz Santiago F, Láinez Ramos-Bossini AJ, Wáng YXJ, Martínez Barbero JP, García Espinosa J, Martínez Martínez A. The value of magnetic resonance imaging and computed tomography in the study of spinal disorders. Quant Imaging Med Surg. 2022;12(7):3947–86. pmid:35782254
  8. 8. Teichner EM, Subtirelu RC, Crutchfield CR, Parikh C, Ashok A, Talasila S, et al. The advancement and utility of multimodal imaging in the diagnosis of degenerative disc disease. Front Radiol. 2025;5:1298054. pmid:40115420
  9. 9. Itri JN, Tappouni RR, McEachern RO, Pesch AJ, Patel SH. Fundamentals of diagnostic error in imaging. Radiographics. 2018;38(6):1845–65. pmid:30303801
  10. 10. Benchoufi M, Matzner-Lober E, Molinari N, Jannot A-S, Soyer P. Interobserver agreement issues in radiology. Diagn Interv Imaging. 2020;101(10):639–41. pmid:32958434
  11. 11. Bhatnagar G, Mallett S, Quinn L, Beable R, Bungay H, Betts M, et al. Interobserver variation in the interpretation of magnetic resonance enterography in Crohn’s disease. Br J Radiol. 2022;95(1134):20210995. pmid:35195444
  12. 12. Brinjikji W, Luetmer PH, Comstock B, Bresnahan BW, Chen LE, Deyo RA, et al. Systematic literature review of imaging features of spinal degeneration in asymptomatic populations. AJNR Am J Neuroradiol. 2015;36(4):811–6. pmid:25430861
  13. 13. Ract I, Meadeb J-M, Mercy G, Cueff F, Husson J-L, Guillin R. A review of the value of MRI signs in low back pain. Diagn Interv Imaging. 2015;96(3):239–49. pmid:24674892
  14. 14. Kasalak Ö, Alnahwi H, Toxopeus R, Pennings JP, Yakar D, Kwee TC. Work overload and diagnostic errors in radiology. Eur J Radiol. 2023;167:111032. pmid:37579563
  15. 15. Siewert B, Bruno MA, Bourland JD, Slanetz PJ, Guillerman P, Schwartz ES, et al. Seven challenges in radiology practice: from declining reimbursement to inadequate labor force: summary of the 2023 ACR intersociety meeting. J Am Coll Radiol. 2025;22(1):129–38. pmid:39480363
  16. 16. Parmar V, Thompson L, Aniq H. Comparison of referrals for lumbar spine magnetic resonance imaging from physiotherapists, primary care and secondary care: how should referral pathways be optimised? Physiotherapy. 2015;101(1):82–7. pmid:25125386
  17. 17. Zhou J, Xie S, Xu S, Zhang Y, Li Y, Sun Q, et al. From pain to progress: comprehensive analysis of musculoskeletal disorders worldwide. J Pain Res. 2024;17:3455–72. pmid:39469334
  18. 18. Liu M, Rong J, An X, Li Y, Min Y, Yuan G, et al. Global, regional, and national burden of musculoskeletal disorders, 1990–2021: an analysis of the global burden of disease study 2021 and forecast to 2035. Front Public Health. 2025;13:1562701.
  19. 19. Huang W, Shu N. AI-powered integration of multimodal imaging in precision medicine for neuropsychiatric disorders. Cell Rep Med. 2025;6(5):102132. pmid:40398391
  20. 20. Jandoubi B, Akhloufi MA. Multimodal artificial intelligence in medical diagnostics. Inf. 2025;16(7):591.
  21. 21. Cascella M, Leoni MLG, Shariff MN, Varrassi G. Artificial intelligence-driven diagnostic processes and comprehensive multimodal models in pain medicine. J Personal Med. 2024;14(9):983.
  22. 22. Cui Y, Zhu J, Duan Z, Liao Z, Wang S, Liu W. Artificial intelligence in spinal imaging: current status and future directions. Int J Environ Res Public Health. 2022;19(18):11708. pmid:36141981
  23. 23. Glenn ER, Seidenstein AH, Savage CH, Zhu AR, Khanna R, Middendorf J, et al. Artificial intelligence for lumbar spine anatomy and pathology detection: a scoping review. J Orthopaed Rep. 2026;5(3):100760.
  24. 24. Jerfy A, Selden O, Balkrishnan R. The growing impact of natural language processing in healthcare and public health. Inquiry. 2024;61. pmid:39396164
  25. 25. Hossain E, Rana R, Higgins N, Soar J, Barua PD, Pisani AR, et al. Natural language processing in electronic health records in relation to healthcare decision-making: a systematic review. Comput Biol Med. 2023;155:106649. pmid:36805219
  26. 26. Jaganathan D, Vadivel A, Jansi Rani S, Thangamuthu V. Integrating real-time data with predictive models for early disease detection in metaverse healthcare. In: Mahajan S, Chatterjee JM, editors. Federated Learning in Metaverse Healthcare. Academic Press; 2026. pp. 265–91.
  27. 27. Dixon D, Sattar H, Moros N, Kesireddy SR, Ahsan H, Lakkimsetti M, et al. Unveiling the influence of AI predictive analytics on patient outcomes: a comprehensive narrative review. Cureus. 2024;16(5):e59954. pmid:38854327
  28. 28. Ambati VS, Saggi S, Dada A, Alan N. Has artificial intelligence in spine surgery lived up to the hype? A narrative review of recent approaches, current challenges, and the path forward. Art Int Surg. 2025;5(1):53–64.
  29. 29. Kita K, Kaito T. Artificial intelligence in spine research: a multimodal perspective beyond imaging. Spine Res. 2025;1(1):7–12.
  30. 30. Lee S, Jung J-Y, Mahatthanatrakul A, Kim J-S. Artificial intelligence in spinal imaging and patient care: a review of recent advances. Neurospine. 2024;21(2):474–86. pmid:38955525
  31. 31. Pinto-Coelho L. How artificial intelligence is shaping medical imaging technology: a survey of innovations and applications. Bioengineering (Basel). 2023;10(12):1435.
  32. 32. Kwong JCC, Khondker A, Lajkosz K, McDermott MBA, Frigola XB, McCradden MD, et al. APPRAISE-AI tool for quantitative evaluation of AI studies for clinical decision support. JAMA Network Open. 2023;6(9):e2335377.
  33. 33. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:q902. pmid:38636956
  34. 34. Gao KT, Tibrewala R, Hess M, Bharadwaj UU, Inamdar G, Link TM, et al. Automatic detection and voxel-wise mapping of lumbar spine Modic changes with deep learning. JOR Spine. 2022;5(2):e1204. pmid:35783915
  35. 35. Cheung JPY, Kuang X, Zhang T, Wang K, Yang C. 5-Year progression prediction of endplate defects: utilizing the EDPP-Flow convolutional neural network based on unbalanced data. J Orthop. 2023;38:7–13. pmid:36910507
  36. 36. van der Graaf JW, Brundel L, van Hooff ML, de Kleuver M, Lessmann N, Maresch BJ, et al. AI-based lumbar central canal stenosis classification on sagittal MR images is comparable to experienced radiologists using axial images. Eur Radiol. 2025;35(4):2298–306. pmid:39299953
  37. 37. Zhang K, Liu C, Pan J, Zhu Y, Li X, Zheng J, et al. Use of MRI-based deep learning radiomics to diagnose sacroiliitis related to axial spondyloarthritis. Eur J Radiol. 2024;172:111347. pmid:38325189
  38. 38. Zhang Z, Pan Y, Lu Y, Ye L, Zheng M, Zhang G, et al. The TabNet model for diagnosing axial spondyloarthritis using MRI imaging findings and clinical risk factors. Int J Rheum Dis. 2024;27(12):e70004. pmid:39690496
  39. 39. Jans LBO, Chen M, Elewaut D, Van den Bosch F, Carron P, Jacques P, et al. MRI-based Synthetic CT in the detection of structural lesions in patients with suspected sacroiliitis: comparison with MRI. Radiology. 2021;298(2):343–9. pmid:33350891
  40. 40. Georgiev R, Novakova M, Bliznakova K. Clinical assessment of CoLumbo Deep Learning System for Central Canal Stenosis Diagnostics. Euras J Med Oncol. 2023;7(1):42.
  41. 41. Lin KYY, Peng C, Lee KH, Chan SCW, Chung HY. Deep learning algorithms for magnetic resonance imaging of inflammatory sacroiliitis in axial spondyloarthritis. Rheumatology (Oxford). 2022;61(10):4198–206. pmid:35104321
  42. 42. Lin Y, Cao P, Chan SCW, Lee KH, Lau VWH, Chung HY. Deep learning algorithm of the SPARCC scoring system in SI joint MRI. J Magn Reson Imaging. 2024;60(4):1390–9. pmid:38168061
  43. 43. Krabbe S, Møller JM, Hadsbjerg AEF, Ewald A, Hangaard S, Pedersen SJ, et al. Detection of structural lesions of the sacroiliac joints in patients with spondyloarthritis: a comparison of T1-weighted 3D spoiled gradient echo MRI and MRI-based synthetic CT versus T1-weighted turbo spin echo MRI. Skeletal Radiol. 2024;53(11):2459–68. pmid:38592521
  44. 44. Badahman F, Alsobhi M, Alzahrani A, Chevidikunnan MF, Neamatallah Z, Alqarni A, et al. Validating the accuracy of a patient-facing clinical decision support system in predicting lumbar disc herniation: diagnostic accuracy study. Diagnostics (Basel). 2024;14(17):1870. pmid:39272655
  45. 45. Bressem KK, Adams LC, Proft F, Hermann KGA, Diekhoff T, Spiller L. Deep learning detects changes indicative of axial spondyloarthritis at MRI of sacroiliac joints. Radiology. 2023;307(3):e239007.
  46. 46. Chen J, Liu S, Li Y, Zhang Z, Liao N, Shi H, et al. Deep learning model for automated detection of fresh and old vertebral fractures on thoracolumbar CT. Eur Spine J. 2025;34(3):1177–86. pmid:39708132
  47. 47. Gao F, Liu S, Zhang X, Wang X, Zhang J. Automated grading of lumbar disc degeneration using a push-pull regularization network based on MRI. J Magn Reson Imaging. 2021;53(3):799–806. pmid:33094867
  48. 48. Lee K-H, Lee R-W, Lee K-H, Park W, Kwon S-R, Lim M-J. The development and validation of an AI diagnostic model for sacroiliitis: a deep-learning approach. Diagnostics (Basel). 2023;13(24):3643. pmid:38132228
  49. 49. Lee KH, Choi ST, Lee GY, Ha YJ, Choi S-I. Method for diagnosing the bone marrow edema of sacroiliac joint in patients with axial spondyloarthritis using magnetic resonance image analysis based on deep learning. Diagnostics (Basel). 2021;11(7):1156. pmid:34202607
  50. 50. Shahzadi T, Ali MU, Majeed F, Sana MU, Diaz RM, Samad MA, et al. Nerve root compression analysis to find lumbar spine stenosis on MRI Using CNN. Diagnostics (Basel). 2023;13(18):2975. pmid:37761342
  51. 51. Liawrungrueang W, Cholamjiak W, Sarasombath P, Jitpakdee K, Kotheeranurak V. Artificial intelligence classification for detecting and grading lumbar intervertebral disc degeneration. Spine Surg Relat Res. 2024;8(6):552–9. pmid:39659374
  52. 52. Lim DSW, Makmur A, Zhu L, Zhang W, Cheng AJL, Sia DSY, et al. Improved productivity using deep learning-assisted reporting for lumbar spine MRI. Radiology. 2022;305(1):160–6. pmid:35699577
  53. 53. Ono Y, Suzuki N, Sakano R, Kikuchi Y, Kimura T, Sutherland K, et al. A deep learning-based model for classifying osteoporotic lumbar vertebral fractures on radiographs: a retrospective model development and validation study. J Imaging. 2023;9(9):187. pmid:37754951
  54. 54. Su ZH, Liu J, Yang MS, Chen ZY, You K, Shen J. Automatic grading of disc herniation, central canal stenosis and nerve roots compression in lumbar magnetic resonance image diagnosis. Front Endocrinol. 2022;13:890371.
  55. 55. Yoo H, Yoo R-E, Choi SH, Hwang I, Lee JY, Seo JY, et al. Deep learning-based reconstruction for acceleration of lumbar spine MRI: a prospective comparison with standard MRI. Eur Radiol. 2023;33(12):8656–68. pmid:37498386
  56. 56. Dorfner FJ, Vahldiek JL, Donle L, Zhukov A, Xu L, Häntze H, et al. Anatomy-centred deep learning improves generalisability and progression prediction in radiographic sacroiliitis detection. RMD Open. 2024;10(4):e004628. pmid:39719299
  57. 57. Roels J, De Craemer A, Renson T, Hooge M, Gevaert A, Van Den Berghe T, et al. Machine learning pipeline for predicting bone marrow edema along the sacroiliac joints on magnetic resonance imaging. Arthritis Rheumatol. 2023;75(12):2169–77.
  58. 58. Zhang K, Luo G, Li W, Zhu Y, Pan J, Li X, et al. Automatic image segmentation and grading diagnosis of sacroiliitis associated with AS using a deep convolutional neural network on CT images. J Digit Imaging. 2023;36(5):2025–34. pmid:37268841
  59. 59. Bordner A, Aouad T, Medina CL, Yang S, Molto A, Talbot H, et al. A deep learning model for the diagnosis of sacroiliitis according to Assessment of SpondyloArthritis International Society classification criteria with magnetic resonance imaging. Diagn Interv Imaging. 2023;104(7–8):373–83. pmid:37012131
  60. 60. Abdollah V, Parent EC, Dolatabadi S, Marr E, Croutze R, Wachowicz K, et al. Texture analysis in the classification of T2 -weighted magnetic resonance images in persons with and without low back pain. J Orthop Res. 2021;39(10):2187–96. pmid:33247597
  61. 61. Faleiros MC, Nogueira-Barbosa MH, Dalto VF, Júnior JRF, Tenório APM, Luppino-Assad R, et al. Machine learning techniques for computer-aided classification of active inflammatory sacroiliitis in magnetic resonance imaging. Adv Rheumatol. 2020;60(1):25. pmid:32381053
  62. 62. Redeker I, Tsiami S, Eicker J, Kiltz U, Kiefer D, Andreica I, et al. Identification of a machine learning-based diagnostic model for axial spondyloarthritis in rheumatological routine care using a random forest approach. RMD Open. 2024;10(4):e004702. pmid:39608866
  63. 63. Chalhoub R, Mouawad A, Aoun M, Daher M, El-sett P, Kreichati G. Will ChatGPT be able to replace a spine surgeon in the clinical setting? World Neurosurg. 2024;185:e648–52.
  64. 64. Zhang W, Chen Z, Su Z, Wang Z, Hai J, Huang C, et al. Deep learning-based detection and classification of lumbar disc herniation on magnetic resonance images. JOR Spine. 2023;6(3):e1276. pmid:37780833
  65. 65. Ke B, Ma W, Xuan J, Liang Y, Zhou L, Jiang W, et al. MRI to digital medicine diagnosis: integrating deep learning into clinical decision-making for lumbar degenerative diseases. Front Surg. 2025;11:1424716. pmid:39834502
  66. 66. Liu G, Wang L, You S-N, Wang Z, Zhu S, Chen C, et al. Automatic detection and classification of modic changes in MRI images using deep learning: intelligent assisted diagnosis system. Orthop Surg. 2024;16(1):196–206. pmid:37933461
  67. 67. Liu L, Zhang H, Zhang W, Mei W, Huang R. Sacroiliitis diagnosis based on interpretable features and multi-task learning. Phys Med Biol. 2024;69(4). pmid:38237177
  68. 68. Liu Z, Zhang H, Zhang M, Qu C, Li L, Sun Y, et al. Compare three deep learning-based artificial intelligence models for classification of calcified lumbar disc herniation: a multicenter diagnostic study. Front Surg. 2024;11:1458569. pmid:39569028
  69. 69. Athertya JS, Saravana Kumar G, Govindaraj J. Detection of Modic changes in MR images of spine using local binary patterns. Biocybern Biomed Eng. 2019;39(1):17–29.
  70. 70. Soin A, Hirschbeck M, Verdon M, Manchikanti L. A pilot study implementing a machine learning algorithm to use artificial intelligence to diagnose spinal conditions. Pain Physician. 2022.
  71. 71. Triantafyllou M, Klontzas ME, Koltsakis E, Papakosta V, Spanakis K, Karantanas AH. Radiomics for the detection of active sacroiliitis using MR imaging. Diagnostics (Basel). 2023;13(15):2587. pmid:37568950
  72. 72. Nigru AS, Benini S, Bonetti M, Bragaglio G, Frigerio M, Maffezzoni F, et al. External validation of SpineNetV2 on a comprehensive set of radiological features for grading lumbosacral disc pathologies. N Am Spine Soc J. 2024;20:100564. pmid:39640208
  73. 73. Lehnen NC, Haase R, Faber J, Rüber T, Vatter H, Radbruch A, et al. Detection of degenerative changes on MR images of the lumbar spine with a convolutional neural network: a feasibility study. Diagnostics (Basel). 2021;11(5):902. pmid:34069362
  74. 74. Bressem KK, Vahldiek JL, Adams L, Niehues SM, Haibel H, Rodriguez VR, et al. Deep learning for detection of radiographic sacroiliitis: achieving expert-level performance. Arthritis Res Ther. 2021;23(1):106. pmid:33832519
  75. 75. Hartley T, Hicks Y, Davies JL, Cazzola D, Sheeran L. BACK-to-MOVE: Machine learning and computer vision model automating clinical classification of non-specific low back pain for personalised management. PLoS One. 2024;19(5):e0302899. pmid:38728282
  76. 76. Lagerstrand K, Hebelka H, Brisby H. Identification of potentially painful disc fissures in magnetic resonance images using machine-learning modelling. Eur Spine J. 2022;31(8):1992–9. pmid:34854974
  77. 77. Miyo R, Yasaka K, Hamada A, Sakamoto N, Hosoi R, Mizuki M, et al. Deep-learning reconstruction for the evaluation of lumbar spinal stenosis in computed tomography. Medicine (Baltimore). 2023;102(23):e33910. pmid:37335676
  78. 78. Seo G, Lee SJ, Park DH, Paeng SH, Koerzdoerfer G, Nickel MD, et al. Image quality and lesion detectability of deep learning-accelerated T2-weighted Dixon imaging of the cervical spine. Skeletal Radiol. 2023;52(12):2451–9. pmid:37233758
  79. 79. Tang H, Hong M, Yu L, Song Y, Cao M, Xiang L, et al. Deep learning reconstruction for lumbar spine MRI acceleration: a prospective study. Eur Radiol Exp. 2024;8(1):67. pmid:38902467
  80. 80. Houssein EH, Gamal AM, Younis EMG, Mohamed E. Explainable artificial intelligence for medical imaging systems using deep learning: a comprehensive review. Cluster Comput. 2025;28(7).
  81. 81. Nugraha HK, Rasmussen AP, Mulford KL, Yang L, Wyles CC, Larson AN. AI in pediatric spine care: clinical, research, and ethical considerations. J Clin Med. 2025;14(22):8115.
  82. 82. Valtonen L, Mäkinen SJ, Kirjavainen J. Advancing reproducibility and accountability of unsupervised machine learning in text mining: importance of transparency in reporting preprocessing and algorithm selection. Organ Res Methods. 2022;27(1):88–113.
  83. 83. Sedlakova J, Daniore P, Horn Wintsch A, Wolf M, Stanikic M, Haag C, et al. Challenges and best practices for digital unstructured data enrichment in health research: a systematic narrative review. PLOS Digit Health. 2023;2(10):e0000347. pmid:37819910
  84. 84. Bensel VA, Habeck A, Brunot MH, Becton E-J, Ray M, Brackett AL, et al. Artificial intelligence in spine care: a scoping review of treatment applications. N Am Spine Soc J. 2025;25:100827. pmid:41536317
  85. 85. Regulatory Compliance Associates, Inc. Proposed Regulatory Framework for Modifications to Artificial Intelligence/Machine Learning (AI/ML)-Based Software as a Medical Device (SaMD) [Internet]. U.S. Food and Drug Administration; 2019 Jun. Report FDA-2019-N-1185-0068. Available from: https://www.regulations.gov/document/FDA-2019-N-1185-0068
  86. 86. Health Data, Technology, and Interoperability: Certification Program Updates, Algorithm Transparency, and Information Sharing. Code of Federal Regulations [Internet]. p. Parts 170, 171. 2024. Available from: https://www.federalregister.gov/documents/2024/01/09/2023-28857/health-data-technology-and-interoperability-certification-program-updates-algorithm-transparency-and
  87. 87. polepole. Understanding AI Manipulation: A Case Study on the “Agitation” Method [Forum]. OpenAI Developer Community [Internet]. 2024. Available from: https://community.openai.com/t/understanding-ai-manipulation-a-case-study-on-the-agitation-method/594003
  88. 88. The Light Collective & Digital Public. Collective digital rights for patients in health AI [Version 1.0] [Internet]. 2024. Available from: https://lightcollective.org/wp-content/uploads/2024/03/Collective-Digital-Rights-For-Patients_v1.0.pdf
  89. 89. Coalition for Health AI. Blueprint for Trustworthy AI Implementation Guidance and Assurance for Healthcare. In. 2023. Available from: https://assets.ctfassets.net/7s4afyr9pmov/4AXIWGIlcrjWDaW2ueTaRS/f98e5cb2528187635895cce6ba5ec309/Blueprint_for_Trustworthy_AI.pdf
  90. 90. Jillson E. Aiming for truth, fairness, and equity in your company’s use of AI [Internet]. Federal Trade Commission; 2021. Available from: https://www.ftc.gov/news-events/blogs/business-blog/2021/04/aiming-truth-fairness-equity-your-companys-use-ai
  91. 91. Jeong C, Lee S, Jeong S, Kim S. A study on the framework for evaluating the ethics and trustworthiness of generative AI. 2025.
  92. 92. Huang Y, Arora C, Huong WC, Kanij T, Madugalla A, Grundy J. Ethical concerns of generative AI and mitigation strategies: A systematic mapping study. Appl Soft Comput. 2026;193:114789.
  93. 93. Cheong BC. Transparency and accountability in AI systems: safeguarding wellbeing in the age of algorithmic decision-making. Front Hum Dyn. 2024;6.