Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

Risk prediction models for inadequate bowel preparation before colonoscopy: A systematic review and meta-analysis

  • Hui-ling Yan ,

    Contributed equally to this work with: Hui-ling Yan, Jia-yu Fu

    Roles Conceptualization, Data curation, Formal analysis, Methodology, Project administration, Writing – original draft

    Affiliation School of Nursing, Hunan Engineering Research Center for Early Diagnosis and Treatment of Liver Cancer, Hunan Province Key Laboratory of Tumor Cellular & Molecular Pathology, Cancer Research Institute, Hengyang Medical School, University of South China, Hengyang, Hunan Province, China

  • Jia-yu Fu ,

    Contributed equally to this work with: Hui-ling Yan, Jia-yu Fu

    Roles Conceptualization, Data curation, Formal analysis, Methodology, Supervision, Visualization, Writing – review & editing

    Affiliation Department of Gastroenterology, The Second Affiliated Hospital, University of South China, Hengyang, Hunan Province, China

  • Mao-ting Huang,

    Roles Conceptualization, Data curation, Formal analysis, Visualization, Writing – original draft

    Affiliation School of Nursing, Hunan Engineering Research Center for Early Diagnosis and Treatment of Liver Cancer, Hunan Province Key Laboratory of Tumor Cellular & Molecular Pathology, Cancer Research Institute, Hengyang Medical School, University of South China, Hengyang, Hunan Province, China

  • Wen-ting Yi,

    Roles Conceptualization, Methodology, Writing – original draft

    Affiliation School of Nursing, Hunan Engineering Research Center for Early Diagnosis and Treatment of Liver Cancer, Hunan Province Key Laboratory of Tumor Cellular & Molecular Pathology, Cancer Research Institute, Hengyang Medical School, University of South China, Hengyang, Hunan Province, China

  • Jia-jun Liu,

    Roles Conceptualization, Data curation, Writing – original draft

    Affiliation School of Nursing, Hunan Engineering Research Center for Early Diagnosis and Treatment of Liver Cancer, Hunan Province Key Laboratory of Tumor Cellular & Molecular Pathology, Cancer Research Institute, Hengyang Medical School, University of South China, Hengyang, Hunan Province, China

  • Xiao-yan Huang,

    Roles Conceptualization, Data curation, Writing – original draft

    Affiliation School of Nursing, Hunan Engineering Research Center for Early Diagnosis and Treatment of Liver Cancer, Hunan Province Key Laboratory of Tumor Cellular & Molecular Pathology, Cancer Research Institute, Hengyang Medical School, University of South China, Hengyang, Hunan Province, China

  • Yao Fang,

    Roles Conceptualization, Data curation, Writing – original draft

    Affiliation School of Nursing, Hunan Engineering Research Center for Early Diagnosis and Treatment of Liver Cancer, Hunan Province Key Laboratory of Tumor Cellular & Molecular Pathology, Cancer Research Institute, Hengyang Medical School, University of South China, Hengyang, Hunan Province, China

  • Ling Zhao,

    Roles Conceptualization, Methodology, Writing – review & editing

    Affiliation School of Nursing, Hunan Engineering Research Center for Early Diagnosis and Treatment of Liver Cancer, Hunan Province Key Laboratory of Tumor Cellular & Molecular Pathology, Cancer Research Institute, Hengyang Medical School, University of South China, Hengyang, Hunan Province, China

  • Yun-shan Chen ,

    Roles Conceptualization, Formal analysis, Methodology, Writing – review & editing

    13025185551@163.com (CYS); zengying2003@126.com (ZY)

    Affiliation School of Nursing, Hunan Engineering Research Center for Early Diagnosis and Treatment of Liver Cancer, Hunan Province Key Laboratory of Tumor Cellular & Molecular Pathology, Cancer Research Institute, Hengyang Medical School, University of South China, Hengyang, Hunan Province, China

  • Ying Zeng

    Roles Conceptualization, Data curation, Formal analysis, Methodology, Project administration, Supervision, Writing – review & editing

    13025185551@163.com (CYS); zengying2003@126.com (ZY)

    Affiliation School of Nursing, Hunan Engineering Research Center for Early Diagnosis and Treatment of Liver Cancer, Hunan Province Key Laboratory of Tumor Cellular & Molecular Pathology, Cancer Research Institute, Hengyang Medical School, University of South China, Hengyang, Hunan Province, China

Abstract

Background

Inadequate bowel preparation (IBP) can impair the safety, efficiency, and diagnostic accuracy of colonoscopy. Numerous multivariable prediction models have been developed to identify patients at high risk of IBP, yet their reported performance varies substantially.

Objectives

To systematically evaluate the predictive performance and methodological quality of existing risk prediction models for IBP before colonoscopy.

Methods

PubMed, Embase, Web of Science, the Cochrane Library, CINAHL, CNKI, Wanfang, and VIP were systematically searched from inception to December 16, 2024, and updated on December 16, 2025. Data extraction was guided by the Checklist for Critical Appraisal and Data Extraction for Systematic Reviews of Prediction Modelling Studies (CHARMS). Areas under the receiver operating characteristic curves (AUCs) with 95% confidence intervals from internal and external validations were pooled separately using random-effects models (Stata 18). Risk of bias and applicability were evaluated with PROBAST. The protocol was registered in PROSPERO (CRD42024627756).

Results

Thirty-one studies reporting 46 prediction models were included, with sample sizes ranging from 152 to 23,456 participants and IBP event incidence rates of 6.3%–50.0%. Frequently used predictors included constipation, diabetes, age, body mass index (BMI), and history of colorectal surgery. Pooled AUCs were 0.76 (95% CI: 0.73–0.79) for internal validation and 0.72 (95% CI: 0.68–0.76) for external validation, both with substantial heterogeneity (I² > 90%). External validation was reported in ten studies, comprising 15 external validation cohorts, including four independent external validation studies of previously published models. Calibration was infrequently reported. Most studies were at high risk of bias, mainly in the analysis domain, while applicability concerns were generally low.

Conclusions

Existing IBP prediction models show moderate to good discrimination on average, but pooled estimates are accompanied by substantial heterogeneity and widespread risk of bias, and robust external validation remains limited. Future studies should standardize outcome definitions and reporting and conduct large, multicenter, independent external validations to improve clinical utility.

1. Introduction

Colorectal cancer (CRC) remains a major global health concern, currently ranking third in incidence and second in cancer-related mortality worldwide [1]. Colonoscopy is widely regarded as the gold standard for CRC screening and prevention because it enables the detection and removal of precancerous lesions, including colorectal adenomas [2, 3]. However, the diagnostic accuracy and procedural safety largely depend on the quality of bowel preparation. Although international guidelines commonly recommend keeping the rate of inadequate bowel preparation (IBP) below 10%, IBP remains common in clinical practice, with reported incidence rates of 18%–35% even with split-dose regimens [4, 5]. IBP reduces adenoma detection rates, prolongs procedure time, increases patient discomfort, and frequently leads to repeat colonoscopies, thereby increasing healthcare utilization and patient burden [68].

Evidence suggests that enhanced patient education, optimized bowel preparation protocols, and reminder-based follow-up can reduce IBP by improving adherence and quality of preparation [911]. However, implementing intensive strategies for all patients is resource-demanding and may be inefficient, potentially increasing the burden on low-risk populations. In addition, real-world implementation may vary across clinicians and healthcare institutions. Moreover, commonly used bowel cleanliness scales, such as the Boston Bowel Preparation Scale (BBPS), assess achieved cleansing at colonoscopy and are not designed for pre-procedural risk prediction, thereby limiting opportunities for proactive optimization [4, 12]. Therefore, risk prediction tools before colonoscopy may help estimate an individual’s probability of developing IBP and support targeted counseling, preparation regimen selection, and follow-up. Clinical prediction models may provide a practical approach by combining readily available predictors to quantify IBP risk in advance and inform individualized preparation planning [13].

Despite the increasing number of IBP risk prediction models, the evidence base remains fragmented, and their clinical applicability remains uncertain. Existing models differ in outcome definitions, predictor selection, modeling methods, and performance reporting, which limits comparability across studies [14]. Furthermore, many models have been evaluated primarily in development cohorts or internal validations, with limited reports of independent external validation, which introduces uncertainty regarding their generalizability across settings and populations [1416]. Importantly, previous reviews primarily summarized individual risk factors rather than quantitatively synthesizing and comparing model performance metrics, which limits their value in informing model selection and updating in clinical settings [17, 18].

Therefore, we conducted a comprehensive systematic review and meta-analysis of IBP risk prediction models before colonoscopy, synthesizing the evidence by summarizing model characteristics, quantitatively synthesizing predictive performance, and appraising methodological quality and reporting transparency to inform model selection and future improvements. The findings will provide valuable references for model selection, updating, and future model development.

2. Methods

This review was registered with PROSPERO (CRD42024627756). It was conducted in accordance with the PRISMA 2020 statement [19] and the TRIPOD-SRMA guidelines [20] to ensure methodological rigor and transparency in the systematic evaluation of prediction models.

2.1. Search strategy

A comprehensive literature search was conducted across eight databases: PubMed, Web of Science, Embase, Cochrane Library, CINAHL, CNKI, Wanfang, and VIP. The time frame spanned from the inception of each database to December 16, 2024, and was updated on December 16, 2025. A combination of MeSH and free-text terms was used to identify studies related to risk prediction models for IBP. Keywords included, but were not limited to, “Colonoscopy,” “Endoscopy,” “Bowel preparation,” “Bowel cleansing,” “Inadequate bowel preparation,” “Risk prediction model,” “Risk factors,” “Predictors,” “Model,” and “Risk score.” Reference lists of included articles and relevant systematic reviews were also manually screened to identify additional eligible studies. We did not impose any language restrictions on the literature searches. For potentially eligible non-English articles, full texts were translated using machine translation (e.g., DeepL/Google Translate) and checked by two reviewers. Any disagreements were resolved by consensus or consultation with a third reviewer when needed. The complete search strategies for each database are detailed in S1 Table.

2.2. Inclusion and exclusion criteria

To guide the formulation of research questions and ensure consistency in eligibility criteria, we adopted the PICOTS framework recommended by the Checklist for Critical Appraisal and Data Extraction for Systematic Reviews of Prediction Modelling Studies (CHARMS) [21].

Studies were included if they met the following criteria: (1) P (Population): Adult patients (≥18 years) undergoing colonoscopy, including both outpatients and inpatients; (2) I (Index model): Studies that developed and/or validated multivariable prediction models for identifying the risk of IBP, regardless of modeling method (e.g., logistic regression, machine learning); (3) C (Comparator): No comparator models were required, as this review did not aim to compare models; (4) O (Outcome): The outcome was IBP before colonoscopy, as defined by the original study using a validated scale (e.g., Boston Bowel Preparation Scale (BBPS), Aronchick Scale); (5) T (Timing): Eligible models must have been constructed using information available before the colonoscopy; studies using intra- or post-procedural predictors were excluded; (6) S (Setting): Studies conducted in secondary or tertiary care settings involving patients undergoing colonoscopy in either outpatient or inpatient contexts; (7) Study design: Observational studies (e.g., cross-sectional, cohort, case-control) that reported model performance metrics such as AUC. We excluded studies based on the following exclusion criteria: (1) Investigated risk factors for IBP without developing or validating a prediction model; (2) Predictor variables < 2; (3) Full text was not available; (4) Duplicate publications.

2.3. Study selection and screening

All references were managed using EndNote X9. Two reviewers independently screened the studies in a standardized, stepwise process. Duplicates were removed first. The remaining records were then assessed based on titles and abstracts to identify potentially eligible studies. Full texts were subsequently retrieved and reviewed against the predefined inclusion and exclusion criteria for final eligibility determination. To enhance comprehensiveness, the reference lists of all included studies and relevant systematic reviews were manually searched to identify additional studies that may have been missed during the initial database search. Any discrepancies in study selection were resolved through discussion with the third author until consensus was reached.

2.4. Data extraction

A standardized data extraction form was developed based on the CHARMS checklist [21]. Extracted information was categorized into two domains: (1) Study characteristics include: first author, year of publication, country, study design, sample size, source of data, participants, outcome assessment tool, IBP prevalence, and bowel preparation regimen. (2) Model characteristics include: methodological details such as handling of continuous variables, treatment of missing data, predictor selection methods, model development, model evaluation, model performance, final predictors included in the model, and model presentation format. Data extraction was performed by one reviewer and independently cross-checked by a second reviewer to ensure accuracy and consistency. Any disagreements were resolved by a third reviewer. For studies with missing or incomplete data, corresponding authors were contacted via email to obtain additional information where possible.

2.5. Risk of bias and certainty of evidence assessment

Two reviewers independently assessed the risk of bias, applicability, and certainty of evidence using the Prediction model Risk Of Bias Assessment Tool (PROBAST) [22, 23] and the GRADE framework [24]. Discrepancies were resolved through discussion with a third reviewer, who made the final judgment when necessary. The PROBAST tool evaluates risk of bias and applicability across four domains: participants, predictors, outcome, and analysis, based on 20 signaling questions. If any domain was rated as “No” or “Probably No,” the study was considered at high risk of bias. If one or more domains were rated as “Unclear” and all other domains were at low risk, the overall risk was judged as unclear. Only when all domains were rated as low risk was the overall study judged to have a low risk of bias. Applicability was assessed using PROBAST across three domains, including participants, predictors, and outcome. Each domain was rated as low concern, high concern, or unclear concern. Overall applicability followed a worst domain rule. Overall applicability was rated as high concern if any domain was rated high concern. Overall applicability was rated as unclear concern when no domain was high concern and at least one domain was unclear concern. Overall applicability was rated as low concern only when all three domains were rated low concern.

The certainty of evidence for pooled AUC values was assessed using the GRADE approach [24, 25], considering five domains: risk of bias, inconsistency, indirectness, imprecision, and publication bias. The overall certainty of the evidence was classified as high, moderate, low, or very low.

2.6. Data synthesis and statistical analysis

All statistical analyses were conducted in Stata version 18. Study and model characteristics were summarized descriptively. Discrimination was assessed using the area under the receiver operating characteristic curve (AUC). We meta-analyzed AUCs separately for results derived from internal validation and external validation to summarize the overall discriminative performance of current IBP prediction models. For internal validation, each eligible model contributed one AUC estimate. For external validation, each distinct validation cohort contributed one effect estimate. When the same model had been externally validated across multiple cohorts/studies, we additionally synthesized those AUCs to summarize model-specific transportability. Before pooling, AUCs were transformed using the logit function and pooled on the logit scale; summary estimates were back-transformed to the original AUC scale for interpretation. Random-effects meta-analyses were fitted using restricted maximum likelihood (REML) to estimate between-study variance, and 95% confidence intervals were calculated using the Hartung–Knapp–Sidik–Jonkman (HKSJ) method to provide more robust inference under heterogeneity [26]. Calibration was not meta-analyzed due to insufficient reporting and was summarized narratively. When quantitative synthesis was not appropriate (e.g., too few estimates or substantial inconsistency), results were presented descriptively.

Heterogeneity was assessed using Cochran’s Q test and the I2 statistic [26]. Publication bias was evaluated using funnel plots and Egger’s regression test for pooled AUCs, in line with GRADE requirements. In accordance with Cochrane guidance, formal publication bias tests were not performed when fewer than 10 studies were included [27]. Sensitivity analyses were conducted using a leave-one-out approach to examine the robustness of the pooled estimates and identify potential influential studies [28]. In cases of substantial heterogeneity (I2 > 50%), subgroup analyses and meta-regression were performed to explore potential sources, with key covariates including country, participants, study design, and other relevant study-level characteristics. Statistical significance was defined as P < 0.05.

3. Results

3.1. Study selection

Fig 1 illustrates the study selection process. A total of 3,011 records were initially retrieved. After removing 1,210 duplicates using EndNote X9, 1,662 records were excluded based on title and abstract screening because they were review articles, irrelevant to the study topic, or otherwise ineligible. The full texts of 139 articles were assessed for eligibility, with 108 studies excluded for reasons including lack of model development, unclear definitions of IBP, or fewer than two predictors. Ultimately, 31 studies were included. Among the included studies, 25 provided sufficient information for inclusion in the meta-analysis.

thumbnail
Fig 1. PRISMA flow diagram for the selection of studies.

https://doi.org/10.1371/journal.pone.0356009.g001

3.2. Study characteristics

Table 1 summarizes the key characteristics of the 31 included studies, covering country, study design, data sources, population, and outcome definition. The studies were published between 2012 and 2025, with 91% published since 2020. The included studies were conducted across 11 countries, with the majority carried out in China (n = 20). Most studies adopted a prospective design (n = 16), followed by retrospective (n = 13) and case–control (n = 2) studies. Data were mainly derived from single-center cohorts (n = 17), with 14 multicenter studies. Several studies focused specifically on outpatient populations (n = 7), inpatients (n = 4), while the rest included general adult colonoscopy populations. The reported proportion of IBP varied widely, ranging from 6.25% to 50%, with sample sizes spanning from 152 to 23,456. IBP was most commonly defined using the Boston Bowel Preparation Scale, with inadequate preparation typically defined as BBPS < 6 and/or any segment score < 2. One study used a modified Aronchick scale, and another used a predefined four-level scale.

thumbnail
Table 1. Overview of the basic data of the included studies.

https://doi.org/10.1371/journal.pone.0356009.t001

3.3. Model development and validation characteristics

Table 2 summarizes the characteristics of model development, validation, and performance evaluation. Most studies employed logistic regression for model construction, while a minority applied machine learning methods such as support vector machines (SVM), decision trees (DT), and extreme gradient boosting (XGBoost). All studies reported performance metrics, primarily the AUC, which ranged from 0.62 to 0.902. Calibration assessment was reported in 18 studies, most commonly using the Hosmer–Lemeshow test, whereas 13 studies did not report calibration assessment. Model presentation varied across studies, including scoring tools, nomograms, risk equations, and mobile applications. Regarding validation, among the 31 included studies, 19 reported both model development and internal validation, 6 conducted external validation following development, 2 focused solely on model development, and 4 were dedicated to external validation of existing models.

thumbnail
Table 2. Overview of the information on the included prediction models.

https://doi.org/10.1371/journal.pone.0356009.t002

Across models, the number of final predictors ranged from 3 to 14, which were broadly grouped into patient characteristics, bowel preparation-related factors, and comorbidities. The most frequently included predictors were constipation (n = 21), diabetes (n = 19), age (n = 14), body mass index (BMI) (n = 9), and history of colorectal surgery (n = 8). Detailed information is presented in S2 Table and Fig 2.

thumbnail
Fig 2. Summary of final predictors included in the prediction model studies.

https://doi.org/10.1371/journal.pone.0356009.g002

3.4. Meta-analysis results

Overall, 31 studies were included in the systematic review, and 25 contributed data to the meta-analysis. Studies were excluded from the meta-analysis if they reported model development without validation, or if the information required to derive the standard error of the AUC was unavailable. Ultimately, a total of 17 studies based on internal validation datasets (31 models) and 10 studies based on external validation datasets (15 models) were included in the meta-analysis. Some studies contributed to both internal- and external-validation syntheses. Meta-analyses were conducted separately for internal- and external-validation results.

3.4.1. Models with internal validation cohorts.

Internally validated model results were pooled using a random-effects meta-analysis, which yielded a pooled AUC of 0.76 (95% CI, 0.73–0.79) (Fig 3), with substantial heterogeneity ( = 95.2%, P < 0.001). Prespecified subgroup analyses were conducted to explore potential sources of heterogeneity (S3 Table).

thumbnail
Fig 3. Forest plot of pooled AUCs for internal validation (random-effects).

https://doi.org/10.1371/journal.pone.0356009.g003

By modeling method, machine learning-based models (n = 13) had a pooled AUC of 0.75 (95% CI, 0.71–0.79;  = 95.0%), whereas non-machine learning models (n = 18; predominantly logistic regression) had a pooled AUC of 0.76 (95% CI, 0.72–0.80;  = 95.0%). No evidence of subgroup differences was observed by the modeling method (P for interaction = 0.745). Discrimination varied by study design (P for interaction < 0.001). Prospective cohort studies (n = 23) yielded a pooled AUC of 0.77 (95% CI, 0.73–0.80), compared with 0.70 (95% CI, 0.65–0.74) for retrospective cohort studies (n = 6). Case-control studies reported a higher pooled AUC of 0.87 (95% CI, 0.83–0.91), but this estimate was based on a small number of models (n = 2) and should be interpreted cautiously. When stratified by predictor composition, models using only patient-related predictors yielded a pooled AUC of 0.73 (95% CI, 0.69–0.77), whereas models combining patient- and preparation-related predictors yielded a pooled AUC of 0.77 (95% CI, 0.74–0.81). Subgroup differences by predictor composition were statistically significant (P for interaction = 0.039). The preparation-only category was represented by a single model and was therefore not informative for between-subgroup inference.

In contrast, subgroup differences were not significant by participant type (P = 0.439) or center type (P = 0.158). Sample size showed a borderline subgroup effect (P = 0.061), with higher pooled AUCs in models with ≥500 participants. Notably, heterogeneity remained substantial within most subgroups, indicating that multiple factors likely contribute to variability in internally validated discrimination.

3.4.2. Models with external validation cohorts.

Using a random-effects meta-analysis, externally validated results across 15 validation cohorts yielded a pooled AUC of 0.72 (95% CI, 0.68–0.76) (Fig 4). Substantial heterogeneity was observed (I² = 93.9%, P < 0.001), and externally validated AUCs varied widely across cohorts, ranging from approximately 0.57 to 0.902. These findings indicated considerable between-cohort variability in discrimination.

thumbnail
Fig 4. Forest plot of pooled AUCs for external validation (random-effects).

https://doi.org/10.1371/journal.pone.0356009.g004

Prespecified subgroup analyses were conducted across 15 external validation cohorts to explore potential sources of heterogeneity (S4 Table). Subgroup differences were statistically significant by bowel preparation regimen (P for interaction = 0.002) and predictor composition (P for interaction = 0.008). Cohorts using a free-choice volume regimen showed the lowest pooled discrimination, with a pooled AUC of 0.67 (95% CI, 0.57–0.76). In contrast, fixed-volume PEG regimens showed higher pooled discrimination, with pooled AUCs of 0.73 (95% CI, 0.66–0.81) for 3 L PEG and 0.72 (95% CI, 0.64–0.79) for 4 L PEG. The 2 L PEG category included only one external validation cohort, which reported an AUC of 0.82 (95% CI, 0.78–0.86); therefore, this estimate should be interpreted cautiously.

When stratified by predictor composition, cohorts evaluating models with patient-related predictors only yielded a pooled AUC of 0.69 (95% CI, 0.63–0.75), whereas models incorporating combined predictors yielded a higher pooled AUC of 0.77 (95% CI, 0.72–0.82); The preparation-related predictors only category included a single external validation cohort, which reported an AUC of 0.80 (95% CI, 0.76–0.84), limiting interpretation of this category. No statistically significant subgroup differences were observed according to participant setting (P = 0.061), study design (P = 0.488), sample size (P = 0.746), or center type (P = 0.398). Heterogeneity remained substantial within most subgroups, indicating that multiple factors likely contribute to variability in externally validated discrimination.

Among externally validated models, only a limited number were evaluated across multiple independent cohorts. Pooled results for repeatedly externally validated models (e.g., Dik and Gimeno-García) are presented in S1 Fig. Meta-regression results suggested that predictor composition was associated with heterogeneity (β = −0.118; P = 0.031). Sample size also showed a statistically significant association (β = 0.0002; P = 0.047), although the effect size was negligible (S5 Table). Given the limited number of external validation cohorts and sparse categories, these meta-regression findings should be interpreted as exploratory.

3.4.3. Comparison of pooled AUCs from internal and external validation.

In the subgroup comparison by validation type, internally validated results yielded a pooled AUC of 0.76 (95% CI, 0.73–0.79) with substantial heterogeneity ( = 95.2%, P < 0.001). Externally validated results yielded a pooled AUC of 0.72 (95% CI, 0.68–0.76), also with substantial heterogeneity ( = 93.9%, P < 0.001). The test for subgroup differences indicated no significant difference between internal and external validation (P = 0.144), suggesting broadly comparable discrimination across validation types despite marked within-group heterogeneity (S4 Table). The pooled AUC across all validation results was 0.75 (95% CI, 0.72–0.77), but this overall estimate should be interpreted as a descriptive summary given the methodological differences between internal and external validation.

3.5. Sensitivity analysis and publication bias

Sensitivity analyses showed that, for both internal and external validation, the pooled AUC estimates remained stable when each study was sequentially excluded (S2 Fig, S3 Fig). The variations were minimal and fell within the 95% confidence intervals of the overall estimates, indicating that the results were robust and not unduly influenced by any single study.

Egger’s tests for both internal and external validation studies indicated no significant publication bias (P = 0.233 and P = 0.960, respectively). The funnel plots appeared approximately symmetric (S4 Fig, S5 Fig).

3.6. Risk of bias and evidence certainty assessment

Risk of bias and applicability were assessed using the PROBAST tool (Table 3) [22, 23]. All included studies were judged to have a high overall risk of bias, indicating methodological shortcomings in the development or validation of the prediction models.

thumbnail
Table 3. PROBAST results of the included studies.

https://doi.org/10.1371/journal.pone.0356009.t003

In the Participants domain, nine retrospective studies were judged at high risk of bias because participants were assembled from pre-existing records without clearly defined consecutive (or population-based) sampling, raising concerns about selection bias and limited representativeness [37, 41, 42, 44, 46, 48, 49, 53, 54]. In the predictor domain, the risk of bias was rated high, as no studies reported quality control measures for predictor assessment. For studies lacking specified assessment methods or clear definitions of predictors, the associated risk of bias was rated unclear. For the outcome domain, one study was rated as high risk due to the use of a non-standardized, investigator-defined bowel preparation scale [15]. Although the remaining studies used standardized tools to assess bowel preparation quality, none explicitly reported blinded outcome assessment, resulting in an overall “unclear” risk rating in this domain.

All studies were rated as high risk of bias in the analysis domain. Several recurring methodological limitations were identified. Five studies failed to meet the recommended minimum events-per-variable (EPV ≥ 20) criterion [29,35,36,48,49], raising concerns about model overfitting. Twelve studies categorized continuous variables without providing a clear rationale or justification [16,30,38,40,42,45,46,49,5457]. Eight studies did not report appropriate handling of missing data [16, 29, 35, 36, 43, 45, 49, 54], and more than half of the studies used univariate analysis to select predictors. Additionally,  thirteen studies did not assess model calibration and reported only discrimination metrics such as the AUC [15,16,29,3234,42,47,4951,53,55]. Eleven studies conducted internal validation using random data splitting, a method that may underestimate the risk of overfitting. In terms of applicability, six studies were judged to raise high concerns because they included only subjects aged ≥ 60 years old or ≥ 40 years old, which restricted the sample case-mix and thus limited generalizability. The remaining studies were considered to raise low concerns regarding applicability.

The certainty of evidence for the pooled AUCs from external validation was assessed using the GRADE framework [25]. All three pooled estimates, including those for all models combined, Dik’s model, and Gimeno-García’s model, were rated as having very low certainty. The primary reasons for downgrading included a high risk of bias and substantial inconsistency across studies. In some cases, concerns related to indirectness and imprecision were also present (S6 Table).

4. Discussion

Bowel preparation quality is a critical determinant of colonoscopy effectiveness, as IBP directly compromises the accuracy and completeness of the procedure [58]. Identifying patients at high risk of IBP before colonoscopy is therefore important for delivering targeted interventions and optimizing preparation pathways. In this systematic review and meta-analysis of 31 studies encompassing multivariable prediction models for IBP, discrimination was moderate to good on average in both internal and external validations, with reported AUCs ranging from 0.620 to 0.902 across studies. Overall, these findings suggest that pre-procedural risk stratification may inform individualized preparation and follow-up by targeting enhanced support for patients at higher risk of IBP. However, PROBAST assessments indicated that all studies were at high risk of bias, largely due to limitations in the analysis domain, and independent external validation was often lacking, which may limit model reliability and transportability in routine practice.

In quantitative synthesis, we separately aggregated the discriminatory capabilities of internal and external validation. The pooled AUC for internal validation model results was 0.76 (95% CI, 0.73–0.79), while the pooled AUC for external validation results was 0.72 (95% CI, 0.68–0.76). Although the difference was not statistically significant (P = 0.144), the trend toward slightly higher performance in internal validation is consistent with optimism commonly observed in development or internal validation settings [60], which may reflect overfitting and differences in case-mix and outcome ascertainment. This underscores the necessity of independent external validation for assessing a model’s generalizability and clinical utility. Notably, both pooled results exhibited substantial heterogeneity (I² > 90%), indicating significant variability in model performance across different cohorts. The very high heterogeneity indicates that the pooled AUC represents an average across highly diverse settings rather than a single transportable estimate.

Exploratory subgroup analyses and meta-regression suggested that heterogeneity in both internal and external validation performance may be partly related to predictor composition. Models incorporating only patient-related factors exhibited relatively lower AUC compared with models that additionally included preparation-related factors. This may relate to preparation-related variables capturing process signals closer to the outcome that are amenable to intervention. Moreover, externally validated performance varied across PEG volume regimens (2 L/ 3 L/ 4 L), raising the possibility that variation in volume-based protocols contributes to between-cohort differences in model performance. Given the limited sample sizes in certain subgroups, these findings should be interpreted with caution and require further confirmation in larger, multi-center, independent validation studies.

From the perspective of predictor composition, models that combined baseline patient characteristics (e.g., age, diabetes, constipation history) with preparation-process variables tended to show higher discrimination than models relying on patient factors alone. Patient factors primarily reflect underlying susceptibility (e.g., age, diabetes, history of constipation), while preparation-related factors indicate modifiable aspects of clinical pathways and implementation processes (e.g., time interval between preparation and examination, preparation type and dosing regimen, patient compliance, and degree of bowel clearance). Such process variables are more closely consistent with the mechanisms underlying outcomes. They not only enhance the discriminatory power in external validation but also provide more direct guidance for nursing interventions and process optimization measures, thereby improving the model’s feasibility for implementation. Future models should standardize the collection and incorporate preparatory process variables based on patients’ baseline risk factors, and assess their incremental value for discrimination, calibration, and clinical net benefit across different healthcare settings.

In the subgroup analysis of external validation, we stratified patients according to PEG bowel preparation solution volume regimens. The results demonstrated significant differences in model discrimination between volume regimens (P = 0.002). Specifically, free-choice volume protocols were associated with lower discriminative performance than fixed-volume PEG regimens, with the lowest pooled AUC observed in free-choice cohorts. Several mechanisms may explain this variation. First, the volume of bowel preparation solution itself may serve as an important predictor of IBP [59]. Second, the preparation-solution volume may affect patient adherence, which in turn influences bowel cleanliness and may compromise predictive accuracy [60]. Although models developed under standardized dosing conditions demonstrated better performance, such uniform regimens are often impractical in real-world settings where individualized dosing is more common. Based on these considerations, we recommend that future studies develop prediction models under diverse dosing contexts or explicitly incorporate preparation volume as a predictor variable to enhance their clinical adaptability and applicability.

To assess the clinical utility of existing IBP risk prediction models and identify those with potential for broader implementation, we systematically compared models that had undergone external validation. Although some models (e.g., Dik’s model and Gimeno’s model) have undergone multiple external validations [34, 51] and reported moderate to good AUC values in individual studies, these models generally showed reduced predictive accuracy across studies and were constrained by methodological limitations. Besides, most models underwent only a single, small-sample external validation and were judged to have a high risk of bias. According to GRADE assessments, none of the models reached moderate or high certainty levels. Taken together, the current evidence does not support universal adoption of a single IBP model across settings. Future research should prioritize large-scale, high-quality, multicenter external validation studies, with careful consideration of the original development context to ensure model applicability and generalizability.

While some IBP prediction models demonstrated acceptable discriminative performance, their overall methodological quality remains suboptimal. According to the PROBAST assessment, most studies revealed a high risk of bias across key domains, including participants, predictors, outcomes, and analysis. Common methodological concerns included small sample sizes with low event-per-variable ratios (EPV < 20), inappropriate handling of missing data, the absence of external validation, and insufficient reporting of essential model performance metrics, such as discrimination and calibration. GRADE assessments further rated the overall certainty of evidence as “very low.” These findings are consistent with prior systematic reviews of IBP prediction models in colonoscopy and further underscore the importance of adopting standardized, methodologically sound practices in future model development and validation [61].

In model evaluation, we observed that existing models generally exhibit shortcomings in reporting standards, with very few studies explicitly adhering to the Transparency Reporting of Multivariate Prediction Models for Individual Prognosis or Diagnosis (TRIPOD) statement. This insufficient reporting undermines the transparency of model development and validation, increasing uncertainty and potentially amplifying the risk of bias. On the one hand, many studies provide insufficient reporting of key methodological details, such as missing data handling and treatment of predictors, making it difficult to assess potential issues within the models, including overfitting, information leakage, or selective reporting. Moreover, the frequent absence of critical reproducible information and comprehensive performance metrics (e.g., calibration) hinders the independent replication and validation of models. This lack of transparency compromises not only the implementability of individual models but also the validity of cross-study comparisons and evidence synthesis, ultimately limiting the value of systematic reviews for informing clinical practice. Future research should strictly adhere to the TRIPOD reporting guidelines to mitigate associated risks, reduce bias, and enhance the certainty of evidence for clinical decision-making.

5. Implications

This study systematically evaluated the performance and quality of the IBP risk prediction model, providing a reference for its evidence base in clinical application and future research. Early identification of patients at high risk for IBP is essential for developing tailored health education and intervention strategies. Such efforts can enhance bowel preparation adequacy and improve the overall quality of colonoscopy. Risk prediction models for IBP may serve as valuable tools to assist pre-procedural assessments, particularly in settings with high patient volumes and limited staffing, by supporting data-informed decision-making and improving workflow efficiency. However, given the current limitations, such as insufficient external validation and limited generalizability, these models should be considered as complementary tools rather than substitutes for clinical judgment. Future research should prioritize large, multicenter prospective external validations, integrate modifiable preparation-related process variables, and report full performance metrics (including calibration and clinical utility) in accordance with TRIPOD to improve transparency, transportability, and real-world applicability.

6. Strengths and limitations

To our knowledge, this is the first systematic review and meta-analysis to quantitatively synthesize discrimination performance (AUC) of multivariable prediction models for inadequate bowel preparation. Following the guidelines of PRISMA and TRIPOD-SRMA, we systematically summarized the characteristics of model development and used PROBAST to assess the risk of bias, providing information for model selection and future optimization.

Nevertheless, several limitations should be acknowledged. First, most included studies were rated as high risk of bias, with generally suboptimal methodological quality, resulting in an overall low certainty of evidence. Second, our meta-analysis focused primarily on discrimination because AUC was the most consistently reported performance measure across the included studies. Calibration information was frequently unavailable or inconsistently reported, limiting our ability to assess whether predicted risks agreed with observed risks. Accordingly, the pooled AUC estimates should be interpreted as summaries of discriminative performance rather than comprehensive evaluations of model performance or clinical applicability. Third, variations in outcome definitions and assessment tools for IBP across studies may have introduced additional heterogeneity, limiting comparability and interpretability. Finally, some studies contributed multiple model estimates, which may challenge the independence assumption of conventional meta-analytic pooling. These limitations highlight the need for future high-quality, standardized studies to confirm and extend our findings.

7. Conclusion

This systematic review included 31 studies reporting 46 IBP prediction models. The findings indicate that existing models demonstrate moderate to good discriminatory performance (AUC) in both internal and external validations, suggesting their potential value in risk stratification and informing individualized preparation pathways. However, confidence in their clinical use is limited by a high risk of bias (mainly in the analysis domain), limited independent external validation, and incomplete TRIPOD reporting. Future studies should prioritize large, prospective multicenter external validation and transparent reporting to support reliable implementation.

Supporting information

S1 Fig. Meta-analysis results for repeatedly externally validated models.

https://doi.org/10.1371/journal.pone.0356009.s001

(DOCX)

S2 Fig. Leave-one-out sensitivity analysis (internal validation).

https://doi.org/10.1371/journal.pone.0356009.s002

(DOCX)

S3 Fig. Leave-one-out sensitivity analysis (external validation).

https://doi.org/10.1371/journal.pone.0356009.s003

(DOCX)

S4 Fig. Publication bias assessment (internal validation).

https://doi.org/10.1371/journal.pone.0356009.s004

(DOCX)

S5 Fig. Publication bias assessment (external validation).

https://doi.org/10.1371/journal.pone.0356009.s005

(DOCX)

S1 Table. Searching strategy(full electronic search strategies).

https://doi.org/10.1371/journal.pone.0356009.s006

(DOCX)

S2 Table. Predictors included in each model (final predictors).

https://doi.org/10.1371/journal.pone.0356009.s007

(DOCX)

S3 Table. Meta-analysis and subgroup analyses of internal validation results.

https://doi.org/10.1371/journal.pone.0356009.s008

(DOCX)

S4 Table. Meta-analysis and subgroup analyses of external validation results.

https://doi.org/10.1371/journal.pone.0356009.s009

(DOCX)

S5 Table. Meta-regression for external validation results.

https://doi.org/10.1371/journal.pone.0356009.s010

(DOCX)

S6 Table. GRADE assessment of certainty of evidence.

https://doi.org/10.1371/journal.pone.0356009.s011

(DOCX)

S8 Table. Study selection details for included and excluded records.

https://doi.org/10.1371/journal.pone.0356009.s013

(DOCX)

Acknowledgments

We are grateful to Professor De-liang Cao and Professor Xi Zeng (Cancer Research Institute, Hengyang Medical School, University of South China) for their advice on manuscript writing. We also gratefully acknowledge Dr. Qi Liu (School of Nursing, The Hong Kong Polytechnic University) for her contribution in editing the manuscript.

References

  1. 1. Bray F, Laversanne M, Sung H, Ferlay J, Siegel RL, Soerjomataram I. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(3):229–63.
  2. 2. Pilonis ND, Bugajski M, Wieszczy P, Franczyk R, Didkowska J, Wojciechowska U, et al. Long-Term Colorectal Cancer Incidence and Mortality After a Single Negative Screening Colonoscopy. Ann Intern Med. 2020;173(2):81–91. pmid:32449884
  3. 3. Săftoiu A, Hassan C, Areia M, Bhutani MS, Bisschops R, Bories E, et al. Role of gastrointestinal endoscopy in the screening of digestive tract cancers in Europe: European Society of Gastrointestinal Endoscopy (ESGE) Position Statement. Endoscopy. 2020;52(4):293–304. pmid:32052404
  4. 4. Hassan C, East J, Radaelli F, Spada C, Benamouzig R, Bisschops R. Bowel preparation for colonoscopy: European Society of Gastrointestinal Endoscopy (ESGE) Guideline– Update 2019. Endoscopy. 2019;51(8):775–94.
  5. 5. Kilgore TW, Abdinoor AA, Szary NM, Schowengerdt SW, Yust JB, Choudhary A, et al. Bowel preparation with split-dose polyethylene glycol before colonoscopy: a meta-analysis of randomized controlled trials. Gastrointest Endosc. 2011;73(6):1240–5. pmid:21628016
  6. 6. Clark BT, Rustagi T, Laine L. What level of bowel prep quality requires early repeat colonoscopy: systematic review and meta-analysis of the impact of preparation quality on adenoma detection rate. Am J Gastroenterol. 2014;109(11):1714–23; quiz 1724. pmid:25135006
  7. 7. Kingsley J, Karanth S, Revere FL, Agrawal D. Cost Effectiveness of Screening Colonoscopy Depends on Adequate Bowel Preparation Rates - A Modeling Study. PLoS One. 2016;11(12):e0167452. pmid:27936028
  8. 8. Pantaleón Sánchez M, Gimeno Garcia A-Z, Bernad Cabredo B, García-Rodríguez A, Frago S, Nogales O, et al. Prevalence of missed lesions in patients with inadequate bowel preparation through a very early repeat colonoscopy. Dig Endosc. 2022;34(6):1176–84. pmid:35189669
  9. 9. Ramprasad C, Saini D, Del Carmen H, Krasnovsky L, Chandra R, Mcgregor R, et al. Text Message System for the Prediction of Colonoscopy Bowel Preparation Adequacy Before Colonoscopy: An Artificial Intelligence Image Classification Algorithm Based on Images of Stool Output. Gastro Hep Adv. 2024;4(2):100556. pmid:39866713
  10. 10. Ganayem R, Alamour O, Cohen DL, Ealiwa N, Abu-Freha N. Enhancing Patient Education for Colonoscopy Preparation: Strategies, Tools, and Best Practices. J Clin Med. 2025;14(12):4375. pmid:40566123
  11. 11. Wang F, Huang X, Wang Z, Yan Z, Wang S, Pan P, et al. One-day versus three-day low-residue diet bowel preparation regimens before colonoscopy: a meta-analysis of randomized controlled trials. J Gastroenterol Hepatol. 2024;39(5):787–95. pmid:38251810
  12. 12. Lai EJ, Calderwood AH, Doros G, Fix OK, Jacobson BC. The Boston bowel preparation scale: a valid and reliable instrument for colonoscopy-oriented research. Gastrointest Endosc. 2009;69(3 Pt 2):620–5. pmid:19136102
  13. 13. Chen L, Kang X, Ren G, Luo H, Zhang L, Wang L, et al. Individualized intervention based on a preparation-related prediction model improves adequacy of bowel preparation: A prospective, multi-center, randomized, controlled study. Dig Liver Dis. 2024;56(3):436–43. pmid:37735023
  14. 14. Fostier R, Tziatzios G, Facciorusso A, Papaefthymiou A, Arvanitakis M, Triantafyllou K, et al. Models and scores to predict adequacy of bowel preparation before colonoscopy. Best Pract Res Clin Gastroenterol. 2023;67:101859. pmid:38103925
  15. 15. Hassan C, Fuccio L, Bruno M, Pagano N, Spada C, Carrara S, et al. A predictive model identifies patients most likely to have inadequate bowel preparation for colonoscopy. Clin Gastroenterol Hepatol. 2012;10(5):501–6. pmid:22239959
  16. 16. Dik VK, Moons LMG, Hüyük M, van der Schaar P, de Vos Tot Nederveen Cappel WH, Ter Borg PCJ, et al. Predicting inadequate bowel preparation for colonoscopy in participants receiving split-dose bowel preparation: development and validation of a prediction score. Gastrointest Endosc. 2015;81(3):665–72. pmid:25600879
  17. 17. Beran A, Aboursheid T, Ali AH, Albunni H, Mohamed MF, Vargas A, et al. Risk Factors for Inadequate Bowel Preparation in Colonoscopy: A Comprehensive Systematic Review and Meta-Analysis. Am J Gastroenterol. 2024;119(12):2389–97. pmid:39225554
  18. 18. Mahmood S, Farooqui SM, Madhoun MF. Predictors of inadequate bowel preparation for colonoscopy: a systematic review and meta-analysis. Eur J Gastroenterol Hepatol. 2018;30(8):819–26. pmid:29847488
  19. 19. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. PLoS Med. 2021;18(3):e1003583. pmid:33780438
  20. 20. Snell KIE, Levis B, Damen JAA, Dhiman P, Debray TPA, Hooft L, et al. Transparent reporting of multivariable prediction models for individual prognosis or diagnosis: checklist for systematic reviews and meta-analyses (TRIPOD-SRMA). BMJ. 2023;381:e073538. pmid:37137496
  21. 21. Moons KGM, de Groot JAH, Bouwmeester W, Vergouwe Y, Mallett S, Altman DG, et al. Critical appraisal and data extraction for systematic reviews of prediction modelling studies: the CHARMS checklist. PLoS Med. 2014;11(10):e1001744. pmid:25314315
  22. 22. Chen R, Wang SF, Zhou JC, Sun F, Wei WW, Zhan SY. Introduction of the prediction model risk of bias assessment tool: a tool to assess risk of bias and applicability of prediction model studies. Zhonghua Liu Xing Bing Xue Za Zhi. 2020;41(5):776–81. pmid:32447924
  23. 23. Wolff RF, Moons KGM, Riley RD, Whiting PF, Westwood M, Collins GS, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51–8. pmid:30596875
  24. 24. Guyatt GH, Oxman AD, Vist GE, Kunz R, Falck-Ytter Y, Alonso-Coello P, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336(7650):924–6. pmid:18436948
  25. 25. Foroutan F, Mayer M, Guyatt G, Riley RD, Mustafa R, Kreuzberger N, et al. GRADE concept paper 8: judging the certainty of discrimination performance estimates of prognostic models in a body of validation studies. J Clin Epidemiol. 2024;170:111344. pmid:38579978
  26. 26. Higgins JPT, Thompson SG. Quantifying heterogeneity in a meta-analysis. Stat Med. 2002;21(11):1539–58. pmid:12111919
  27. 27. Higgins J, Thomas J, Chandler J, Cumpston M, Li T, Page M. Cochrane handbook for systematic reviews of interventions version 6.5 (updated August 2024). Cochrane. 2024.
  28. 28. StataCorp. Meta-analysis reference manual. Stata Press. 2021.
  29. 29. Berger A, Cesbron-Métivier E, Bertrais S, Olivier A, Becq A, Boursier J, et al. A predictive score of inadequate bowel preparation based on a self-administered questionnaire: PREPA-CO. Clin Res Hepatol Gastroenterol. 2021;45(4):101693. pmid:33852957
  30. 30. Chen L, Ren G, Luo H, Zhang L, Wang L, Zhao J, et al. Superiority of a preparation-related model for predicting inadequate bowel preparation in patients undergoing colonoscopy: A multicenter prospective study. J Gastroenterol Hepatol. 2022;37(12):2297–305. pmid:36181263
  31. 31. Fuccio L, Frazzoni L, Spada C, Mussetto A, Fabbri C, Manno M, et al. Factors that affect adequacy of colon cleansing for colonoscopy in hospitalized patients. Clin Gastroenterol Hepatol. 2021;19(2):339-348.e7. pmid:32200083
  32. 32. Gu F, Xu J, Du L, Liang H, Zhu J, Lin L, et al. The machine learning model for predicting inadequate bowel preparation before colonoscopy: a multicenter prospective study. Clin Transl Gastroenterol. 2024;15(5):e00694. pmid:38441136
  33. 33. Gimeno-García AZ, Baute JL, Hernandez G, Morales D, Gonzalez-Pérez CD, Nicolás-Pérez D, et al. Risk factors for inadequate bowel preparation: a validated predictive score. Endoscopy. 2017;49(6):536–43. pmid:28282690
  34. 34. Gkolfakis P, Kapizioni C, Tziatzios G, Facciorusso A, Frazzoni L, Thomopoulos K, et al. Comparative performance and external validation of three different scores in predicting inadequate bowel preparation among Greek inpatients undergoing colonoscopy. Ann Gastroenterol. 2023;36(1):25–31. pmid:36593808
  35. 35. Malkin D, Cohen DL, Richter V, Ariam E, Vosko S, Shirin H, et al. A novel model to predict inadequate bowel preparation prior to colonoscopy incorporating patients’ reactions to drinking the laxative. J Clin Med. 2023;12(23):7335. pmid:38068387
  36. 36. Okamoto N, Tanaka S, Abe H, Miyazaki H, Nakai T, Tsuda K, et al. Predicting Inadequate Bowel Preparation When Using Sodium Picosulfate plus Magnesium Citrate for Colonoscopy: Development and Validation of a Prediction Score. Digestion. 2022;103(6):462–9. pmid:36380621
  37. 37. Sninsky JA, Toups V, Cotton C, Peery AF, Arora S. An electronic medical record prediction model to identify inadequate bowel preparation in patients at outpatient colonoscopy. Tech Innov Gastrointest Endosc. 2024;26(2):130–7. pmid:38911129
  38. 38. Zhang N, Xu M, Chen X. Establishment of a risk prediction model for bowel preparation failure prior to colonoscopy. BMC Cancer. 2024;24(1):341. pmid:38486227
  39. 39. Zhao X, Pan Y, Hao J, Feng J, Cui Z, Ma H, et al. Development and validation of a novel scoring system based on a nomogram for predicting inadequate bowel preparation. Clin Transl Oncol. 2024;26(9):2262–73. pmid:38565812
  40. 40. Liu Y, Liu XQ, Yang XN, Wang P, Liu XK, Luo D. Construction and validation of a risk predictive model for the bowel preparation failure in colonoscopy patients. Chinese Journal of Nursing. 2024;59(9):1091–8.
  41. 41. Huang MQ, Han DJ, Zhu HJ, Yang F. Analysis of predictive factors and development of model predicting inadequate intestinal preparation in outpatients undergoing colonoscopy by clinical pharmacists. Pharmaceutical and Clinical Research. 2023;31(2):172–6.
  42. 42. Li JM, Liu TW, Fu SY, Li Y, Lin YF, Zhang BP. Using the optimal subset method to establish a prediction model for bowel preparation. Chinese Journal of Practical Internal Medicine. 2020;40(3):231–6.
  43. 43. Rao W, Peng SY, Xu T, Re ZY, Huang XL. Analysis of predictive factors and model establishment of inadequate bowel preparation. Modern Digestion & Intervention. 2021;26(11):1378–83.
  44. 44. Wang GH, Chen J, Sheng ZJ, Xi MJ, Zhou YT. Establishing and evaluating a risk prediction model for colonoscopy bowel preparation failure based on automated machine learning. China Journal of Endoscopy. 2024;30(5):36–47.
  45. 45. Guo SL, Zhu T, Lin WN, Chen XR, Chen ZQ, Xia MY. Establishment and validation of risk predictive model for inadequate bowel preparation before colonoscopy in the elderly. Chinese Nursing Research. 2023;37(3):392–8.
  46. 46. Wang JX, Suo LN, Yu FF. Factors influencing quality of bowel preparation before colonoscopy in elderly patients and construction and validation of a risk prediction model for bowel preparation failure. Chin J Coloproctol. 2024;44(11):51–4.
  47. 47. Kutyla MJ, O’Connor S, Hourigan LF, Kendall B, Whaley A, Meeusen V, et al. An Evidence-based Approach Towards Targeted Patient Education to Improve Bowel Preparation for Colonoscopy. J Clin Gastroenterol. 2020;54(8):707–13. pmid:31764487
  48. 48. Zhao YY, Li ZQ, Ma X, Gao F. Establishment and validation of nomograms prediction model for bowel preparation quality in inpatients. Modern Digestion & Intervention. 2022;27(9):1100–5.
  49. 49. Song ZH, Zhang PK, Gao XZ. Development and validation of a risk nomogram for inadequate bowel preparation. Chin J Gastrointestinal Endoscopy (Electronic Edition). 2024;11(2):100–4.
  50. 50. Afecto E, Ponte A, Fernandes S, Gomes C, Correia JP, Carvalho J. Validation and Application of Predictive Models for Inadequate Bowel Preparation in Colonoscopies in a Tertiary Hospital Population. GE Port J Gastroenterol. 2021;30(2):134–40. pmid:37008528
  51. 51. Yuan X, Gao H, Liu C, Wang W, Xie J, Zhang Z, et al. External validation of two prediction models for adequate bowel preparation in Asia: a prospective study. Int J Colorectal Dis. 2022;37(6):1223–9. pmid:35467123
  52. 52. Fang L, Bingbing W, Wang Q, Yinchuan X. Development and Validation of a Nomogram Predictive Model and Scoring Tool for Assessing the Risk of Inadequate Bowel Preparation in Colonoscopy Patients Over 40 Years Old: A Retrospective Observational Study. J Clin Nurs. 2025;34(10):4398–414. pmid:40011667
  53. 53. Apaer Z, Tian X, Wei R, Feng G, Ling HX. External validation of a prediction model of inadequate bowel preparation in Xinjiang. Chin J Gastroenterol Hepatol. 2024;33(4):385–9.
  54. 54. Xu MM, Fu XR, Zhang N, Wang Y, Shi FZ. Development and validation of a risk score model for inadequate bowel preparation for colonoscopy in elderly patients. Chinese Journal of Nursing. 2022;57(11):1337–44.
  55. 55. Liu J, Jiang W, Yu Y, Gong J, Chen G, Yang Y, et al. Applying machine learning to predict bowel preparation adequacy in elderly patients for colonoscopy: development and validation of a web-based prediction tool. Ann Med. 2025;57(1):2474172. pmid:40065741
  56. 56. Wu H, Wu R, Zhang Y, Lu X, Zhao W, Xu B, et al. Development and validation of a gut motility based model for predicting bowel preparation quality. Sci Rep. 2025;15(1):28265. pmid:40753140
  57. 57. Yin H, Wang Y, Wang H, Li T, Xu X, Li F, et al. Derivation and validation of a prediction model for inadequate bowel preparation in Chinese outpatients. Sci Rep. 2025;15(1):1430. pmid:39789134
  58. 58. Rex DK, Anderson JC, Butterly LF, Day LW, Dominitz JA, Kaltenbach T, et al. Quality indicators for colonoscopy. Gastrointest Endosc. 2024;100(3):352–81. pmid:39177519
  59. 59. Shahini E, Sinagra E, Vitello A, Ranaldo R, Contaldo A, Facciorusso A, et al. Factors affecting the quality of bowel preparation for colonoscopy in hard-to-prepare patients: Evidence from the literature. World J Gastroenterol. 2023;29(11):1685–707. pmid:37077514
  60. 60. Waldmann E, Penz D, Majcher B, Zagata J, Šinkovec H, Heinze G, et al. Impact of high-volume, intermediate-volume and low-volume bowel preparation on colonoscopy quality and patient satisfaction: An observational study. United European Gastroenterol J. 2019;7(1):114–24. pmid:30788123
  61. 61. Guo SL, Zhu T, Lin WN, Chen XR, Chen ZQ, Xia MY. Risk prediction model for inadequate bowel preparation before colonoscopy: a systematic review. Journal of Nursing. 2022;29(1):35–40.