Fig 1.
This study (orange) originates from a source review (blue). Articles included in the source review were used for data extraction with ChatGPT-4o. The ’Consensus on extracted data’ from the source review was considered as the reference standard, against which ChatGPT-4o’s responses were tested for validity and reproducibility (black).
Table 1.
Response assessment categories.
Fig 2.
ChatGPT-4o response assessment.
Evaluation of ChatGPT-4o responses compared to the reference standard. Response categories included ’Completely correct,’ ’Satisfactory,’ ’Lacking information,’ ’Missing all information,’ and ’False information’.
Fig 3.
Stratification of ChatGPT-4o responses by outcome reporting status.
The proportion of ChatGPT-4o responses in each response category stratified based on whether the articles reported the outcome or not.
Fig 4.
Proportion of correct and satisfactory ChatGPT-4o responses across four data domains.
The proportion of ’Completely correct’ and ’Satisfactory’ responses from ChatGPT-4o compared to the reference standard across four data domains: General information, Population, Intervention and comparator, and Outcome.
Table 2.
Overall and domain-specific results of reproducibility of data extraction by ChatGPT-4o.