Figures
Abstract
This paper examines how successive Intergovernmental Panel on Climate Change (IPCC) synthesis reports represent scenario knowledge when they are queried through the same scenario-based questions. The purpose is not to test whether the reports predict the same future, because IPCC scenarios are conditional what-if explorations rather than forecasts. Instead, the paper asks how the evidentiary, causal, and uncertainty structure of the reports changes across assessment cycles. We analyze six synthesis reports from the First Assessment Report to the Sixth Assessment Report, using a retrieval-augmented generation (RAG) pipeline and layered analytical prompts. Four theoretically selected pillars organize the comparison: mitigation-adaptation pathway divergence, emerging technologies and scale-up constraints, compound socio-economic risks and governance stress, and carbon dioxide removal feasibility with policy lock-in risks. The results show substantial continuity in the high-level logic of climate assessment: all reports treat delayed mitigation as increasing future risk and all reports caution that technological potential depends on policy, finance, infrastructure, and institutional conditions. The strongest change lies in representation. Early reports often rely on qualitative conditional reasoning; later reports increasingly use quantitative signposts, calibrated uncertainty language, integrated socio-economic pathways, and explicit treatment of limits to adaptation and net-zero pathways. For example, carbon dioxide removal is nearly absent as a policy-relevant pathway in the early assessments, appears in AR5 as a model-dependent negative-emissions assumption, and becomes in AR6 a necessary but limited component of net-zero pathways whose overuse can create mitigation-deterrence and lock-in risks. The paper contributes a replicable method for auditing longitudinal consistency in large assessment corpora and a substantive account of how IPCC scenario knowledge has moved from descriptive climate futures toward integrated, uncertainty-calibrated, and policy-conditioned knowledge representation.
Citation: Warin T, Bisson C (2026) AI-assisted longitudinal comparison of scenario knowledge representation in IPCC synthesis reports. PLOS Clim 5(7): e0000965. https://doi.org/10.1371/journal.pclm.0000965
Editor: Jingyu Wang, Nanyang Technological University, SINGAPORE
Received: December 22, 2025; Accepted: May 28, 2026; Published: July 14, 2026
Copyright: © 2026 Warin, Bisson. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All relevant data are within the paper and its Supporting Information files.
Funding: The author(s) received no specific funding for this work.
Competing interests: The authors have declared that no competing interests exist.
Introduction
Climate-change risks are now discussed not only as physical changes in temperature, precipitation, and sea level, but also as interacting pressures on ecosystems, infrastructure, health, finance, and governance. Recent climate assessments and indicator-based syntheses show that the scale and interdependence of these risks have become more visible over time [1,2]. In this context, climate scenarios are indispensable tools for organizing uncertainty. They do not predict the future. They describe conditional futures: if emissions, technologies, institutions, and socio-economic trajectories follow specified assumptions, then certain climate and impact pathways become more or less plausible.
This distinction matters for the present study. The IPCC defines scenarios as plausible descriptions of how the future may develop, based on internally coherent assumptions about driving forces and relationships [3]. Across its assessment cycles, the IPCC has used different scenario architectures, from early emissions pathways and the Special Report on Emissions Scenarios to Representative Concentration Pathways and Shared Socioeconomic Pathways [4–8]. These frameworks differ in their assumptions, variables, and policy relevance. Consequently, comparing IPCC reports is not a simple exercise in checking whether they said the same thing at different moments. It is a comparison of how the IPCC has represented climate futures under changing scientific evidence, changing modelling capacity, and changing policy contexts.
Comparing the six synthesis reports is important for three reasons. First, the reports are among the most authoritative cumulative assessments of climate science, impacts, adaptation, and mitigation. Second, policy actors often cite IPCC conclusions across time as if the underlying scenario categories were directly interchangeable. Third, longitudinal comparison can reveal whether apparent divergence reflects disagreement, improved evidence, revised scenario architecture, or a new policy problem that earlier assessments were not yet positioned to address.
The task is also difficult. The reports differ in length, terminology, structure, uncertainty conventions, and baseline assumptions. The early assessments were written before calibrated confidence language, before the RCP/SSP architecture, before the widespread use of net-zero framing, and before carbon dioxide removal became central to mitigation pathways. Manual comparison remains possible but becomes increasingly fragile as the corpus grows and as the relevant evidence is distributed across thousands of pages and multiple assessment styles.
A retrieval-augmented generation approach is therefore a logical methodological choice. RAG allows the analyst to hold the question constant while grounding each answer in text retrieved from a particular report [9,10]. The method does not ask a language model to supply climate knowledge from its parametric memory. It asks the model to organize and compare passages from a fixed corpus under a stable analytical protocol. Layered prompts further separate analytical tasks, such as identifying causal claims, classifying uncertainty, extracting mechanisms, and comparing the result with later scenario frameworks. This makes the comparison more transparent than an unconstrained summary and more semantically flexible than simple keyword counts.
Existing comparative IPCC research has usually focused on particular report components, sectors, or modelling parameters. Fløttum et al. [11], for instance, compared topics and frames across the AR5 summaries for policymakers. Jaczewski et al. [12] compared temperature indices under SRES scenarios for Poland. Luo et al. [13] compared radiative-forcing climate footprints calculated under AR5 and AR6 models. These studies are valuable, but they do not provide a longitudinal, all-assessment comparison of how scenario knowledge itself is represented in the synthesis reports. Nor do they combine RAG with a fixed set of analytical prompts to test cross-report consistency.
This study addresses that gap. It compares the six IPCC synthesis reports from 1990 to 2023 through four scenario pillars. The pillars were selected because together they cover the main analytical space of climate scenario use: the strategic relationship between mitigation and adaptation; the socio-technical feasibility of changing emissions and resilience trajectories; the interaction between climate hazards and socio-economic governance systems; and the feasibility and risks of carbon dioxide removal in net-zero pathways. These pillars are not intended to exhaust all IPCC content. They are designed to be sufficient for answering the three research questions because each question concerns longitudinal consistency, depth of analysis, and policy implications in the representation of scenario knowledge.
The first pillar, mitigation-adaptation divergence, was selected because it expresses the core conditional structure of climate policy. Mitigation changes future forcing and therefore future risk; adaptation changes the exposure and vulnerability of systems that face those risks. The relationship between the two is not static. Delayed mitigation can increase adaptation needs and residual damages, while adaptation limits can make mitigation more urgent. A prompt that asks about inflection points up to 2050 therefore tests whether a report represents climate futures as pathway-dependent rather than as a linear accumulation of impacts.
The second pillar, technology and scale-up constraints, was selected because scenario analysis often turns on feasibility assumptions. Technologies may be physically possible, economically competitive under certain conditions, socially contested, infrastructure-dependent, or limited by supply chains and institutions. This pillar therefore tests whether the reports represent technological change as a list of options or as a socio-technical process embedded in finance, policy, infrastructure, innovation systems, and social acceptance. It also captures a historically important shift from broad technology categories to system-level transformation.
The third pillar, compound socio-economic risk and governance stress, was selected because climate risk is increasingly understood as systemic. A heatwave is not only a meteorological event; its consequences depend on health systems, energy systems, housing, labour conditions, inequality, emergency management, and governance capacity. This pillar tests whether the reports represent climate impacts as sectoral damages or as interacting risks mediated by institutions and social vulnerability. It is therefore central to the paper’s claim that later assessments show greater causal integration.
The fourth pillar, CDR feasibility and policy lock-in, was selected because it is the strongest test of topic emergence and intertemporal policy dependence. CDR is now central to net-zero pathways, but it was not central to early assessment frameworks. Asking all reports the same CDR question makes the temporal asymmetry visible. It also tests whether reports represent future removals as a technical possibility, a model assumption, a governance challenge, or a source of mitigation-deterrence risk. No other pillar in the design produces as clear a contrast between early silence, intermediate model dependence, and later policy salience.
The comparison is therefore not a historical curiosity. It bears directly on the way scientific assessments are used in public decision-making. IPCC synthesis reports circulate beyond the scientific community: they are quoted in national adaptation plans, corporate transition strategies, litigation, financial-risk disclosure, education, and media accounts of climate futures. In those settings, readers often treat the reports as cumulative statements of one evolving body of knowledge. That intuition is broadly correct, but it can become analytically misleading when the scenario language of one assessment cycle is read as if it were directly equivalent to the scenario language of another. A stabilization scenario in the early assessment period, an SRES storyline in the TAR and AR4 period, an RCP-based pathway in AR5, and an SSP-RCP pathway in AR6 do not merely differ by label. They embody different assumptions about socio-economic development, emissions, land use, technology, policy timing, and the representation of risk.
This is why the object of analysis in the present article is scenario knowledge representation. By representation we mean the form through which climate futures are made intelligible in the reports: the type of evidence offered, the causal relations emphasized, the degree of quantification, the form of uncertainty language, and the extent to which physical, technological, socio-economic, and governance dimensions are integrated. The same underlying proposition may be represented in different ways across assessment cycles. For example, the claim that delayed mitigation raises future risks may appear in an early report as a qualitative conditional statement, in AR4 as a quantified risk-and-cost trade-off, in AR5 as a carbon-budget problem, and in AR6 as a pathway problem linked to overshoot, adaptation limits, equity, and net-zero feasibility. The continuity of the proposition does not remove the importance of the changing representational apparatus.
A longitudinal comparison also has a practical value for scenario users. It helps distinguish three situations that are easily conflated. The first is substantive convergence, where successive reports answer a question in broadly similar terms despite changes in language. The second is representational deepening, where later reports retain the same general conclusion but express it with more evidence, uncertainty calibration, causal detail, and policy relevance. The third is genuine topic emergence, where a later report addresses a matter that earlier reports could not yet assess in a comparable form. Carbon dioxide removal is the clearest example in the present study. In such cases, the absence of a topic from early assessments should not be read as an IPCC judgment that the topic was negligible. It may indicate that the modelling, empirical literature, or policy architecture needed to assess the topic had not yet matured.
The paper therefore adopts a conservative interpretation of divergence. A difference between reports is not automatically treated as a contradiction. It may reflect a new scenario architecture, a different assessment mandate, improved evidence, a changed policy environment, or the adoption of calibrated uncertainty conventions. This position is essential for avoiding a retrospective reading in which AR6 categories are imposed uncritically on earlier reports. At the same time, the paper does not make the opposite mistake of treating all differences as merely semantic. The use of identical prompts across report-specific retrieval windows makes it possible to detect where later reports provide genuinely new kinds of analytical content, especially in relation to adaptation limits, CDR dependence, compound risks, and governance stress.
The four-pillar design follows from this interpretation. Mitigation-adaptation divergence tests how reports represent conditional pathway forks. Technology and scale-up constraints test how reports represent socio-technical feasibility rather than technical potential alone. Compound socio-economic risks and governance stress test how reports represent risk propagation through social and institutional systems. Carbon dioxide removal and lock-in test how reports represent the temporal displacement of mitigation effort into future removal obligations. Together, the four pillars do not claim to exhaust the IPCC corpus. They provide a theoretically sufficient cross-section of scenario knowledge because they cover pathway structure, feasibility, systemic risk, and intertemporal policy dependence, which are the dimensions most relevant to the research questions.
The use of RAG follows the same logic of restraint. A general-purpose language model asked to summarize the IPCC could easily produce fluent but untraceable answers. In the present design, the model is not allowed to function as a climate expert. It is used as an evidence-organizing device. Each answer is conditioned on retrieved passages from one report at a time. The same question is repeated across reports, and the subsequent comparison is performed on the resulting report-specific outputs. This procedure cannot eliminate all interpretive risk, but it reduces a major source of variation: the analyst does not change the question while moving from one report to another. The procedural constancy of the prompt helps expose the historical variability of the reports themselves.
The paper asks three research questions. First, to what extent do IPCC synthesis reports provide consistent answers to the same high-level scenario prompts, and where do their answers converge or diverge? Second, how has the depth of analysis changed across assessment cycles, especially in the use of quantitative evidence, causal mechanisms, and calibrated uncertainty language? Third, what do these changes imply for the use of climate scenarios in research and policy, particularly when users compare conclusions across reports produced under different scenario architectures?
The remainder of the paper is organized as follows. The next section reviews literature on climate scenario frameworks, IPCC comparison, and NLP-based analytical prompting. The methods section then specifies the scoping review protocol, the document corpus, the scenario prompts, the RAG pipeline, the contextual prompting design, and the comparative coding approach. The results section presents cross-report findings with an emphasis on the evolution of scenario knowledge representation. The discussion connects these results to scenario use, methodological contribution, limitations, and indicator design for future assessments. The conclusion summarizes the contribution and identifies next steps for AI-assisted assessment research.
Literature review
Review design and search protocol
The literature review was designed as a scoping review rather than a full systematic review. This choice is appropriate because the objective is to map an emerging interdisciplinary field at the intersection of climate scenario assessment, IPCC comparison, foresight, NLP, and RAG, not to estimate an intervention effect or exhaustively synthesize a homogeneous body of empirical evidence [14–17]. The protocol follows the logic of scoping review methodology: identify the research question, identify relevant studies, screen for conceptual and methodological relevance, chart the evidence, and synthesize the literature narratively.
Searches were conducted in Scopus on 12 April 2026 using title, abstract, and keyword fields. The search strings combined terms related to climate scenarios, assessment frameworks, IPCC comparisons, NLP, prompts, RAG, and foresight. The initial counts were used to map the state of the field. Broad searches such as natural language processing and prompt produced large results, while climate-scenario and RAG combinations produced no title or abstract matches in the Scopus query table. After screening, the final qualitative synthesis retained 34 sources: 24 peer-reviewed articles or proceedings papers, one scholarly book, and nine IPCC reports or guidance documents. The retained sources were included because they directly informed scenario frameworks, uncertainty language, comparative IPCC analysis, AI-assisted foresight, RAG, retrieval, or prompt-based reasoning.
The inclusion criteria were English-language scholarly sources, IPCC or IPCC-related institutional sources, and works directly relevant to at least one of four categories: climate scenario frameworks, comparative analysis of IPCC materials, NLP/RAG/prompting for document analysis, or foresight and scenario methodology. The exclusion criteria were sectoral impact papers that used IPCC scenarios only as exogenous inputs without comparing assessment logic, general AI papers without relevance to document-grounded analysis, news items, non-scholarly commentary, and duplicate records. Quality appraisal was based on peer-reviewed status, methodological transparency, conceptual relevance, and fit with the research questions. The synthesis method was thematic: sources were grouped into scenario frameworks, uncertainty and assessment language, IPCC comparison, and AI-assisted document analysis.
The search-map reported in Table 1 is used in this article as a positioning instrument rather than as a claim of exhaustive bibliometric coverage. It identifies the relative density of adjacent literatures and, more importantly, the scarcity of work at their intersection. The large number of results for natural language processing and prompting confirms that prompt-based text analysis is no longer marginal. The low or null counts for combinations involving climate scenarios, IPCC, RAG, and layered analytical prompts indicate that the specific object of the present article remains underdeveloped. This contrast helps justify the methodological contribution without requiring the literature review to become a comprehensive review of all NLP research or all climate-scenario research.
Screening proceeded in two stages. The first stage removed records that used the relevant keywords only incidentally, for example articles that mentioned IPCC scenarios as inputs for a local impact model but did not address assessment logic, scenario representation, uncertainty language, or document comparison. The second stage retained works that directly informed at least one element of the research design: the historical development of scenario frameworks, the treatment of uncertainty in assessment reports, methods for comparing IPCC materials, foundations of RAG and semantic retrieval, or the use of generative AI in foresight and document analysis. This two-stage strategy is appropriate for a scoping review because the purpose is to clarify an interdisciplinary problem-space rather than to calculate a pooled effect size.
The quality appraisal was correspondingly pragmatic but explicit. IPCC reports and IPCC guidance documents were treated as primary institutional sources for scenario definitions, uncertainty language, and assessment conventions. Peer-reviewed journal articles and conference proceedings were preferred for methodological and comparative claims. Foundational books were retained when they established scenario-planning concepts or information-retrieval principles that remain in use. Works whose main relevance was speculative, journalistic, or primarily promotional were excluded unless they served only as examples of a wider concern, such as the risk of hallucination in generative systems. This hierarchy of evidence is important because the article uses AI-assisted analysis to study authoritative scientific reports; it would be inconsistent to ground that methodological choice in weak or anecdotal sources.
The synthesis was narrative and thematic. Sources were not aggregated by statistical effect but by conceptual function. The scenario-framework literature explains why the reports cannot be compared as if they used a stable scenario taxonomy. The uncertainty literature explains why changes in confidence language are part of the phenomenon being studied. The comparative-IPCC literature shows that previous work has addressed selected report components, topics, or model outputs but has not provided a prompt-controlled longitudinal analysis of scenario knowledge representation across all synthesis reports. The NLP and RAG literature explains how retrieval grounding and structured prompting can support a disciplined form of assisted document analysis.
Climate scenarios and IPCC assessment frameworks
Scenario planning has a long tradition in strategic foresight and climate assessment [18,19]. In climate research, scenarios are most useful when they clarify the consequences of assumptions rather than imply unconditional prediction. Early IPCC assessments relied on relatively simple emissions scenarios and qualitative descriptions of impacts and responses. Subsequent cycles incorporated SRES families, RCPs, and SSPs, enabling more explicit links among radiative forcing, socio-economic development, mitigation challenges, adaptation challenges, and risk distribution.
Moss et al. [5] described the transition toward a new generation of climate scenarios designed to support integrated analysis of mitigation, adaptation, impacts, and vulnerability. Van Vuuren et al. [6] summarized the RCP architecture as a set of pathways spanning different 2100 radiative-forcing levels. O’Neill et al. [7] and Riahi et al. [8] developed the SSP narratives and associated energy, land-use, and emissions implications, making socio-economic assumptions more explicit. These developments changed not only the content of climate scenarios but also the grammar of assessment: reports could increasingly express conditional futures in terms of pathways, feasibility, risks, limits, and co-produced socio-economic futures.
The treatment of uncertainty also changed. IPCC uncertainty guidance formalized calibrated confidence and likelihood language to communicate evidence and agreement across working groups [20–22]. This development is crucial for the present comparison because early reports may express caution or conditionality without using the formal terms that later reports use. A longitudinal study must therefore avoid treating the absence of calibrated language in early reports as equivalent to the absence of uncertainty. It should instead compare how uncertainty is represented under the conventions available in each assessment cycle.
The transition from emissions scenarios to more integrated pathway architectures is central to the present argument. In the early assessment period, scenarios primarily organized assumptions about emissions trajectories and the associated physical climate response. The SRES framework later introduced internally coherent socio-economic storylines, but these storylines were not designed to span the full contemporary policy grammar of mitigation ambition, adaptation limits, net-zero timing, overshoot, and CDR. The RCP framework then separated radiative-forcing pathways from socio-economic assumptions, while the SSP framework reconnected climate futures to development narratives, mitigation challenges, and adaptation challenges. Each transition enlarged the set of questions that an assessment could answer.
This enlargement should not be understood as a simple movement from imprecision to precision. Earlier assessments were not deficient merely because they lacked later vocabulary. They operated under different evidentiary and modelling constraints. The methodological problem is that readers today often encounter the entire sequence of reports from the standpoint of AR6. Terms such as net zero, overshoot, residual emissions, maladaptation, soft and hard limits to adaptation, and CDR feasibility now structure the conversation. When these terms are absent from early reports, one must ask whether the relevant issue was conceptually absent, expressed under another vocabulary, or not yet supported by a sufficiently developed assessment literature. The RAG design allows that distinction to be examined systematically by asking the same general question while allowing each report to answer in its own terminology.
The scenario literature also suggests that the policy relevance of scenarios depends on the credibility of their assumptions. Scenarios are not evaluated only by internal coherence. They are also evaluated by their usefulness for decision-making under uncertainty. This distinction is particularly relevant for technologies and CDR. A pathway may be internally coherent in a model while remaining institutionally, materially, or politically difficult to implement. The later IPCC reports increasingly attend to such feasibility questions. This movement is visible in AR5 and AR6, where mitigation pathways are discussed not only as emissions trajectories but also as sets of assumptions about land, energy systems, finance, governance, distributional effects, and technological deployment.
Uncertainty language is one of the most visible markers of representational change. The IPCC’s calibrated language does not simply add rhetorical caution. It provides a standardized way to connect evidence, agreement, confidence, and likelihood. Earlier reports expressed uncertainty, but often through ordinary-language caution or conditional phrasing. Later reports increasingly distinguish what is likely, what is highly confident, what is model-dependent, and what remains uncertain because evidence is limited or agreement is low. The present study therefore treats uncertainty articulation as an analytical dimension rather than as a secondary stylistic feature. A report that contains the same substantive claim but embeds it in calibrated uncertainty language represents scenario knowledge differently from a report that states the claim qualitatively.
The reasons-for-concern tradition illustrates the same pattern. As climate risk assessment matured, the IPCC increasingly represented risk through categories that could connect physical thresholds to ecosystems, extreme events, distributional impacts, aggregate damages, and large-scale singular events. This did not eliminate uncertainty; it provided a more communicable structure for uncertainty. The evolution from broad impact categories toward risk frameworks is important for P3, because compound risk and governance stress require a vocabulary in which hazards, exposure, vulnerability, and institutional capacity can be considered together. The later reports are therefore expected to show greater causal integration, not merely more references to more sectors.
NLP, RAG, and analytical prompting in scientific texts
Natural language processing methods have increasingly been used to extract, classify, and summarize complex scientific and policy texts. Traditional techniques such as term-frequency analysis, topic modelling, and vector-space retrieval can identify lexical patterns, but they often struggle to reconstruct causal argumentation and implicit uncertainty [23]. Recent transformer-based language models can generate more context-sensitive analyses, but unconstrained generation risks hallucination and weak provenance [9,10]. RAG addresses this problem by combining generative models with explicit retrieval from a controlled text corpus.
RAG was introduced for knowledge-intensive NLP tasks as a way to combine parametric language models with non-parametric retrieved evidence. Subsequent work showed that retrieval augmentation can reduce hallucination in knowledge-grounded generation. For the present study, this architecture is not used to automate scientific judgment. It is used to structure document analysis: each report is queried through the same prompts, and the model output is constrained by retrieved passages from that report.
Prompting research is also relevant. Chain-of-thought and related structured prompting approaches show that complex tasks can be decomposed into intermediate analytical steps [24]. In this study, the layered prompts do not ask the model to reveal free-form reasoning. They ask it to produce auditable analytical outputs: causal claims, evidence, uncertainty categories, mechanisms, external consistency judgments, and research gaps. The resulting outputs become intermediate objects for human comparison, not substitutes for expert interpretation. This positioning is consistent with recent AI-assisted scenario and foresight work, which treats generative AI as a support for synthesis, exploration, or misinformation checking rather than as an autonomous source of domain authority [25–28].
RAG is especially suitable for this task because the corpus is both authoritative and heterogeneous. The reports are authoritative because they are produced through formal assessment procedures; they are heterogeneous because their length, internal structure, terminology, and conventions differ substantially across time. A purely generative model would risk importing knowledge from outside the report being analyzed. A retrieval-and-generation design addresses this problem by first selecting report-specific evidence windows and then constraining the language model to organize those retrieved passages into a structured analytical answer. Because the executed scripts use sparse tf-idf retrieval and do not implement dense embeddings, the method is reported as sparse retrieval-augmented generation rather than hybrid retrieval.
The role of prompting is therefore methodological rather than ornamental. The prompt is the instrument that defines the analytical task. In a conventional manual content analysis, the coding manual specifies the categories through which coders identify, classify, and compare textual evidence. In the present study, layered prompts play a similar role. They separate the identification of causal claims from the classification of uncertainty, the reconstruction of mechanisms, the evaluation of external consistency, and the formulation of research gaps. This separation reduces the likelihood that one broad summarizing prompt will collapse evidence, inference, and interpretation into a single opaque paragraph.
The design also deliberately avoids treating model fluency as evidence. A fluent model output is not, by itself, a finding. It becomes analytically useful only when it can be traced back to retrieved passages and when the same procedure is applied across comparable documents. This point is important because generative AI has often been criticized for hallucination, overgeneralization, and weak provenance. The present approach does not remove those risks, but it changes the burden of proof. Outputs are treated as intermediate analytical artefacts. Claims that are unexpected, highly specific, or central to the interpretation must be checked against the source text. The final result is therefore not the model’s answer; it is the authors’ comparison of source-grounded outputs.
A further advantage of the layered design is that it makes absence analytically meaningful while still requiring caution. If the same prompt retrieves little relevant evidence from an early report but retrieves detailed evidence from AR6, that pattern may indicate the emergence of a new assessment topic. Yet the conclusion remains conditional on retrieval quality and corpus choice. The study therefore distinguishes between silence in a synthesis report and silence in the broader assessment literature. A topic that is absent from a synthesis report may have appeared in a Working Group chapter. The synthesis-report corpus was chosen for comparability, not because it contains every relevant statement in IPCC history.
Finally, the approach creates a bridge between qualitative interpretation and operational reproducibility. It does not reduce the IPCC reports to token counts, but it does impose stable analytical procedures. Another research team could use the same documents, prompts, retrieval settings, and coding matrix to examine whether the same cross-report patterns emerge. This is a modest but important contribution, particularly for meta-assessment research. Large assessment corpora are too important to be compared only through impressionistic reading, but they are too complex to be responsibly summarized by unconstrained automation.
Comparative IPCC research and the remaining gap
Previous comparative work has examined particular IPCC materials or scenario uses. Fløttum et al. [11] compared topics and frames across AR5 summaries for policymakers. Luo et al. [13] compared radiative-forcing climate footprint calculations under AR5 and AR6 models in the agricultural sector. Jaczewski et al. [12] compared temperature indices under three SRES scenarios for Poland. These studies demonstrate that IPCC-based comparison is useful, but their object of comparison differs from ours. They compare frames, parameters, or sectoral outputs. The present study compares the representation of scenario knowledge across all synthesis reports.
The specific gap is therefore methodological and substantive. Methodologically, there is limited work using RAG to query successive scientific assessments through identical prompts. Substantively, there is limited longitudinal analysis of how scenario knowledge evolves from qualitative conditional statements to calibrated, quantitative, and pathway-sensitive representations. This paper addresses both gaps.
This gap is not merely technical. It concerns the sociology of assessment knowledge. IPCC reports are iterative documents, but each assessment cycle is also shaped by the literature available at that moment, the division of labour among working groups, the policy questions occupying governments, and the methodological conventions accepted by the scientific community. A longitudinal comparison therefore needs to ask how the assessment genre itself changes. The synthesis report in 1990 did not perform the same communicative task as the synthesis report in 2023. Both are authoritative, but they organize authority differently. The earlier reports establish the basic scientific and policy problem; the later reports increasingly manage a dense landscape of pathways, feasibility conditions, residual risks, and implementation constraints.
The present study contributes by making these changes observable under a controlled query structure. Rather than reading each report for whatever it happens to emphasize, the method asks each report to speak to the same four scenario problems. This creates a counterfactual comparison of sorts: what happens when an early assessment is asked a question that later became central? The answer may be a partial answer, an indirect answer, or a declaration of insufficient evidence. Each outcome is informative. It shows not only what was known, but how the relevant problem could or could not be represented at that stage of assessment history.
Methods
Research design
The study is a longitudinal qualitative document analysis supported by a retrieval-augmented LLM pipeline. It does not claim inferential statistical testing. Quantification is descriptive and operational: counts of search results, retrieved passages, prompt-output combinations, and coded representational features. This clarification responds to the concern that the statistical analysis should not be overstated. The analytical claim is comparative and interpretive, grounded in a fixed corpus and reproducible prompting procedure.
The unit of analysis is the report-prompt output: one IPCC synthesis report queried through one scenario pillar and one analytical layer. This unit was chosen because it makes comparison symmetrical. The analysis does not compare all paragraphs in FAR with all paragraphs in AR6. It compares how each report responds when placed under the same analytical demand. The design therefore resembles a controlled qualitative comparison more than a conventional topic model. The control is not statistical randomization, but procedural invariance: the same retrieval logic, prompt structure, and coding dimensions are applied across all assessment cycles.
The study also separates three levels of inference. At the first level, retrieved passages indicate what the report makes available for analysis. At the second level, the LLM organizes those passages into claims, mechanisms, uncertainty classifications, and gaps. At the third level, the authors compare outputs across reports and interpret longitudinal patterns. This separation is important because it prevents the model’s intermediate organization from being mistaken for final evidence. The primary evidence remains the IPCC report text. The model output is a structured aid for examining that text, and the comparative interpretation is a human scholarly judgment based on repeated patterns across the corpus.
The descriptive coding matrix was designed to capture representational change rather than substantive agreement alone. A report may converge with later reports on the general direction of a claim while diverging in evidence specificity, uncertainty articulation, causal integration, or scenario-framework alignment. The ordinal scale therefore does not rank reports by quality. It describes the form of representation available in the synthesis report. A score of 1 in an early assessment may be entirely appropriate for the conventions of that period. The value of the scale lies in making visible the historical movement from qualitative conditional statements toward integrated pathway representation.
The workflow is shown in Fig 1. The design holds constant the scenario prompt and analytical prompt while varying the report. This allows differences in output to be interpreted as differences in the retrieved source material and its representational conventions, subject to the limitations discussed later.
Corpus and document preparation
The corpus comprises six IPCC assessment synthesis files: the 1990/1992 First Assessment materials, the 1995 Second Assessment synthesis, the 2001 Third Assessment synthesis, the 2007 Fourth Assessment synthesis, the 2014 Fifth Assessment synthesis, and the 2023 Sixth Assessment synthesis [1,29–33]. The focus on synthesis reports was deliberate. These reports provide the highest-level assessment genre most directly comparable across cycles, even though their structure, terminology, and uncertainty conventions differ across time.
Each report was converted to text with pdftools::pdf_text(), pages were collapsed into a single string, whitespace was normalized, and the document was segmented into deterministic overlapping character chunks. The first-pass execution script used chunk_chars = 1800 and chunk_overlap = 200, which implies a 1,600-character step between chunk starts. The method therefore uses character windows rather than approximate word-count windows. This matters because chunk length determines both retrieval granularity and the probability that a relevant paragraph is split across two retrieval units.
The exclusive focus on synthesis reports has both strengths and costs. Its principal strength is genre comparability. Each synthesis report is designed to condense the assessment cycle into a high-level account of major findings for policymakers and informed users. This makes it a suitable object for longitudinal analysis. The cost is topical compression. Some issues, especially those involving sectoral impacts, regional vulnerability, mitigation technologies, or CDR options, are treated in much greater detail in the underlying Working Group volumes. The present study therefore asks how scenario knowledge is represented in the synthesis reports, not how the entire IPCC literature has treated each topic across all chapters.
Text preparation was treated as a methodological step rather than a clerical operation. Headers, footers, page numbers, repeated captions, and extraction artefacts can bias retrieval because they introduce tokens that are not part of the substantive argument. Deterministic overlapping character segmentation was used to reduce the risk of detaching a claim from its warrant or uncertainty qualifier. The 1,800-character chunk size with 200-character overlap reflects a compromise between retrieval precision, contextual continuity, and manual auditability.
Report-specific indexing is a crucial design feature. A single pooled index across all reports would risk allowing later vocabulary to dominate retrieval, especially for concepts such as SSPs, net zero, overshoot, and CDR. By indexing each report separately, the method asks whether that report contains evidence relevant to the prompt in its own terms. This does not solve every comparability problem, but it prevents the most obvious form of temporal contamination. The comparison is conducted after output generation, not during retrieval.
Scenario prompts and theoretical sufficiency of the four pillars
The executed first-pass code used ten scenario prompts. The present article treats the first four prompts as the focal analytical pillars because they map directly onto the research questions: pathway divergence, socio-technical feasibility, compound socio-economic risk, and carbon-removal or lock-in feasibility. The six additional first-pass prompts were retained as exploratory and triangulation prompts on energy-system transitions, financial-stability channels, urban infrastructure, the food-water-energy nexus, international cooperation, and early-warning signposts. Table 2 gives the four focal prompts and their analytical scope; S1 Text reports the full executed prompt bank so that the distinction between focal analysis and broader generated output is transparent; Table 3 provides the four pillars and their comprehensiveness regarding the research question.
Retrieval-augmented generation pipeline
The implemented retrieval stage used a transparent sparse vector-space procedure. For each IPCC report and each first-pass prompt, the script created a quanteda corpus containing the query and all chunks from that report, tokenized the corpus while removing punctuation and numbers, constructed a document-feature matrix, and applied tf-idf weighting with scheme_tf = “prop” and scheme_df = “inverse.” Relevance was then ranked by cosine similarity between the query row and each chunk row [23]. Dense embeddings were not used in the executed analysis; they remain a possible extension for a labelled future ablation study.
The equations are retained to make the retrieval logic explicit. In the implemented procedure, quanteda computes proportional term frequency and inverse document frequency internally; the equations below state the conceptual weighting and similarity operations used to rank report chunks.
In Equations 1 and 2, is the weight of term t in chunk c, tfprop is proportional term frequency, idf is inverse document frequency, N is the number of chunks in the report-specific index, df_t is the number of chunks containing term t, q is the prompt vector, and c is the chunk vector. The implementation also protects against zero vector norms before computing similarity.
The retrieval design was intentionally report-specific. Each report was indexed and queried separately so that later terminology could not contaminate earlier evidence windows. This is central to the design: an AR6 passage could not be retrieved to answer a FAR prompt. Cross-report comparison occurred only after the report-specific outputs were produced.
The sparse-only retrieval choice has a clear implication for interpretation. It privileges terms that appear explicitly in the reports and therefore keeps retrieval auditable, but it may under-retrieve passages in which an earlier report describes a later concept using different vocabulary. This is most relevant for P4: early reports may discuss sinks, stabilization, forests, or land-use options without using the contemporary term carbon dioxide removal. The limitation is acknowledged in the discussion and partly mitigated by manual review of consequential absence claims.
The top-k evidence window was fixed across report-prompt pairs. The code set top_k_chunks = 8 and max_ctx_chars = 12000. Because chunks are selected only while the cumulative context remains below the character ceiling, the actual number of chunks passed to the model may be fewer than eight when retrieved chunks are long. This fixed ceiling keeps the model input comparable across reports: later reports are not allowed unlimited evidence while earlier reports receive only a few sentences. Where a report contains sparse evidence, the appropriate output is an insufficiency statement rather than an artificially elaborate answer.
Retrieval quality was assessed through plausibility and source checking. Retrieved passages were expected to contain terms or concepts directly related to the scenario pillar. Where the model produced a surprising interpretation, the underlying passages were inspected to determine whether the claim was supported, overstated, or produced by inference. The method therefore combines computational retrieval with human audit. This is essential because RAG does not guarantee correctness. It improves grounding, but a retrieved passage can still be misread or overextended by a generative model.
Contextual LLM prompting and multi-prompt design
The retrieved evidence window was submitted to a locally hosted Ollama-compatible language model through the ollamar R package. The analysis scripts specify the model tag as xara:latest and call generate(host = req_url, model = model, prompt = prompt, stream = FALSE, output = “text”). For security and reproducibility reasons, the private local-area-network address used during development is omitted from the public record. The exact model digest, Ollama version, operating system, hardware configuration, and run dates were not recorded in the execution artifacts and are therefore treated as missing execution metadata. The scripts do not explicitly set temperature, top_p, random seed, repeat penalty, model-level context length, or maximum generated tokens.
Ten second-pass analytical prompts were applied to the first-pass outputs. They covered internal validity, uncertainty classification, causal mechanisms to 2050, external consistency with AR6 constructs, distributional and justice dimensions, evidence sufficiency for policy design, contradictions and tensions, operational early-warning indicators, boundary conditions, and a compact research agenda.
The design generated two structured output files. The first-pass file, answers.csv, has ten prompt rows, six report-answer columns, and two prompt-identification columns, yielding 60 document-specific first-pass answers. The second-pass file, comparative-analyses.csv, has 600 rows, corresponding to ten first-pass prompts, ten analytical prompts, and six documents. The article interprets the first four first-pass prompts as the focal four-pillar analysis; the remaining generated outputs are useful for audit, triangulation, and possible Supporting Information, but they should not be confused with the theoretical core of the article.
The output files confirm that the minimum-length safeguard operated as intended. First-pass answers exceed the 250-word floor, and second-pass analytical outputs are longer still. This does not make word count a measure of quality, but it documents that the cells contain substantive analytical text rather than short completions. The decisive validity criterion remains evidence anchoring and human verification against the retrieved IPCC passages.
The prompt template was intentionally restrictive. It instructed the model to use only the retrieved context, to avoid importing later knowledge unless the external-consistency layer explicitly asked for comparison with AR6 constructs, and to state when evidence was insufficient. This last requirement is important. Without an insufficiency option, the model may produce a plausible answer even when the retrieved evidence is weak. In this study, insufficiency is not treated as failure. It is a legitimate result that helps identify historical gaps in scenario knowledge representation.
The ten analytical layers were designed to separate claims from interpretation. The internal-validity layer asks whether the report provides a causal claim, supporting evidence, and a warrant. The uncertainty layer asks how strongly the claim is qualified. The causal-mechanism layer asks whether the report describes a temporal chain, feedback, threshold, or path dependency. The external-consistency layer asks how the report’s representation relates to later scenario frameworks, while recognizing that earlier reports cannot be expected to use later terminology. The remaining layers address distributional and justice dimensions, evidence sufficiency for policy design, contradictions and rival hypotheses, operational early-warning indicators, boundary conditions, and future research priorities. Taken together, the layers provide a multi-dimensional representation of each report’s answer.
The methodological value of this layered structure is clearest in cases where a simple summary would conceal important differences. Two reports may both say that technology matters, but one may describe only options while another describes cost curves, infrastructure constraints, institutional capacity, and social acceptance. Two reports may both mention risk, but one may list sectoral impacts while another reconstructs cascading mechanisms across systems. The layered prompts are designed to capture these differences without requiring a pre-existing fixed dictionary of terms.
The first-pass script records the model tag, retrieval parameters, prompt templates, minimum word threshold, output path, and error-handling logic. Specifically, it sets model = “xara:latest,” writes the first-pass outputs to rag_outputs/answers.csv, uses chunk_chars = 1800, chunk_overlap = 200, top_k_chunks = 8, max_ctx_chars = 12000, and min_words = 250, and applies a second generation call when an answer falls below the minimum word threshold. The second-pass script uses the same local model tag and endpoint configuration, reads rag_outputs/answers.csv, applies ten analytical prompts to each document-prompt cell, trims long input cells to max_ctx_chars = 12000, enforces the same minimum word threshold, and writes the comparative outputs to rag_outputs/comparative-analyses.csv.
Comparative coding and validation
Outputs were compared through a descriptive coding matrix. Each report-prompt pair was coded on four representational dimensions: specificity of evidence, uncertainty articulation, causal integration, and scenario-framework alignment. The coding scale was ordinal and descriptive: 0 = absent or insufficient evidence; 1 = qualitative representation; 2 = quantitative or uncertainty-calibrated representation; 3 = integrated pathway representation linking quantitative signposts, uncertainty, mechanisms, and socio-economic conditions. The purpose of the scale was to make qualitative comparison more transparent, not to estimate a statistical model.
Validity safeguards included three checks. First, outputs were required to cite or paraphrase only retrieved evidence. Second, unexpected claims were checked against the original report passages. Third, conclusions were based on repeated cross-report patterns, not on isolated model phrasing. The final interpretation therefore combines model-assisted extraction with human assessment of the underlying IPCC texts.
The coding scheme was applied after the model generated report-specific outputs. It was not used to train the model and did not alter retrieval. This sequencing preserves the distinction between extraction and interpretation. First, each report is queried and summarized under the same prompt structure. Second, the outputs are coded for representational dimensions. Third, the coded patterns are interpreted in relation to the research questions. The coding categories were chosen because they correspond directly to the attributes that determine whether a scenario statement is usable for comparison across time: evidence specificity, uncertainty articulation, causal integration, and scenario-framework alignment.
The ordinal scores should be read as transparent descriptors rather than as measurements with interval properties. The difference between 1 and 2 is not assumed to be equal to the difference between 2 and 3. The purpose is to discipline interpretation by making explicit why a later report is judged to provide a more integrated representation. A report coded as 3 on causal integration, for example, does not simply mention more sectors; it connects them through a mechanism that includes physical hazards, exposure, vulnerability, response options, and temporal consequences.
Validation focused on the most consequential claims. Claims were treated as consequential when they supported one of the main longitudinal findings, when they involved quantitative signposts, or when they implied that a topic was absent from an earlier report. Such claims were checked against the retrieved evidence and, where necessary, against the original report location. This procedure does not replace a full double-coded manual content analysis. It does, however, provide a practical validity safeguard for a method whose purpose is to identify broad patterns across long authoritative documents (See Table 4).
Results
Overview: from scenario content to scenario knowledge representation
The results show that the most important longitudinal change is not simply the addition of new scenario content. It is the transformation of how scenario knowledge is represented. Across the assessment history, IPCC synthesis reports move from broad conditional statements about emissions and impacts toward more integrated representations that connect forcing pathways, socio-economic assumptions, adaptation limits, mitigation feasibility, carbon budgets, net-zero constraints, and calibrated uncertainty.
Table 5 summarizes the main convergence and divergence patterns across the four pillars. Table 6 then synthesizes the representational evolution by assessment cycle.
The four coding dimensions clarify this transformation. Evidence specificity increases because later reports more often attach scenario claims to temperature thresholds, emissions trajectories, concentration levels, carbon budgets, adaptation-cost estimates, or technology-feasibility parameters. Uncertainty articulation increases because later reports rely more consistently on calibrated confidence and likelihood language. Causal integration increases because later reports connect physical hazards to socio-economic systems, institutions, policy sequencing, and path dependence. Scenario-framework alignment increases because later reports operate within progressively more explicit scenario architectures. The shift is therefore cumulative but uneven. Not all pillars evolve at the same pace, and not all forms of representation appear together.
The strongest continuity appears in P2. Across the full assessment history, the IPCC consistently resists a purely technology-driven account of climate response. The strongest divergence appears in P4, because CDR becomes central only when stringent mitigation pathways, net-zero framing, and overshoot scenarios make future removals a policy-relevant assumption. P1 and P3 occupy an intermediate position. Both show early conceptual continuity but later representational elaboration. Mitigation-adaptation divergence is present from the beginning as a conditional logic, but it becomes more quantitative and pathway-sensitive over time. Compound risk is also present early in sectoral form, but only later does it become a systemic and governance-mediated risk grammar.
These findings are consistent with the historical development of climate assessment. Early reports had to establish the scientific basis of anthropogenic climate change, broad physical consequences, and the need for mitigation and adaptation. Middle reports increasingly connected emissions scenarios to impacts, vulnerability, costs, and policy response. Later reports operate in a world where climate change is already observed, mitigation pathways are evaluated against temperature goals, adaptation limits are documented, and socio-economic pathways are used to frame both vulnerability and transformation. The language of the reports changes because the assessment problem changes.
P1: Mitigation-adaptation divergence
Prompt P1 asked each report to identify plausible inflection points up to 2050 where mitigation and adaptation pathways structurally diverge. The main convergence is conceptual. All six reports, in different language, represent climate futures through conditional divergence: if emissions continue to rise or mitigation is delayed, future adaptation needs and residual damages increase. This logic is present even when early reports do not use later terms such as carbon budget, overshoot, or adaptation limits.
The main divergence is representational. FAR and SAR tend to express divergence qualitatively, using doubling of CO2, stabilization, and broad impact categories. TAR gives stronger attention to vulnerability and timing but remains transitional. AR4 introduces more explicit thresholds and economic signposts, including mitigation-cost ranges and risk escalation. AR5 strengthens the link between carbon budgets, delayed action, adaptation limits, and negative emissions. AR6 gives the most integrated response: near-term emissions reductions, 1.5 °C and 2 °C thresholds, adaptation limits, overshoot risks, and SSP-conditioned pathways are linked in a single scenario grammar.
The evolution is therefore not a simple shift from wrong to right. It is a shift from qualitative conditional reasoning to quantitative and uncertainty-calibrated pathway reasoning. For users of IPCC scenarios, this means that older reports can support broad historical comparison, but later reports should be used when precise timing, pathway feasibility, or confidence language is required.
In FAR, the representation of divergence is anchored in the relationship between greenhouse-gas accumulation and future physical impacts. The report provides a conditional structure: higher emissions lead to greater warming and therefore greater impacts. However, this structure is not yet expressed through the later policy grammar of adaptation limits, overshoot, residual risk, or carbon budgets. The relevant knowledge representation is therefore qualitative and physically oriented. It establishes the direction of the problem but leaves the timing, socio-economic differentiation, and policy trade-offs relatively underdeveloped.
SAR begins to introduce more policy-relevant signposts. Stabilization, concentration levels, and the timing of emissions change become more visible. The report is still cautious and often qualitative, but it provides a more explicit basis for thinking about divergence before mid-century. Adaptation appears as a necessary response, while mitigation appears as a means of reducing the future scale of adaptation needs. The two are not yet integrated into the kind of pathway comparison that appears in AR5 and AR6, but the basic conditional relationship is already present.
TAR represents an intermediate stage. It gives greater attention to vulnerability, adaptive capacity, and the consequences of delayed action. This matters because divergence is no longer represented only as a physical climate threshold. It is increasingly represented as a relationship between climate hazards and the capacity of human and ecological systems to respond. TAR therefore broadens the causal field, even if it does not yet provide the full quantitative and calibrated representation of later reports.
AR4 marks a stronger transition toward quantified risk and response comparison. It uses more systematic risk framing and more explicit economic and physical signposts. Mitigation costs, stabilization ranges, and reasons-for-concern language make pathway divergence more visible as a policy problem. The report can therefore answer P1 with more than a general warning. It can connect delayed mitigation to higher future risk, higher adaptation burden, and higher probability of severe impacts under particular warming levels. The representation is still less integrated than AR6, but it is more structured than in the early assessments.
AR5 deepens the pathway logic through carbon budgets, RCPs, and a clearer treatment of adaptation limits. The report’s representation of divergence is no longer only that high emissions produce high impacts. It is that different emissions pathways imply different temperature outcomes, different adaptation burdens, different residual risks, and different later reliance on negative emissions. This is a significant representational shift because mitigation timing becomes connected to future technological and adaptation feasibility. A delayed pathway is not simply a later version of an early pathway; it can become structurally different because it demands faster transformation or future removals.
AR6 provides the most integrated representation of P1. It links near-term mitigation, 1.5 °C and 2 °C thresholds, overshoot, adaptation limits, feasibility, equity, and socio-economic pathways. The result is a scenario grammar in which divergence is multidimensional. Pathways diverge not only by emissions and temperature but also by development conditions, institutional capacity, exposure, vulnerability, technology deployment, and residual damage. This makes AR6 the most useful report for current policy users who need to understand how near-term decisions shape mid-century adaptation possibilities. It also explains why AR6 should not be used to retroactively criticize earlier reports for lacking categories that had not yet become available.
The P1 result therefore supports the first and second research questions simultaneously. The reports converge on the broad claim that mitigation delay increases later risk and adaptation burden. They diverge in the depth and form of representation. The evolution is from conditional physical reasoning to integrated pathway reasoning. For scenario users, the implication is that older reports can be used to establish historical continuity, but later reports are required for decisions that depend on timing, thresholds, confidence, and feasibility.
P2: Emerging technologies and scale-up constraints
Prompt P2 showed the strongest cross-report convergence. Each assessment recognizes that technology can alter emissions or resilience trajectories, but none treats technological change as sufficient by itself. The reports consistently identify boundary conditions: cost, infrastructure, learning, policy support, finance, social acceptability, institutional capacity, and market design. This continuity is important because it shows that the IPCC has not moved from technological pessimism to technological optimism; rather, it has increasingly specified the socio-technical conditions under which technologies scale.
The representational change lies in granularity. Early reports discuss energy efficiency, renewables, and nuclear energy in broad terms. SAR and TAR add learning, market barriers, and policy conditions. AR4 and AR5 provide more systematic attention to technology portfolios, costs, and path dependency. AR6 adds a more explicitly system-level representation: renewable integration, storage, hydrogen, CCS, critical minerals, supply chains, technology transfer, and institutional coordination appear as interacting constraints rather than isolated barriers.
This pattern supports a methodological conclusion. When the same prompt is held constant, later reports do not merely list more technologies. They represent technologies as embedded in systems. That is a relevant result for scenario users because a technology pathway cannot be interpreted without its enabling conditions.
The early reports already contain the essential caution that technological options do not implement themselves. FAR discusses energy efficiency, renewable energy, nuclear energy, and other mitigation options in broad terms, but it does not represent technology as an autonomous solution. Cost, infrastructure, and social or institutional conditions are part of the feasibility problem from the beginning. This finding is important because it counters a common reading in which early climate assessment is imagined as technologically naive. The early reports are less granular, but they are not simply techno-optimistic.
SAR and TAR add more explicit attention to policy and market conditions. Technological change is increasingly connected to learning, diffusion, institutional arrangements, and economic incentives. The representation remains less detailed than in later assessments, but the problem is already socio-technical. Technologies are not only devices; they are embedded in systems of investment, regulation, adoption, and infrastructure. This broadening anticipates later work on transition pathways, even though the later vocabulary of energy-system transformation is not yet fully present.
AR4 and AR5 give the technology pillar a more systematic form. They discuss mitigation portfolios, cost ranges, deployment constraints, and the consequences of locking in emissions-intensive infrastructure. AR5 in particular represents technology in relation to pathway feasibility. The question is not whether a technology exists, but under which conditions it can contribute at sufficient scale and speed. This shift is central to scenario knowledge representation because it connects technical potential to temporal feasibility. A technology that arrives too late, scales too slowly, or depends on unrealistic enabling conditions cannot play the same role in all pathways.
AR6 extends this logic to a more complex system representation. Renewable energy, storage, electrification, hydrogen, carbon capture, critical minerals, demand-side measures, finance, innovation systems, and technology transfer appear as interdependent elements. The representational change is not only that more technologies are listed. It is that technology is increasingly framed as a system transformation problem. This matters for policy because the bottleneck may not be laboratory performance. It may be grid integration, land-use conflict, mineral supply, permitting, workforce capacity, social legitimacy, or policy credibility.
Across P2, the high convergence is therefore substantive. Every report treats technology as important but conditional. The longitudinal difference is in the density of boundary conditions. Early assessments identify cost and general feasibility constraints. Middle assessments add market and policy barriers. Later assessments embed technologies in socio-technical systems and pathway feasibility. This is one of the clearest cases in which the IPCC’s core judgment is stable while its representational apparatus becomes more useful for decision-making.
The P2 findings also illustrate the value of the research-agenda prompt. When earlier outputs identify missing data on deployment costs, regional conditions, learning rates, infrastructure, or policy instruments, they point to precisely the kinds of evidence that later reports increasingly incorporate. The model-assisted procedure therefore does not only summarize content; it helps identify how the assessment literature’s evidentiary needs changed over time. This should be useful for future assessment design, especially when emerging technologies are discussed before their scale-up conditions are well documented.
P3: Compound socio-economic risks and governance stress
Prompt P3 asked whether the reports identify compound risks, second-order effects, and governance stress. The cross-report pattern is partial convergence. All reports recognize that climate risks interact with human systems, but the structure of this representation changes substantially. Early reports describe sectoral impacts such as health, water, agriculture, and coastal risk. SAR is notable because it already identifies insurance and financial vulnerability, even though it does not connect these effects into a fully developed cascading-risk framework.
TAR and AR4 broaden the representation through vulnerability, capacity, and reasons-for-concern language. AR5 introduces more explicit treatment of adaptation limits, emergent risks, conflict and migration with careful uncertainty language. AR6 provides the most integrated account by linking hazards, exposure, vulnerability, inequality, infrastructure, governance, and sequential or compound events. In AR6, climate risk is represented less as a set of sectoral damages and more as a systemic interaction between hazards and social organization.
The key finding is that governance becomes increasingly endogenous to scenario knowledge. In the earlier assessments, governance is largely a response context. In later assessments, governance quality, institutional capacity, and policy coordination become determinants of risk outcomes. This representational shift is central to interpreting the evolution of IPCC scenario knowledge.
FAR represents socio-economic risk primarily through sectoral impact categories. Health, water, agriculture, ecosystems, and coastal zones appear as domains affected by climate change. This organization is understandable for an early assessment whose central task was to establish the climate problem and its broad consequences. Yet it means that compound risk is not fully represented as a cascade. The report identifies multiple affected sectors but provides limited treatment of the mechanisms through which impacts in one sector might propagate into another or stress institutions.
SAR shows that the seeds of compound-risk thinking were present early. It recognizes climate change as an additional stress on already stressed systems and includes concern for health, water, food, and economic vulnerability. Its discussion of insurance and disaster losses is especially relevant because it anticipates a later concern with financial risk and institutional resilience. At the same time, SAR does not yet integrate these elements into a formal governance-stress framework. The representation is broad but still relatively enumerative: risks are listed and described, but their second-order interactions remain underdeveloped.
TAR and AR4 progressively expand the role of vulnerability and adaptive capacity. This is a major representational shift. Risk is no longer merely the product of a physical hazard; it depends on exposure, sensitivity, capacity, development conditions, and institutional context. AR4’s reasons-for-concern framing further helps connect warming levels to categories of risk and therefore provides a more structured basis for comparing future conditions. Nevertheless, the synthesis-level representation still often separates sectors more than later reports do.
AR5 introduces stronger attention to adaptation limits, emergent risks, and the interaction between climate change and socio-economic stressors. Conflict and migration are treated with caution, which is methodologically important. The report does not simply assert deterministic causal links between climate and social instability. It communicates the role of mediating conditions and uncertainty. This is a more sophisticated representation because it recognizes both the possibility of second-order effects and the evidentiary limits of attributing them to climate change alone.
AR6 gives the most integrated response to P3. It represents climate risk as the interaction of hazards, exposure, vulnerability, inequality, infrastructure, ecosystems, finance, and governance. Compound and sequential events become more prominent, as do limits to adaptation and maladaptation. Governance is not merely the arena in which responses occur; it is part of the causal structure of risk. Weak institutions, fragmented planning, inadequate finance, and inequitable exposure can intensify impacts. Conversely, inclusive governance and effective adaptation can reduce risk. This is a substantial representational change from sectoral impact lists toward systemic risk analysis.
The governance finding is particularly important for scenario use. If governance is treated only as an implementation variable, scenarios may underestimate the ways institutional capacity shapes the future itself. If governance is treated as endogenous to risk, then scenario users must consider policy credibility, coordination capacity, fiscal space, legitimacy, and distributional vulnerability as part of climate futures. AR6 moves most strongly in this direction. It thereby makes climate scenarios more relevant to decision-makers, but also more complex to interpret.
P3 therefore shows partial convergence and strong representational expansion. All reports recognize that climate change affects human systems. Later reports increasingly explain how those effects interact, cascade, and become mediated by institutions. The movement from impacts to systemic risk is one of the central historical developments in the synthesis reports. It also supports the paper’s methodological claim: identical prompts can reveal not only what topics appear, but how the structure of explanation changes across assessments.
P4: Carbon dioxide removal feasibility and policy lock-in
Prompt P4 produced the greatest divergence. FAR and SAR offer little explicit discussion of carbon dioxide removal as the term is used today. They discuss emissions reduction, sinks, and afforestation in broad terms, but they do not represent CDR as a central policy pathway. TAR remains focused on emissions reductions and only indirectly touches related issues such as leakage and long-term stabilization. AR4 introduces CCS and land-based mitigation options more clearly, but large-scale removal feasibility remains underdeveloped.
AR5 marks a major representational change. CDR enters the synthesis through mitigation pathways that require net negative emissions in the second half of the century, especially for stringent temperature goals. Yet AR5 also frames this dependence as uncertain because technologies such as BECCS had not been demonstrated at large scale and because land, water, sustainability, and governance constraints remained unresolved.
AR6 treats CDR as unavoidable for balancing residual emissions and for achieving net-zero CO2, while emphasizing that CDR cannot substitute for rapid emissions reductions. It also makes the policy lock-in problem explicit: if policymakers rely on future removals to justify delayed mitigation, then pathways become vulnerable to failure if CDR cannot scale sustainably. In P4, the RAG comparison makes visible how an issue moves from marginal or implicit to central and highly conditioned. This is the clearest example of why absence in an earlier assessment should not be interpreted as evidence that a later concern is unimportant. It may simply mean that the scientific and scenario infrastructure needed to represent the concern had not yet matured.
The early reports are best interpreted as pre-CDR assessments rather than as assessments that rejected CDR. FAR and SAR discuss emissions reductions, sinks, and land-use options, but they do not represent large-scale carbon dioxide removal as a central component of mitigation pathways. This is not surprising. The policy problem of net zero had not yet crystallized in the form that now dominates mitigation pathways, and the modelling literature on large-scale negative emissions was not yet central to synthesis-level assessment. The appropriate result for P4 in these reports is therefore insufficiency with limited indirect relevance, not substantive disagreement with AR6.
TAR and AR4 occupy a transitional position. They discuss stabilization, mitigation portfolios, carbon sinks, land-based mitigation, and carbon capture and storage more clearly than the earliest reports. Yet the synthesis-level treatment remains closer to emissions reduction and storage than to a contemporary CDR framework. The feasibility question is present only in fragments. Costs, storage, institutional capacity, and leakage can be discussed, but the reports do not yet frame future removals as a pathway condition for meeting stringent temperature goals. The representation is therefore partial and technologically dispersed.
AR5 changes the analytical status of removals. Negative emissions become visible as a model-dependent assumption in many stringent mitigation pathways, especially those aiming to limit warming with delayed near-term action. This is a major representational event. CDR is no longer only a land-use or sink issue. It becomes an intertemporal mitigation device: present emissions trajectories can imply future removal requirements. AR5 also communicates uncertainty about that dependence. Technologies such as BECCS are important in models, but their large-scale feasibility is constrained by land, water, biodiversity, governance, and social acceptability. The report therefore introduces both the promise and the fragility of removal-dependent pathways.
AR6 goes further by connecting CDR to net-zero CO2, residual emissions, overshoot, and mitigation deterrence. The report represents CDR as necessary for balancing residual emissions but not as a substitute for rapid emissions reductions. This distinction is central. It prevents CDR from being treated as an escape hatch that allows high emissions to continue. AR6 also makes the lock-in problem more explicit. If policy choices preserve fossil infrastructure or delay mitigation in expectation of future removals, then the pathway becomes dependent on technologies and governance systems that may not scale sustainably. The risk is not only technical failure. It is institutional and political dependence on a future corrective capacity.
The P4 result is therefore the strongest evidence that scenario knowledge representation can change because a policy problem becomes newly assessable. Earlier reports did not have the same net-zero framing, the same overshoot vocabulary, or the same evidence base for evaluating CDR at scale. Later reports do. This creates a large divergence in output when the same prompt is applied across reports. The divergence should not be interpreted as inconsistency in the IPCC’s core climate science. It should be interpreted as the emergence and consolidation of a new scenario problem.
For policymakers, the implication is immediate. Reliance on future removals must be read as a pathway assumption that carries feasibility and governance risks. A pathway that requires large future CDR is not equivalent to a pathway that achieves rapid near-term emissions reductions and uses removals only for residual emissions. The RAG comparison makes this difference historically visible. It shows how the assessment system moved from sinks as a general mitigation concept to CDR as a constrained, necessary, and potentially risky component of net-zero strategy.
Discussion
Main interpretation: the evolution of scenario knowledge representation
The results show that the main object of comparison is the evolution of scenario knowledge representation rather than scenario content alone. Across P1-P4, the longitudinal change is best understood as a change in representational capacity. Early assessments establish the conditional logic of climate risk. Middle assessments introduce stronger vulnerability, mitigation-cost, and uncertainty structures. Later assessments integrate pathways, socio-economic futures, adaptation limits, net-zero constraints, CDR feasibility, and governance stress into a more coherent framework.
This matters because users often compare IPCC reports as if each report were a new version of the same static document. The comparison shows instead that each report is a historically situated assessment instrument. Its conclusions are shaped by available evidence, available models, accepted scenario architectures, and assessment conventions. A statement in FAR or SAR may remain broadly consistent with AR6 while being less precise, less calibrated, and less integrated. Divergence should therefore be interpreted before it is judged.
The theoretical contribution of this finding is to treat assessment change as a change in knowledge representation rather than as a simple accumulation of facts. Climate assessment does accumulate facts: observations improve, models develop, impacts are better documented, and mitigation options are evaluated with greater empirical detail. Yet the more consequential change for scenario users is that facts are organized through changing conceptual architectures. The reports progressively connect emissions, warming, impacts, vulnerability, feasibility, uncertainty, and policy sequencing. The result is a more integrated grammar of climate futures.
This interpretation also protects the analysis from two symmetrical errors. The first error is presentism: judging early reports by the categories and expectations of AR6. The second error is flattening: treating all reports as if they merely repeat the same message with minor updates. The evidence supports neither position. The early reports contain durable conditional reasoning, but later reports represent that reasoning with far more explicit scenario architecture, uncertainty calibration, and policy relevance. The history is therefore one of continuity with transformation.
The findings also clarify what should be meant by consistency across IPCC reports. Consistency does not require identical wording, identical scenario labels, or identical quantitative thresholds. A more appropriate standard is functional consistency: do reports sustain the same high-level causal logic when asked comparable questions? On that standard, P1 and P2 show substantial consistency, P3 shows conceptual continuity with stronger systemic elaboration, and P4 shows topic emergence rather than straightforward inconsistency. This distinction should be useful for users who cite IPCC reports across time.
Implications for climate scenario use
The first implication is that scenarios should be used as conditional decision tools, not as forecasts. This is especially important when comparing reports across time. A high-emissions scenario in an earlier report is not automatically equivalent to an SSP5-8.5 pathway, and an early qualitative impact statement is not equivalent to an AR6 statement with calibrated confidence. Scenario users should translate across frameworks before drawing policy conclusions.
The second implication concerns technology and CDR. The high convergence in P2 supports the durable IPCC conclusion that technological change requires enabling socio-technical conditions. The low convergence in P4 shows that some policy-relevant topics become assessable only when the modelling and evidence base matures. Users should therefore treat emerging scenario elements as dynamic: their absence from older reports may reflect the stage of assessment science rather than a negative judgment on relevance.
The third implication concerns governance. AR6 increasingly represents governance not as an external implementation detail but as part of the risk system. This should encourage scenario users to track institutional capacity, policy credibility, finance, and distributional vulnerability alongside physical climate indicators.
The analysis also implies that scenario translation is now a necessary skill for policy users. A scenario category from one assessment period should not be mapped mechanically onto a category from another. Translation requires attention to baseline assumptions, socio-economic narratives, forcing levels, technology assumptions, and policy context. Without such translation, an apparent inconsistency may simply be an artefact of framework change. Conversely, a superficial continuity may conceal an important change in feasibility assumptions or uncertainty language.
For research users, the main implication is that longitudinal comparisons should specify the level at which comparison is being made. One can compare physical climate response, impact categories, socio-economic vulnerability, mitigation pathways, uncertainty treatment, or policy relevance. These are not the same object. The present study compares scenario knowledge representation at the synthesis-report level. It therefore complements, rather than replaces, sectoral comparisons, model intercomparison studies, and detailed Working Group analyses.
For policy users, the strongest substantive message concerns delayed action. Across assessment cycles and across pillars, delay tends to increase future constraints. It raises adaptation burdens, narrows technological and institutional flexibility, increases the possibility of residual damages, and can create dependence on future removals. Later reports make this logic more precise, but the basic structure is stable. That stability should increase confidence in the general direction of the policy implication, even though the exact quantitative estimates and scenario labels have changed over time.
Methodological contribution of layered RAG prompting
The methodological contribution is a structured use of RAG for longitudinal assessment comparison. The design is not a free-form conversation with an LLM. It is a controlled comparison in which the same scenario prompt is applied to each report, relevant passages are retrieved from that report only, and the output is produced through analytical layers that correspond to explicit research tasks. This makes the method closer to assisted content analysis than to automated interpretation.
The method can support hypothesis generation and audit. For example, one can test whether uncertainty language becomes more calibrated after AR4, whether CDR becomes central only after AR5, or whether governance appears increasingly as a determinant of risk. The method does not prove causal explanations for these changes, but it makes the evidence traceable and comparison-ready. It also provides a practical tool for future assessment authors who need to track continuity and change across prior reports.
The method is particularly useful where comparison requires semantic sensitivity. Keyword searches can show when a term appears, but they cannot reliably determine whether an earlier report expressed a related concept in different words. Manual expert reading can do this, but it is time-consuming and difficult to reproduce across long corpora. A report-specific RAG pipeline occupies an intermediate position. It can surface relevant passages, organize them under stable analytical categories, and produce comparison-ready outputs while leaving final interpretation to human researchers.
The approach is also useful for auditing assessment continuity. Future assessment teams could query prior reports on a given topic to identify what has remained stable, what has changed, and what was previously underdeveloped. This does not mean that AI should write assessments. It means that AI-assisted retrieval and prompting can help maintain institutional memory in long assessment cycles. In domains where reports are recurrent, cumulative, and policy-relevant, such an audit function may be valuable.
At the same time, the method should not be overstated. It does not discover truth independent of the corpus. It cannot correct omissions in the synthesis reports. It cannot replace expert judgment about scenario feasibility. It can miss relevant passages, overinterpret retrieved text, or reproduce biases in the model’s language. The methodological contribution is therefore bounded: it provides a disciplined, transparent, and scalable aid to comparison, not an autonomous assessment system.
Indicators for future assessments
The analysis suggests several indicators that could be used to monitor the maturity of scenario knowledge representation in future assessment cycles. One indicator is the share of key scenario claims accompanied by quantitative signposts. A second is the share accompanied by calibrated uncertainty language. A third is the degree of cross-sectoral integration, measured by whether causal chains connect physical hazards, exposure, vulnerability, institutions, finance, and policy response. A fourth is topic emergence, measured by first appearances and subsequent integration of concepts such as net zero, CDR, overshoot, adaptation limits, and compound risks.
These indicators should not replace expert judgment. Their value is diagnostic. They can help authors, reviewers, and users detect whether important scenario claims remain qualitative, whether uncertainty is unevenly represented, and whether emerging risks have been integrated into the assessment structure.
These indicators can be operationalized in several ways. Evidence specificity could be measured by the proportion of scenario claims linked to numerical thresholds, ranges, or model-based quantities. Uncertainty articulation could be measured by the proportion of key claims accompanied by calibrated confidence or likelihood language. Causal integration could be measured by the number of distinct system components connected within a claim, for example hazards, exposure, vulnerability, institutions, finance, infrastructure, and response options. Scenario-framework alignment could be measured by the degree to which claims are explicitly linked to recognized pathway families, mitigation categories, or socio-economic assumptions.
Such indicators should be interpreted as meta-assessment tools. They do not evaluate whether a claim is true. They evaluate how a claim is represented for users. A highly important claim may remain qualitative because the evidence does not support quantification. Conversely, a quantified claim may still rest on uncertain assumptions. The value of the indicators lies in revealing unevenness: some topics may receive strong quantification and calibrated language, while others remain narrative, fragmented, or weakly integrated. That unevenness is itself useful information for assessment design.
The topic-emergence indicator is especially relevant for AR7 and subsequent assessments. Concepts such as CDR, overshoot, maladaptation, adaptation limits, just transition, climate-related financial risk, and systemic resilience did not all enter the assessment vocabulary at the same time. A longitudinal text-audit can help identify when a concept first appears, when it becomes integrated into scenario logic, and when it becomes connected to policy decisions. This would help distinguish early-warning concepts from mature assessment categories.
Limitations and validity safeguards
Several limitations must be acknowledged. First, the analysis depends on the quality of retrieval. If relevant passages were not retrieved, a report may appear silent when the topic was addressed elsewhere. This limitation is partly mitigated by fixed top-k retrieval, source checking of unexpected findings, and manual inspection of consequential absence claims, but it cannot be eliminated without full-document manual coding.
Second, the LLM outputs are interpretations of retrieved passages, not primary evidence. The model may infer a causal relation or confidence level more strongly than the source text warrants. For that reason, the study treats outputs as analytical aids and bases conclusions on patterns verified against the source reports.
Third, the corpus is limited to synthesis reports. Working Group chapters contain more detailed evidence, and some findings that appear absent from a synthesis report may be present in the underlying Working Group volumes. This choice improves cross-cycle comparability but reduces topical coverage.
Fourth, the scenario baselines change across time. The study compares responses to identical prompts, but the underlying scenario architecture differs by assessment cycle. This is not a flaw in itself; it is part of what the study investigates. It does mean, however, that direct numerical equivalence across reports should not be assumed.
Fifth, the study is descriptive and interpretive. It does not perform inferential statistics and should not claim statistical representativeness beyond the corpus analyzed. Its rigor comes from transparent retrieval, fixed prompts, coding rules, and source verification.
Sixth, the method is sensitive to prompt design. A different set of prompts might foreground equity, biodiversity, regional adaptation, finance, or loss and damage. The four pillars used here are theoretically sufficient for the stated research questions, but they are not comprehensive. They should be understood as a structured sample of scenario problems rather than a complete map of climate assessment knowledge.
Seventh, the analysis is sensitive to the distinction between terminology and concept. The sparse retrieval design makes lexical overlap auditable, but it cannot always determine whether two differently named concepts are genuinely comparable. For example, early discussions of sinks are relevant to CDR, but they are not equivalent to contemporary CDR pathways. Human interpretation is required to avoid false equivalence, and future work could test dense-retrieval ablations as a sensitivity extension.
Eighth, the method cannot determine why a report changed. It can show that later reports contain more calibrated uncertainty, more integrated causal mechanisms, or more detailed CDR discussion. It cannot prove whether those changes were caused by scientific advances, policy demand, modelling innovation, author decisions, or the evolving mandate of the IPCC. The discussion offers plausible interpretations grounded in the history of scenario frameworks, but causal attribution remains limited.
Ninth, full auditability depends on reproducibility materials. Readers require access not only to the source reports but also to the exact prompts, retrieved passages, model outputs, coding matrix, retrieval parameters, and model configuration. Without those materials, the study remains transparent in design but only partially reproducible in execution. S1 Text therefore separates implemented parameters from missing execution metadata and proposed defaults for controlled reruns.
Conclusion
This paper shows that the six IPCC synthesis reports exhibit both continuity and transformation in their answers to identical scenario-based prompts. The continuity lies in high-level causal logic: delayed mitigation increases future risk; technologies require enabling conditions; climate impacts interact with socio-economic systems; and long-term stabilization depends on sustained policy action. The transformation lies in scenario knowledge representation. Across assessment cycles, IPCC synthesis reports become more quantitative, more uncertainty-calibrated, more pathway-sensitive, and more explicit about feasibility, governance, adaptation limits, and CDR-related lock-in.
The revised contribution is therefore twofold. Substantively, the paper provides a longitudinal interpretation of how IPCC scenario knowledge matures across assessment cycles. Methodologically, it demonstrates a replicable RAG-based approach for auditing large assessment corpora under stable prompts. The method is not a replacement for expert review. It is a disciplined aid for making comparison more transparent and for identifying where evidence, uncertainty, and causal mechanisms have become more developed over time.
Future work should extend the corpus to Working Group chapters, test robustness across model families and retrieval settings, publish complete retrieval logs and coding matrices, and invite assessment experts to evaluate the validity of model-assisted interpretations. Such extensions would strengthen the use of AI-assisted methods in climate assessment while preserving the central role of expert judgment.
The main empirical lesson is that longitudinal comparison requires translation rather than simple alignment. The reports share a durable causal architecture, but the language and evidence through which that architecture is represented have changed substantially. A reader who compares FAR and AR6 must therefore ask not only whether the two reports agree, but also whether they are capable of answering the same question at the same level of specificity. In many cases, early reports answer the question qualitatively, middle reports begin to quantify and systematize it, and later reports integrate it into a pathway-sensitive and uncertainty-calibrated framework.
The methodological lesson is that AI-assisted analysis is most defensible when it is narrow, grounded, and auditable. RAG does not make the comparison objective by itself. It makes the comparison more structured by holding prompts constant, retrieving report-specific evidence, and producing intermediate outputs that can be checked. The human analyst remains responsible for interpreting what the outputs mean, correcting overstatements, and situating differences in the history of assessment practice.
The practical lesson is that future scenario use should be historically aware. Older IPCC reports remain valuable because they show the persistence of central climate-risk logics. Later reports are indispensable because they contain the current representational architecture through which those logics are connected to carbon budgets, adaptation limits, net-zero pathways, feasibility, governance, and CDR. The most responsible use of the assessment record is therefore cumulative but not flattening: each report should be read as part of a sequence, and the sequence should be translated across changing scenario frameworks.
Supporting information
S1 Text. Reproducibility, prompt, and parameter audit.
This supporting information file contains the source-corpus inventory, retrieval formulas, prompt templates, executed first- and second-pass prompt banks, parameter audit, missing execution metadata, proposed defaults for future controlled reruns, output inventory, and manual verification rules.
https://doi.org/10.1371/journal.pclm.0000965.s001
(DOCX)
S1 Table. First-pass model outputs generated from the ten scenario prompts across the six IPCC synthesis-report files (answers.csv).
https://doi.org/10.1371/journal.pclm.0000965.s002
(CSV)
S1 Code. Redacted R scripts used for first-pass retrieval/generation, with the private local-area-network endpoint removed and replaced by a local Ollama-compatible endpoint configuration.
https://doi.org/10.1371/journal.pclm.0000965.s003
(R)
S2 Table. Second-pass comparative outputs generated from the ten analytical prompts applied to the first-pass outputs (comparative-analyses.csv).
https://doi.org/10.1371/journal.pclm.0000965.s004
(CSV)
S2 Code. Redacted R scripts used for second-pass comparative analysis, with the private local-area-network endpoint removed and replaced by a local Ollama-compatible endpoint configuration.
https://doi.org/10.1371/journal.pclm.0000965.s005
(R)
References
- 1.
Intergovernmental Panel on Climate Change. Climate Change 2023: Synthesis Report. Contribution of Working Groups I, II and III to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change. Geneva: IPCC; 2023.
- 2. Ripple WJ, Wolf C, Mann ME, Rockström J, Gregg JW, Xu C, et al. The 2025 state of the climate report: a planet on the brink. BioScience. 2025;75(12):1016–27.
- 3.
Pörtner H-O, Roberts DC, Tignor M, Poloczanska ES, Mintenbeck K, Alegría A. Climate Change 2022: Impacts, Adaptation and Vulnerability. Contribution of Working Group II to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change. Cambridge and New York: Cambridge University Press; 2022.
- 4.
Nakicenovic N, Swart R. Special report on emissions scenarios. Cambridge: Cambridge University Press; 2000.
- 5. Moss RH, Edmonds JA, Hibbard KA, Manning MR, Rose SK, van Vuuren DP, et al. The next generation of scenarios for climate change research and assessment. Nature. 2010;463(7282):747–56. pmid:20148028
- 6. van Vuuren DP, Edmonds J, Kainuma M, Riahi K, Thomson A, Hibbard K, et al. The representative concentration pathways: an overview. Clim Change. 2011;109(1–2):5–31.
- 7. O’Neill BC, Kriegler E, Ebi KL, Kemp-Benedict E, Riahi K, Rothman DS, et al. The roads ahead: Narratives for shared socioeconomic pathways describing world futures in the 21st century. Glob Environ Change. 2017;42:169–80.
- 8. Riahi K, van Vuuren DP, Kriegler E, Edmonds J, O’Neill BC, Fujimori S. The shared socioeconomic pathways and their energy, land use, and greenhouse gas emissions implications: an overview. Glob Environ Change. 2017;42:153–68.
- 9. Lewis P, Perez E, Piktus A, Petroni F, Karpukhin V, Goyal N, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. Adv Neural Inf Process Syst. 2020;33:9459–74.
- 10.
Shuster K, Poff S, Chen M, Kiela D, Weston J. Retrieval augmentation reduces hallucination in conversation. Findings of the Association for Computational Linguistics: EMNLP 2021. 2021. p. 3784–803.
- 11. Fløttum K, Gasper D, St Clair AL. Synthesizing a policy-relevant perspective from the three IPCC worlds: a comparison of topics and frames in the SPMs of the Fifth Assessment Report. Glob Environ Change. 2016;38:118–29.
- 12. Jaczewski A, Brzoska B, Wibig J. Comparison of temperature indices for three IPCC SRES scenarios based on RegCM simulations for Poland in 2011–2030 period. Meteorol Z. 2015;24(1):99–106.
- 13. Luo X, Xia T, Huang J, Xiong D, Ridoutt B. Radiative forcing climate footprints in the agricultural sector: Comparison of models from the IPCC 5th and 6th Assessment Reports. Farm Syst. 2023;1(3):100057.
- 14. Arksey H, O’Malley L. Scoping studies: towards a methodological framework. Int J Soc Res Methodol. 2005;8(1):19–32.
- 15. Levac D, Colquhoun H, O’Brien KK. Scoping studies: advancing the methodology. Implement Sci. 2010;5:69. pmid:20854677
- 16. Munn Z, Peters MDJ, Stern C, Tufanaru C, McArthur A, Aromataris E. Systematic review or scoping review? Guidance for authors when choosing between a systematic or scoping review approach. BMC Med Res Methodol. 2018;18(1):143. pmid:30453902
- 17. Tricco AC, Lillie E, Zarin W, O’Brien KK, Colquhoun H, Levac D, et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Ann Intern Med. 2018;169(7):467–73. pmid:30178033
- 18. Bradfield R, Wright G, Burt G, Cairns G, van der Heijden K. The origins and evolution of scenario techniques in long range business planning. Futures. 2005;37(8):795–812.
- 19. Amer M, Daim TU, Jetter A. A review of scenario planning. Futures. 2013;46:23–40.
- 20.
Mastrandrea MD, Field CB, Stocker TF, Edenhofer O, Ebi KL, Frame DJ, et al. Guidance note for lead authors of the IPCC Fifth Assessment Report on consistent treatment of uncertainties. Geneva: Intergovernmental Panel on Climate Change; 2010.
- 21. Mastrandrea MD, Mach KJ, Plattner G-K, Edenhofer O, Stocker TF, Field CB, et al. The IPCC AR5 guidance note on consistent treatment of uncertainties: a common approach across the working groups. Clim Change. 2011;108(4):675–91.
- 22. Yohe G, Oppenheimer M. Evaluation, characterization, and communication of uncertainty by the intergovernmental panel on climate change—an introductory essay. Clim Change. 2011;108(4):629–39.
- 23.
Salton G, McGill MJ. Introduction to Modern Information Retrieval. New York: McGraw-Hill; 1983.
- 24. Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, Xia F, et al. Chain-of-thought prompting elicits reasoning in large language models. Adv Neural Inf Process Syst. 2022;35:24824–37.
- 25. Spaniol MJ, Rowland NJ. AI-assisted scenario generation for strategic planning. Futures Foresight Sci. 2023;5(2):e148.
- 26. Ellermann K, Seidenberg T, Asmar L, Knepler J, Dumitrescu R. Leveraging GenAI for technology foresight. Proc Des Soc. 2025;5:2221–30.
- 27. Calleo Y, Pilla F, Di Zio S. Generative pre-trained transformers for climate scenarios: a statistical coefficient for future policy development. Qual Quant. 2025;60(2):3895–921.
- 28. Shahbazi Z, Behnamian S. Using large language models to detect and debunk climate change misinformation. Big Data Cogn Comput. 2026;10(1):34.
- 29.
Intergovernmental Panel on Climate Change. Climate Change: The IPCC 1990 and 1992 Assessments. Geneva: World Meteorological Organization and United Nations Environment Programme; 1992.
- 30.
Intergovernmental Panel on Climate Change. IPCC Second Assessment: Climate Change 1995. Geneva: IPCC; 1995.
- 31.
Intergovernmental Panel on Climate Change. Climate Change 2001: Synthesis Report. Watson RT, Core Writing Team, editors. Cambridge: Cambridge University Press; 2001.
- 32.
Intergovernmental Panel on Climate Change. Climate Change 2007: Synthesis Report. Core Writing Team, Pachauri RK, Reisinger A, editors. Geneva: IPCC; 2007.
- 33.
Intergovernmental Panel on Climate Change. Climate Change 2014: Synthesis Report. Core Writing Team, Pachauri RK, Meyer LA, editors. Geneva: IPCC; 2014.