Figures
Abstract
Public support for climate action hinges on the integrity of information environments. In an online experiment, U.S. adults used ChatGPT to evaluate climate-related claims. We first conducted computational text analysis of GPT conversation logs, assessing valence, formality, bias, recency, authority cues, and semantic richness of ChatGPT’s responses using language-model based classifiers and link-level metadata. Concurrently, we measured credibility judgments for both the claims and ChatGPT’s responses through participant self-reports. Credibility perceptions were driven primarily by individual differences; greater acceptance of the scientific consensus, attention to climate change, prior familiarity with ChatGPT, and younger age predicted higher perceived credibility. Once these factors were accounted for, authority cues were associated with slightly less extreme climate attitudes, while positive valence and semantic richness correlated positively with perceived credibility. These results support transparent, unbiased, audience-tailored engagement as a pathway to strengthen responsible climate communication and help bridge the gap between scientific consensus and public understanding.
Citation: Tsang SJ, Wang D (2026) Countering climate misinformation with large language models: Evidence from ChatGPT. PLOS Clim 5(9): e0000930. https://doi.org/10.1371/journal.pclm.0000930
Editor: Oscar Brousse, University College London The Bartlett Faculty of the Built Environment, UNITED KINGDOM OF GREAT BRITAIN AND NORTHERN IRELAND
Received: January 3, 2026; Accepted: July 21, 2026; Published: September 2, 2026
Copyright: © 2026 Tsang, Wang. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All relevant data are within the paper and its Supporting Information files.
Funding: This work is partially supported by the Initiation Grant for Faculty Niche Research Areas, Hong Kong Baptist University (RC-FNRA-IG/21-22/ARTS/01 to TS). The funder had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. The grant funded WD to assist with this project.
Competing interests: The authors have declared that no competing interests exist.
Main
Public support for climate mitigation and adaptation depends not only on evidence and policy, but also on the integrity of the information environment that shapes citizens’ beliefs. Climate misinformation, false or misleading content shared without intent to deceive, and climate disinformation, deliberately deceptive content, circulate widely through social platforms and messaging applications [1]. These narratives claim that global warming is natural, prediction models are unreliable, scientists lack consensus, or that solutions are too costly, ultimately eroding public trust, deepening polarization, and undermining support for effective action [2]. In response, while traditional fact-checking remains valuable, it cannot compete with the speed, scale, and personalization of digital platforms [3]. Rapid advances in large language models (LLMs) now offer new avenues for real-time verification.
Importantly, AI-generated corrections have the potential to act as persuasive messages, influencing user receptivity and encouraging informed opinion change. Artificial intelligence (AI) chatbots, such as ChatGPT, can assess the authenticity of climate-related claims in real time, delivering timely, salient, and scalable corrections at the moment of exposure. This capability may increase user engagement, as immediate feedback is often more effective at countering misinformation than delayed responses, and allows for personalization and contextual relevance. However, the deployment of AI for climate information also carries significant risks. Chatbots can sometimes generate hallucinated evidence, introduce subtle biases, perform unevenly across different topics, or display unwarranted overconfidence, all of which may undermine user trust and credibility [4]. Whether these systems consistently correct falsehoods and strengthen climate-supportive attitudes remains an open question with significant implications for sustainability.
Given ChatGPT’s widespread use, we focus on its responses to identify key heuristics such as fairness, accountability, transparency, and explainability, guided by the Elaboration Likelihood Model (ELM) [5]. We examine which response features most effectively reduce belief in misinformation, improve claim accuracy judgments, increase recognition of falsehoods, bolster confidence in AI-assisted verification, and strengthen climate-supportive attitudes. These findings inform chatbot design, enhance communication about uncertainty and sources, and support the responsible deployment of AI in climate communication.
Climate misinformation
Climate communication is complex because climate-related claims span multiple domains, from physical science (e.g., attribution of extreme events and projections of sea level rise; [6]), to economics (e.g., the effects of carbon pricing and discount rates; [7,8]), to technology and infrastructure (e.g., the reliability of renewable energy, nuclear power, and energy storage; [9]), and to social dimensions (e.g., equity, employment transitions, and climate justice; [10]). The accuracy of these claims spans a spectrum from well supported to misleading to false [8,11]. Much misinformation hinges on framing rather than outright falsehoods [12], using tactics such as cherry-picking short time series, conflating weather with climate, equating uncertainty with ignorance, emphasizing costs while omitting benefits, and presenting fringe views as equivalent to established consensus.
Research questions
Credibility refers to an individual’s judgment about the trustworthiness, expertise, and accuracy of a source or a message [13,14]. Prior research distinguishes message credibility, meaning whether content appears accurate and well supported, from source credibility, meaning whether the provider seems knowledgeable, honest, and reliable [12,15]. Both shape whether people accept information and whether it influences beliefs and behaviors. This study assesses perceived credibility in two places: the false climate claim (whether participants find it believable despite its inaccuracy) and the LLM-generated response (whether its explanation, evidence, and guidance seem trustworthy and authoritative). By analyzing which response attributes drive credibility judgments, we aim to identify when AI-mediated verification mitigates misinformation and when it may inadvertently reinforce it.
Research question 1: How do variations in LLM-generated response features influence participants’ perceptions of the credibility of the false climate claim and the credibility of the response itself?
Beyond credibility judgments, the outcome that ultimately matters for climate action, such as pro-environmental behavior and policy support, is attitude. In this context, attitude refers to a relatively stable evaluation of climate change that integrates beliefs, feelings, and intentions [16]. Key elements include whether people believe climate change is happening, whether they view human activities as a primary cause, how serious they judge the risks to be, and whether they support mitigation and adaptation. Attitudes also encompass willingness to endorse policies and are linked to the adoption of personal behaviors and support for collective responses [17]. By linking features of LLM-generated responses to changes in climate attitudes, this study aims to clarify how conversational AI can be designed to support accurate understanding and constructive engagement with climate issues.
Research question 2: How do variations in LLM-generated response features influence participants’ climate-supportive attitudes?
LLMs and conversational AI
AI comprises computational methods that enable machines to perform tasks associated with human intelligence, including perception, reasoning, and language use [18]. LLMs learn statistical patterns from vast corpora to understand and generate text [19,20]. A prominent example is ChatGPT, a conversational system built on LLMs that can answer questions, summarize documents, write code, and support research and creative work [21]. However, these systems have limitations, including hallucinations, biases inherited from training data, and sensitivity to prompt phrasing [22]. Prior work has examined retrieval‑augmented generation, domain‑specific fine‑tuning, and safety guardrails to address these challenges [23]. Building on this, we analyze LLM responses to climate prompts and translate these findings into design recommendations for AI‑assisted verification tools that enhance credibility, usability, and climate‑supportive outcomes.
Model of AI Credibility Evaluation
In the context of climate misinformation and AI-assisted verification, we combined objective computational analyses of ChatGPT’s responses, capturing features such as valence, formality, bias, recency, authority cues, and semantic richness, with subjective participant ratings from survey instruments. Predicting credibility judgments based on LLM-generated responses is appropriate, given the central role of language models like ChatGPT in shaping public information environments. This approach allows us to evaluate both the textual characteristics of LLM-generated climate messaging and how audiences perceive their credibility.
This approach aligns with Shin’s FATE framework for establishing the credibility of AI-generated responses [24], identifying fairness, accountability, transparency, and explainability as essential heuristics in line with the Elaboration Likelihood Model (ELM) [5]. Fairness ensures unbiased and equitable treatment; accountability enables verification through authoritative sources; transparency openly communicates relevant details, such as recency; and explainability clarifies the AI’s reasoning. Shin argued that these dimensions act as heuristic cues to foster user trust [24]. In this study, we examine how these elements influence the perceived credibility of AI-generated information by mapping features like valence, bias, authority, recency, formality, and semantic richness to each dimension (see Table 1). This mapping illustrates how specific AI response features support user trust and perceived credibility, guiding the ethical and effective design of human-AI interactions.
Valence. Valence reflects the evaluative tendency of a message and is central to persuasion and risk communication. Negative language tends to heighten attention and perceived severity by foregrounding harm, risk, loss, costs, constraints, or failure. Positive language, by contrast, enhances efficacy and support by highlighting benefits, opportunities, solutions, gains, resilience, and success. Research consistently finds that negative framing often attract more attention [25] and is more diagnostic for risk judgments [26,27], whereas positive framing increases perceived feasibility and willingness to support solutions [28,29]. This perspective aligns with the practical aim of distinguishing responses emphasizing risk versus those emphasizing solutions.
Bias. Bias refers to one-sided or prejudicial framing and is distinct from factual accuracy or sentiment. Biased language can manifest as loaded or partisan terms, selective emphasis, pejoratives, stereotyping, or normative claims presented as objective facts. Such framing often activates prior identities and heuristics, heightening motivated reasoning and polarization [30,31]. Detecting bias at the level of phrasing and emphasis allows us to examine whether apparent partiality, rather than the truthfulness of the content per se, diminishes credibility or dampens climate-supportive attitudes.
Authority. Authority denotes the prominence of a source in the broader information ecosystem. Structural signals such as domain authority and link-based metrics correlate with visibility and perceived credibility [32], reflecting how often a site is cited and its standing within the online information network. High authority suggests a better-linked, more established source and often aligns with institutional trust, though it is not a guarantee of accuracy or neutrality [33]. Treating authority as a proxy, rather than a definitive measure, emphasizes the need to complement it with qualitative assessments of source quality when interpreting results or making policy recommendations.
Recency. Recency denotes the timeliness of cited information. In rapidly evolving domains such as climate science and policy, currency functions as a heuristic for relevance and reliability, with up-to-date sources generally perceived as more credible [34]. Timeliness does not guarantee accuracy, evergreen materials may remain valid for years, but outdated sources can undermine trust, particularly when audiences expect current data or recent consensus statements. Measuring recency by publication date helps assess whether responses anchor claims in contemporary evidence, which is likely to matter for both credibility judgments and attitude change.
Formality. Formality captures the stylistic register of a response and is closely linked to perceived professionalism, objectivity, and competence. Formal text typically uses an impersonal tone, precise and technical vocabulary, complete and well-structured sentences, and cautious or hedged claims, while minimizing first- or second-person address. Informal text, in contrast, relies more on conversational markers, contractions, colloquialisms, and direct address. Formal style can bolster perceptions of expertise and trust, especially in scientific and policy contexts [35], though overly technical language may reduce accessibility or warmth. Conversely, informal style can enhance engagement and relatability but may appear less authoritative [36]. Measuring formality helps clarify whether stylistic choices contribute to credibility and acceptance independent of content.
Semantic richness. Semantic richness represents the diversity and unpredictability of word choice, roughly, how information-dense a response is [37]. Lexical diversity and entropy have been used as proxies for elaboration, specificity, and topical depth [38]. Higher lexical entropy typically indicates richer vocabulary, greater conceptual breadth, and inclusion of named entities or quantitative details. Such richness can increase perceptions of expertise and informativeness but may reduce processing fluency if language becomes overly complex or jargon-laden. Conversely, lower entropy often signals generic or boilerplate text that feels vague or noncommittal. Quantifying semantic richness thus helps determine whether more detailed language is associated with higher credibility and more robust attitude change.
Methods
Ethics statement
The project was reviewed and approved by the Research Ethics Committee of Hong Kong Baptist University (REC/24–25/0265). All participants provided written informed consent prior to their participation in the study.
Sample and procedures
A total of 810 adult participants were recruited between April 3 and 7, 2025 through a private survey company PureSpectrum, and written informed consent was obtained prior to survey initiation. Participants were informed that they would be asked to “engage with ChatGPT,” described as “an advanced computer program developed by OpenAI, capable of engaging in conversations, answering questions, and providing information in a manner similar to that of a human being.” They were then instructed to engage with ChatGPT in a minimum of three and up to ten conversational turns (i.e., ten user inputs). Beyond these instructions, we did not provide detailed guidance on how to formulate prompts (e.g., specific wording or level of detail) in order to approximate naturalistic use and allow participants to interact with the system in ways that reflected their own information‑seeking practices. All conversations in which participants engaged with ChatGPT via Qualtrics were logged and subsequently analyzed to compute various linguistic and content-based metrics, as documented at https://github.com/Fact-checkingclub/human-LLM-interaction-in-climate-misinformation-correction.
They were asked to review a claim related to climate change, after which they were asked to check and verify its accuracy, determining whether the claim is true or false using ChatGPT. Each participant was randomly assigned one of four false claims: “Cold Snaps Disprove Climate Change,” “Stable Ice Caps Disprove Climate Change,” “Natural Causes of Climate Change,” or “Human Actions Not Causing Climate Change.” Each claim included a brief 60-word description elaborating on the misinformation. These claims were selected for their widespread circulation online and their key partisan divides in climate change discourse. Preliminary analyses indicated that there were no significant differences among the four misinformation claim conditions on our main variables (all p values > .05). Therefore, data from all conditions were combined for subsequent analyses.
Upon engaging with ChatGPT, participants were asked to answer questions about their experience, including perceived credibility of misinformation, perceived credibility of the ChatGPT response, attitudes towards climate change, and several demographic items. Twenty participants who did not complete the full engagement with ChatGPT were excluded from the analysis, resulting in a final analytic sample of 790 adults (48.5% male, 50.9% female, 0.6% other). The mean age was 48.76 years (SD = 17.05; range = 18–86). Mean educational attainment fell between “some college, no degree” and “associate degree (2-year)” categories; 19.4% reported some college without a degree and 13.7% reported an associate degree. Overall, 64.1% had completed at least an associate degree, including 21.6% with a bachelor’s degree.
LLM-generated responses
Valence. Sentiment probabilities for positive and negative classes were estimated with the RoBERTa-base model [39,40]. For each claim, class probabilities were averaged across its responses to yield claim-level information valence (positivity: M = 0.31; SD = 0.23; negativity: M = 0.11; SD = 0.09; see Fig 1).
Bias. Bias detection was performed using the DistilBERT binary classifier (biased vs. unbiased) [41]. Bias was quantified as the proportion of responses labeled biased for each claim (M = 0.89; SD = 0.18).
Authority. Domain authority was obtained from Semrush on a logarithmic 1–100 scale [42,43]. Claim-level authority was computed as the mean domain authority across sources associated with the responses (M = 75.80; SD = 15.39).
Recency. To assess source recency [44,45], each response was queried via Google Search to identify the associated website domain and publication date. Sources published between October 1, 2018, and October 1, 2023, were classified as current, reflecting a five-year window aligned with the most recent update of ChatGPT-4o (Mini) prior to the study (October 2023). Claim-level recency was defined as the proportion of sources classified as current (M = 0.95; SD = 0.10).
Formality. The probability that a response used formal language was estimated with the DeBERTa-large model [44,46]. Claim-level formality was computed as the mean predicted probability across responses (M = 0.87; SD = 0.08).
Semantic richness. For each ChatGPT-generated response to the target claim, naïve Shannon entropy was computed as: , where
denotes the number of unique words in the response and
stands for the relative frequency of word
. Higher entropy indicates a larger set of unique words and a more even distribution of word use, reflecting greater lexical richness and complexity [38]. Claim-level semantic richness was defined as the mean lexical entropy across all responses for that claim (M = 6.02; SD = 1.05).
Outcome variables
Perceived credibility of misinformation. After engaging with ChatGPT, participants assessed the credibility of the target claim using two items on five-point scales: “How would you rate the accuracy of the claim you were asked to review?” (1 = Not at all accurate; 5 = Extremely accurate), and “In general, how truthful is that claim?” (1 = Not at all truthful; 5 = Extremely truthful). Items were averaged to create a composite index of perceived claim credibility, with higher scores indicating greater perceived accuracy/truthfulness of the (misleading) claim (M = 3.43; SD = 1.16; Cronbach’s α = .90).
Perceived credibility of ChatGPT response. After viewing the ChatGPT-generated response, participants evaluated its quality and credibility using five items on a five-point scale (1 = Not at all; 5 = Extremely): accurate, correct, logical, reasonable, and credible. Items were averaged to form a composite index of perceived response credibility, with higher scores indicating more favorable evaluations (M = 3.58; SD = 1.08. Cronbach’s α = .96).
Attitude towards climate change. Participants completed three items on a six-point agreement scale (1 = Strongly disagree; 6 = Strongly agree): “There is solid evidence that the average temperature on Earth has been getting warmer,” “Earth is getting warmer mostly because of human activity such as burning fossil fuels,” and “The government should impose stricter environmental laws and regulations.” Items were averaged to create a composite index of pro-consensus, pro-regulation climate attitudes, with higher scores indicating stronger endorsement (M = 4.47; SD = 1.25; Cronbach’s α = .86). The pairwise relationships among the three outcome variables are presented in Fig 2. To avoid a neutral midpoint and encourage respondents to take a clear position, we used a 6-point Likert scale to measure climate attitudes. This forced-choice approach provides a more definitive assessment of participants’ views on climate issues.
Each point was colored according to local density (Gaussian kernel density estimation). Ordinary least squares (OLS) regression lines with 95% confidence intervals were overlaid to illustrate the linear trend.
Control variables
Control variables, including climate attitudes and prior awareness of ChatGPT, were measured pre-treatment to avoid any influence from the intervention. By contrast, attention to climate change was assessed post-treatment, reflecting our view of it as a relatively stable disposition and thus unlikely to be affected by the manipulation.
Stance on climate change. Participants rated agreement with two statements on a five-point scale (1 = Strongly disagree; 5 = Strongly agree): “There is solid evidence that the Earth is getting warmer,” and “The Earth is getting warmer mostly because of human activity such as burning fossil fuels.” Items were averaged to form a composite index of pro-consensus stance on climate change, with higher scores indicating stronger endorsement of anthropogenic warming (α = .83; M = 3.78; SD = 1.02).
Attention to climate change. Participants answered a single item, “To what extent do you think the controversy on climate change is important?”, on a five-point scale (1 = Not at all; 5 = Extremely). Higher scores indicate greater perceived importance (M = 3.32; SD = 1.26).
Prior awareness of ChatGPT. Participants reported how much they had heard about ChatGPT, described as “an artificial intelligence (AI) program used to create text,” on a five-point scale (1 = None at all; 5 = A great deal). Higher scores reflect greater prior awareness (M = 2.93; SD = 1.22).
Results
Perceived credibility of ChatGPT responses
We first conducted a multiple linear regression to predict perceived credibility of ChatGPT’s responses (see Table 2). The model included control variables (age, attention to climate change, personal stance on climate change, and prior awareness of ChatGPT) and ChatGPT-response features (negativity, positivity, bias, authority, recency, formality, and semantic richness). No evidence of multicollinearity was detected, as all Variance Inflation Factor (VIF) values were below 10 and Tolerance values exceeded 0.1.
Age was negatively associated with perceived credibility (β = −.07, p < .05), whereas attention to climate change (β = .19, p < .001), personal stance on climate change (β = .30, p < .001), and prior awareness of ChatGPT (β = .15, p < .001) were positively associated. In other words, younger participants, those more attentive to climate issues, those aligned with the scientific consensus on climate change, and those already familiar with ChatGPT tended to rate its misinformation‑debunking responses as credible.
Among ChatGPT-response features, bias (β = −.09, p < .05) was negatively associated with credibility, while positive valence (β = .16, p < .01) and semantic richness (β = .15, p < .01) were positively associated. Thus, responses that were less biased, more positive in valence, and richer in content received higher credibility ratings.
Perceived credibility of climate misinformation
A parallel multiple regression model examined credibility ratings of climate misinformation, using the same control variables and response-feature predictors. Age again showed a negative association (β = −.14, p < .001), whereas attention to climate change (β = .18, p < .001), stronger acceptance of the scientific consensus on climate change (β = .08, p < .05), and prior awareness of ChatGPT (β = .09, p < .05) were positive predictors.
Among ChatGPT-response features, positive valence (β = .22, p < .001) and semantic richness (β = .15, p < .05) were positively associated with credibility ratings of climate misinformation. Consistent with the prior model, a more positive valence and richer content were linked to higher credibility judgments, even for misinformation.
Attitudes towards climate change
Finally, a third multiple regression model predicted participants’ attitudes toward climate change. Using the same set of controls and response-feature predictors, we found that, as expected, personal stance (β = .67, p < .001) and attention to the issue (β = .24, p < .001) were key predictors. Age was negatively associated with climate attitudes (β = -.05, p < .05).
Among ChatGPT-response features, authority showed a small but significant negative association (β = -.08, p < .01), suggesting that responses referencing more authoritative sources were associated with slightly less extreme pro-climate attitudes.
Discussion
This study investigated how message-level features of LLM-generated climate content shape people’s credibility judgments and climate attitudes at a time when LLMs increasingly mediate information flows. Building on prior research in LLMs [19–22] and climate communication [7,10], our findings demonstrate that attributes of LLM-generated responses substantially influence credibility evaluations. Among the examined features, such as valence, formality, bias, recency, authority, and semantic richness, three key patterns stand out. First, positive valence and greater semantic richness consistently increased perceived credibility. These effects held for perception of both ChatGPT’s response and climate misinformation, suggesting that positivity and semantic richness serve as generalized heuristics of trustworthiness. Second, bias reduced perceived credibility of ChatGPT response, consistent with social norms favoring neutrality and impartiality. Third, authoritative cues were associated with slightly less extreme pro-climate attitudes, implying a modest calibrating effect on opinion intensity. Collectively, these results indicate that readers treat positivity and semantic richness as cues of expertise and effort, while discounting biased content.
Several mechanisms may explain these findings. A lower presence of bias signals fairness and impartiality, reducing defensiveness and fostering trust, which in turn bolsters credibility [47]. However, the role of bias warrants nuance. It reduced perceived credibility of ChatGPT responses but did not significantly predict perceived credibility of climate misinformation. One possibility is that bias may be less noticeable when participants evaluate stand-alone climate claims rather than system-generated responses, or that bias interacts with prior beliefs: when misinformation aligns with an individual’s worldview, bias might be experienced as confirmation rather than as a credibility cost.
A more positive valence conveys efficacy and a solution-oriented outlook, which can enhance perceived usefulness and, consequently, credibility [48]. Yet this stylistic advantage has a double-edged nature. The same positive and semantically rich language that enhances the perceived credibility of AI generated responses can also increase the perceived credibility of climate misinformation, with positive valence exerting an even stronger effect in that context. Upbeat language that emphasizes opportunities, solutions, gains, resilience, and success, coupled with detailed and information-dense content, may lead users to judge both ChatGPT’s debunking and misinformation as credible. Similarly, greater semantic richness, captured by lexical entropy, signals expertise by providing specific details, varied concepts, and concrete evidence, making both ChatGPT responses and misinformation appear thorough and well-informed. This effect warrants both optimism and caution. While richer content can facilitate learning and enhance credibility through elaboration and contextualization, it can also create an illusion of understanding or illusory of explanatory depth, even when claims are false. Thus, LLM-based systems require design guardrails to ensure that stylistic fluency is clearly distinguished from evidentiary support, preventing polished presentation from amplifying falsehoods.
The negative association between authoritative cues and pro-climate attitudes may seem counterintuitive, but it aligns with theories of psychological reactance and autonomy [49]. Highly authoritative language, whether assertive, prescriptive, or certainty laden, can trigger resistance in contentious domains like climate change. Users may perceive it as constraining autonomy or signaling an agenda, leading them to temper their positions. Although authority can convey expertise, it can also dampen persuasion and reduce endorsement of pro-climate positions, especially when the response advocates for the reality of climate change. This dynamic is problematic if it weakens support for evidence-based climate action. It also highlights that autonomy-supportive, collaborative, and transparently uncertain language may be more conducive to constructive attitude change than authority focused messaging.
Nonetheless, credibility assessments of ChatGPT-generated climate content appear to depend more on audience characteristics, such as stable orientations and issue engagement, than on specific message attributes. Individuals who are already attentive to climate change and endorse the scientific consensus tend to judge related information as credible, whether it is expert rebuttal or misinformation. This pattern reflects motivated reasoning [50,51], wherein prior beliefs and involvement shape accuracy judgments more strongly than stylistic or tonal features of messages. An age gradient is also evident: older participants were less trusting of both ChatGPT-generated debunking content and the original climate misinformation, and they reported less pro-climate attitudes, consistent with generational divides in climate opinions and trust in digital technologies. Additionally, prior awareness of ChatGPT was positively associated with perceived credibility of both misinformation and its correction. This “spillover trust” underscores a key risk for AI-mediated communication: familiarity with a tool can generalize into elevated trust in climate-related claims regardless of their accuracy.
These findings should be interpreted in light of several limitations. First, like most nonprobability online samples, our data are not fully representative of U.S. adults; respondents skew younger and more educated, so results reflect an online U.S. adult sample rather than a population-representative one. Since we focus on relationships among variables rather than population-level estimates, this limitation is less consequential, though generalizability should be interpreted with caution. Second, limited variability in key measures reduced statistical power. Because the ChatGPT task used a narrow set of climate statements and a fixed fact-checking prompt, several semantic features (e.g., authority, formality) clustered within a restricted range; user-side variables (e.g., prior climate stance) also showed limited dispersion. As a result, estimates are likely conservative and may miss effects that could emerge with more diverse prompts, message types, and user populations; future work should systematically vary both LLM inputs and survey materials to increase variability. Third, language model-based coding of response attributes may introduce measurement error. Fourth, we did not systematically model variation in participant engagement with the task, even though participants more concerned about climate change may have engaged in longer or more detailed interactions that could have influenced ChatGPT responses and the distribution of study variables. Future work should measure and manipulate engagement explicitly. Finally, as public understanding and use of LLMs, such as trust, awareness of hallucinations, and usage patterns, evolve rapidly, the patterns documented here should be viewed as a snapshot of early adoption, underscoring the need for ongoing monitoring of LLM response attributes as platforms continue to change.
Against this backdrop, LLMs in climate communication can serve as information-support or persuasion-oriented systems. This study emphasizes the former: an accuracy-first, verification task that prioritizes correct, transparent information and clear communication of uncertainty [52]. We discuss tailored engagement only as a complement to information-support, helping users interpret climate information given their prior knowledge and concerns, not as a call to deploy unconstrained persuasive agents [53]. Given tailoring sits close to persuasion, it raises ethical questions. Aligning message style with users’ beliefs or engagement can aid comprehension but risks undisclosed influence. Empirical evidence and regulation for AI-mediated tailoring remain limited. Our results show that LLM messages can shape credibility and attitudes, highlighting dual-use risks (e.g., normalizing climate delay, amplifying misinformation, exploiting vulnerable users). Any deployment should be guided by commitments to climate integrity and public welfare, and constrained by accuracy‑first design, respect for user autonomy, and transparency about when and how content is adapted, so that efforts to close the climate information gap do not erode informed consent or trust [54].
In conclusion, LLMs should employ positive language to highlight genuine efficacy and feasible solutions, but use it judiciously to avoid downplaying risks or conveying unwarranted certainty. Uncertainty must be communicated plainly and consistently, with explicit confidence levels, stated assumptions, and limitations, so users can calibrate their trust. Transparency, verification, and accountability are central to the ethical and responsible deployment of LLMs. Outputs should include semantically rich content (e.g., specific numbers, named entities, and concrete evidence) only when those details are traceable to identifiable sources and consistent with established syntheses of the climate science literature (e.g., major assessment reports and peer‑reviewed summaries). Effective practice therefore entails clearly flagging contested issues, presenting multiple well-sourced viewpoints, explaining the rationale for evidence-based positions, and distinguishing areas of genuine consensus from legitimate debate. We recognize that decisions about which sources and syntheses to treat as authoritative are partly epistemological and ethical, and ultimately require ongoing external calibration to evolving evidence and expert review. Given that perceived credibility and attitudes toward LLM-generated climate communication are driven largely by audience characteristics, such as prior beliefs, engagement level, age, and familiarity with AI, responsible practice should integrate accuracy‑first design that calibrates trust with audience segmentation and tailored engagement strategies that meet users where they are. Together, these principles can advance responsible AI-mediate climate communication and help bridge the gap between scientific consensus and public understanding.
Supporting information
S1 Data. Climate–ChatGPT Credibility and Attitudes Dataset.
This dataset contains the study’s key variables, including participant demographics (gender, age, education); outcome measures (perceived credibility of misinformation, perceived credibility of ChatGPT responses, attitudes towards climate change); control variables (stance on climate change, attention to climate change, prior awareness of ChatGPT); and features of LLM-generated responses (valence, formality, bias, recency, authority, semantic richness).
https://doi.org/10.1371/journal.pclm.0000930.s001
(CSV)
References
- 1. Storani S, Falkenberg M, Quattrociocchi W, Cinelli M. Relative engagement with sources of climate misinformation is growing across social media platforms. Sci Rep. 2025;15(1):18629. pmid:40436972
- 2. Falkenberg M, Galeazzi A, Torricelli M, Di Marco N, Larosa F, Sas M, et al. Growing polarization around climate change on social media. Nat Clim Chang. 2022;12(12):1114–21.
- 3.
Tsang S, Zhou L, Huang Y. Everything about fact-checking: The Commercial Press (HK); 2024.
- 4. Al Khourdajie A. The role of artificial intelligence in climate change scientific assessments. PLOS Clim. 2025;4(9):e0000706.
- 5. Petty RE, Cacioppo JT. The Elaboration Likelihood Model of Persuasion. Advances in Experimental Social Psychology. Elsevier. 1986. p. 123–205.
- 6. Covi MP, Kain DJ. Sea-Level Rise Risk Communication: Public Understanding, Risk Perception, and Attitudes about Information. Environmental Communication. 2015;10(5):612–33.
- 7. Mengesha I, Roy D. Carbon pricing drives critical transition to green growth. Nat Commun. 2025;16(1):1321. pmid:39900895
- 8. Tsang SJ. Attention over content: evaluating the effectiveness of science education in countering climate misinformation. International Journal of Science Education, Part B. 2026;:1–15.
- 9. Ho SS, Chuah ASF, Kim N, Tandoc Jr EC. Fake news, real risks: How online discussion and sources of fact-check influence public risk perceptions toward nuclear energy. Risk Analysis. 2022;42(11):2569–83.
- 10. Xie JJ, Martin M, Rogelj J, Staffell I. Distributional labour challenges and opportunities for decarbonizing the US power system. Nat Clim Chang. 2023;13(11):1203–12.
- 11. Cook J, Ellerton P, Kinkead D. Deconstructing climate misinformation to identify reasoning errors. Environ Res Lett. 2018;13(2):024018.
- 12. Tsang SJ. Misinformation, disinformation, and fake news? Proposing a typology framework of false information. Journalism. 2024;27(3):719–39.
- 13.
Hocevar KP, Metzger M, Flanagin AJ. Source credibility, expertise, and trust in health and risk messaging. Oxford University Press. 2017.
- 14.
Hovland CI. Communication and persuasion: Psychological studies of opinion change. Yale University Press. 1953.
- 15.
Metzger MJ, Flanagin AJ. Psychological approaches to credibility assessment online. The Handbook of the Psychology of Communication Technology. 2015. p. 445–66.
- 16. Albarracin D, Shavitt S. Attitudes and attitude change. Annual Review of Psychology. 2018;69(1):299–327.
- 17. Wong-Parodi G, Berlin Rubin N. Exploring how climate change subjective attribution, personal experience with extremes, concern, and subjective knowledge relate to pro-environmental attitudes and behavioral intentions in the United States. Journal of Environmental Psychology. 2022;79:101728.
- 18. Montemayor C. Language and Intelligence. Minds & Machines. 2021;31(4):471–86.
- 19. Raiaan MAK, Mukta MdSH, Fatema K, Fahad NM, Sakib S, Mim MMJ, et al. A Review on Large Language Models: Architectures, Applications, Taxonomies, Open Issues and Challenges. IEEE Access. 2024;12:26839–74.
- 20. Wang D, Tsang SJ, Zhou Y. Performance unfairness of large language models in cross-language fact-checking. Information Processing & Management. 2026;63(4):104616.
- 21. OpenAI. Introducing ChatGPT. https://openai.com/index/chatgpt/ 2022.
- 22. Huang L, Yu W, Ma W, Zhong W, Feng Z, Wang H, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Trans Inf Syst. 2025;43(2):Article 42.
- 23. Joshi S. Mitigating LLM hallucinations: A comprehensive review of techniques and architectures. 2025.
- 24. Shin D. User Perceptions of Algorithmic Decisions in the Personalized AI System:Perceptual Evaluation of Fairness, Accountability, Transparency, and Explainability. Journal of Broadcasting & Electronic Media. 2020;64(4):541–65.
- 25. Chao F, Wang X, Yu G. Determinants of debunking information sharing behaviour in social media users: perspective of persuasive cues. Internet Research. 2023;34(5):1545–76.
- 26. Zhu D, Xie X, Gan Y. Information source and valence: How information credibility influences earthquake risk perception. Journal of Environmental Psychology. 2011;31(2):129–36.
- 27. Terpstra T, Zaalberg R, de Boer J, Botzen WJW. You have been framed! How antecedents of information need mediate the effects of risk communication messages. Risk Anal. 2014;34(8):1506–20. pmid:24593227
- 28. Van Kleef GA, van den Berg H, Heerdink MW. The persuasive power of emotions: Effects of emotional expressions on attitude formation and change. J Appl Psychol. 2015;100(4):1124–42. pmid:25402955
- 29. Van Damme I, Smets K. The power of emotion versus the power of suggestion: memory for emotional events in the misinformation paradigm. Emotion. 2014;14(2):310–20. pmid:24219394
- 30. Linos E, Lasky-Fink J, Larkin C, Moore L, Kirkman E. The formality effect. Nat Hum Behav. 2024;8(2):300–10. pmid:37996499
- 31. Rennekamp KM, Witz PD. Linguistic Formality and Audience Engagement: Investors’ Reactions to Characteristics of Social Media Disclosures*. Contemporary Accting Res. 2021;38(3):1748–81.
- 32. Steiglechner P, Smaldino PE, Moser D, Merico A. Social identity bias and communication network clustering interact to shape patterns of opinion dynamics. J R Soc Interface. 2023;20(209):20230372. pmid:38086404
- 33. Guilbeault D, Becker J, Centola D. Social learning and partisan bias in the interpretation of climate trends. Proc Natl Acad Sci U S A. 2018;115(39):9714–9. pmid:30181271
- 34. Filieri R, Hofacker CF, Alguezaui S. What makes information in online consumer reviews diagnostic over time? The role of review relevancy, factuality, currency, source credibility and ranking score. Computers in Human Behavior. 2018;80:122–31.
- 35. Dennis AR, Moravec PL, Kim A. Search & Verify: Misinformation and source evaluations in Internet search results. Decision Support Systems. 2023;171:113976.
- 36. Hasanain M, Elsayed T. Studying effectiveness of Web search for fact checking. Asso for Info Science & Tech. 2021;73(5):738–51.
- 37. Zheng W. Lexical richness viewed through lexical diversity, density, and sophistication. Digital Scholarship in the Humanities. 2025;40(2):692–708.
- 38. Shi Y, Lei L. Lexical Richness and Text Length: An Entropy-based Perspective. Journal of Quantitative Linguistics. 2020;29(1):62–79.
- 39.
Camacho-Collados J, Rezaee K, Riahi T, Ushio A, Loureiro D, Antypas D. TweetNLP: Cutting-edge natural language processing for social media. In: 2022. https://doi.org/arXiv:14774
- 40.
Loureiro D, Barbieri F, Neves L, Anke LE, Camacho-Collados J. TimeLMs: diachronic language models from Twitter. In: 2022. https://arxiv.org/abs/2203.829
- 41. Raza S, Reji DJ, Ding C. Dbias: detecting biases and ensuring fairness in news articles. Int J Data Sci Anal. 2022;:1–21. pmid:36065448
- 42.
Halibas AS, Cherian AM, Pillai IG, Reazol LB, Delvo EG, Sumondong GH. Web Ranking of Higher Education Institutions: An SEO Analysis. In: 2020 International Conference on Computation, Automation and Knowledge Management (ICCAKM), 2020. 411–5. https://doi.org/10.1109/iccakm46823.2020.9051481
- 43. Reyes-Lillo D, Morales-Vargas A, Rovira C f. Reliability of domain authority scores calculated by Moz, Semrush, and Ahrefs. Profesional de la información. 2023;32(4).
- 44. Yu W, Shen F, Min C. Correcting science misinformation in an authoritarian country: An experiment from China. Telematics and Informatics. 2022;66:101749.
- 45. Hristidis V, Ruggiano N, Brown EL, Ganta SRR, Stewart S. ChatGPT vs Google for Queries Related to Dementia and Other Cognitive Decline: Comparison of Results. J Med Internet Res. 2023;25:e48966. pmid:37490317
- 46.
Dementieva D, Trifinov I, Likhachev A, Panchenko A. Detecting text formality: A study of text classification approaches. In: 2022. https://arxiv.org/abs/08975
- 47. Lijiang Shen, Monahan JL, Rhodes N, Roskos-Ewoldsen DR. The Impact of Attitude Accessibility and Decision Style on Adolescents’ Biased Processing of Health-Related Public Service Announcements. Communication Research. 2009;36(1):104–28.
- 48. TAN H, YING WANG E, ZHOU B. When the Use of Positive Language Backfires: The Joint Effect of Tone, Readability, and Investor Sophistication on Earnings Judgments. J of Accounting Research. 2014;52(1):273–302.
- 49. Pavey L, Sparks P. Reactance, autonomy and paths to persuasion: Examining perceptions of threats to freedom and informational value. Motiv Emot. 2009;33(3):277–90.
- 50. Bayes R, Druckman JN. Motivated reasoning and climate change. Current Opinion in Behavioral Sciences. 2021;42:27–35.
- 51. Tsang SJ. Biased, not lazy: assessing the effect of COVID-19 misinformation tactics on perceptions of inaccuracy and fakeness. Online Media and Global Communication. 2022;1(3):469–96.
- 52. Atkins C, Girgente G, Shirzaei M, Kim J. Generative AI tools can enhance climate literacy but must be checked for biases and inaccuracies. Commun Earth Environ. 2024;5(1).
- 53. Bai H, Voelkel JG, Muldowney S, Eichstaedt JC, Willer R. LLM-generated messages can persuade humans on policy issues. Nat Commun. 2025;16(1):6037. pmid:40593786
- 54. Schäfer MS, Chen K, Mahl D, Painter J, Volk SC. Climate Change Communication in the Age of Artificial Intelligence. WIREs Climate Change. 2026;17(3).