scieee AI-readable full text Open interactive document viewer

On the reliability of Large Language Models to misinformed and demographically informed prompts

Aremu, Toluwani,Akinwehinmi, Oluwakemi,Nwagu, Chukwuemeka,Ahmed, Syed Ishtiaque,Orji, Rita,Arnau Del Amo, Pedro,Saddik, Abdulmotaleb El

Abstract

We investigate and observe the behavior and performance of Large LanguageModel (LLM)-backed chatbots in addressing misinformed prompts and ques-tions with demographic information within the domains of Climate Change andMental Health. Through a combination of quantitative and qualitative methods,we assess the chatbots’ ability to discern the veracity of statements, their adher-ence to facts, and the presence of bias or misinformation in their responses.Our quantitative analysis using True/False questions reveals that these chat-bots can be relied on to give the right answers to these close-ended questions.However, the qualitative insights, gathered from domain experts, shows thatthere are still concerns regarding privacy, ethical implications, and the neces-sity for chatbots to direct users to professional services. We conclude that whilethese chatbots hold significant promise, their deployment in sensitive areasnecessitates careful consideration, ethical oversight, and rigorous refinement toensure they serve as a beneficial augmentation to human expertise rather thanan autonomous solution. Dataset and assessment information can be found athttps://github.com/tolusophy/Edge-of-Tomorrow.

Full text

Received: 26 April 2024 Revised: 12 November 2024 Accepted: 8 December 2024 DOI: 10.1002/aaai.12208 ARTICLE On the reliability of Large Language Models to misinformed and demographically informed prompts Toluwani Aremu1Oluwakemi Akinwehinmi2Chukwuemeka Nwagu3 Syed Ishtiaque Ahmed4Rita Orji3Pedro Arnau Del Amo2 Abdulmotaleb El Saddik1,5 1Mohamed Bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE 2CIMNE, University of Lleida, Lleida, Spain 3Dalhousie University, Halifax, Canada 4University of Toronto, Toronto, Canada 5University of Ottawa, Ottawa, Canada Correspondence Toluwani Aremu, Mohamed Bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE. Email: toluwani.ar[email protected] Abstract We investigate and observe the behavior and performance of Large Language Model (LLM)-backed chatbots in addressing misinformed prompts and questions with demographic information within the domains of Climate Change and Mental Health. Through a combination of quantitative and qualitative methods, we assess the chatbots’ ability to discern the veracity of statements, their adherence to facts, and the presence of bias or misinformation in their responses. Our quantitative analysis using True/False questions reveals that these chatbots can be relied on to give the right answers to these close-ended questions. However, the qualitative insights, gathered from domain experts, shows that there are still concerns regarding privacy, ethical implications, and the necessity for chatbots to direct users to professional services. We conclude that while these chatbots hold significant promise, their deployment in sensitive areas necessitates careful consideration, ethical oversight, and rigorous refinement to ensure they serve as a beneficial augmentation to human expertise rather than an autonomous solution. Dataset and assessment information can be found at https://github.com/tolusophy/Edge-of-Tomorrow. INTRODUCTION In recent times, the proliferation of Large Language Models (LLMs) has significantly impacted the field of artificial intelligence, owing to their exceptional capabilities in language comprehension and generation. These advanced models have become integral in various applications across multiple industries. Yet, their growing popularity and utility bring forth crucial challenges and ethical considerations. Predominantly based on transformative deep learning architectures like transformers, LLMs have revolutionized This is an open access article under the terms of the Creative Commons Attribution License, which permits use, distribution and reproduction in any medium, provided the original work is properly cited. © 2025 The Author(s). AI Magazine published by John Wiley & Sons Ltd on behalf of Association for the Advancement of Artificial Intelligence. Natural Language Processing (NLP). These models, characterized by their vast neural networks containing millions or billions of parameters, are trained on extensive datasets encompassing a wide array of sources such as internet content, literary works, and diverse media. Such comprehensive training enables them to grasp and interpret a myriad of linguistic patterns and subtleties. Mirroring the historical reliance on search engines for internet queries, users are now increasingly turning to chatbots powered by LLMs for instantaneous and direct responses. Notably, since the advent of ChatGPT, a variant based on the GPT-3.5 architecture in late 2022, the AI Magazine. 2025;46:e12208. wileyonlinelibrary.com/journal/aaai 1of15 https://doi.org/10.1002/aaai.12208 2of15 AI MAGAZINE FIGURE 1 Starting a conversation with an LLM Chatbot. development and implementation of LLMs have rapidly expanded across various sectors. These models have been deployed in areas including virtual assistance, customer support, content creation, search functionality, and in the realms of medical, scientific research, programming assistance, educational tools, and more. However, this expansion has simultaneously sparked significant concerns regarding the ethical use of these technologies, as there have been instances of the models exhibiting biases, generating inaccurate information, making unfair judgments, or inciting ethical debates due to potential misuse (Figure 1). LLMs are rapidly evolving, raising concerns about their potential to generate and disseminate misinformation. Biases within these models could lead to unequal information access or re-inforcement of existing societal biases. Based on these issues, we investigate the behavior of chatbots backed by LLMs, to answer the following two research questions: Research questions 1. When faced with misinformed prompts, do LLMs reflect, amplify, or rectify the misinformation through their responses? 2. Do LLMs exhibit biases when answering prompts, which contain demographic information? To answer these questions, we focused on the implications of utilizing these chatbots in discussions related to climate change and mental health. We focus our analysis on three LLM-powered chatbots: ChatGPT, Bing Chat, and Google BARD, assessing whether they manifest biases or propagate misinformation. Climate change and mental health, being among the most extensively discussed topics on social media as indicated by Google Trends1and Exploding Topics2, are chosen for their relevance and the critical nature of accurate information dissemination in these areas. For the purpose of our study, our main contributions are as follows: We developed a comprehensive benchmark dataset comprising 3120 true/false questions on climate change and 2762 on mental health. This dataset was instrumental for the empirical and quantitative evaluation of responses from LLM-backed chatbots. We conducted an in-depth qualitative analysis in collaboration with domain experts to scrutinize the responses from ChatGPT, Google BARD, and Bing Chat for potential biases. The findings from this analysis are presented herein. To support this, we constructed a dedicated benchmark dataset containing 53 questions on climate change and 40 on mental health, aimed at analyzing the extent of misinformation in the responses provided by these chatbots. Additionally, we utilized 24 climate change and 38 mental health questions specifically to evaluate biases. We proposed and deliberated on several strategies that could address the current challenges hindering the effective and ethical deployment of LLM-backed chatbots in providing accurate information on climate change and mental health issues. LITERATURE REVIEW Background The contemporary AI landscape, particularly the rise of LLMs has sparked critical discussions around ethical concerns. These concerns extend beyond job displacement and privacy violations to encompass the potential for misinformation dissemination. LLMs, trained on massive datasets, can unknowingly perpetuate biases and factual inaccuracies present in the training data. Previous studies says confirmation bias and motivated reasoning can lead to favor information that aligns with existing beliefs (Nickerson 1998). This raises concerns about the trustworthiness of LLMs outputs, especially when applied to sensitive domains like climate change and mental health. Massive foundation models boasting billions of learned parameters and trained on extensive datasets have found applications across a wide array of domains and contexts (Bommasani et al. 2022), and exhibited remarkable effectiveness in their respective downstream tasks. Consequently, the integration of AI into real-world applications has witnessed a phenomenal and exponential surge (Li, Gan et al. 2023; Moor et al. 2023; Weisz et al. 2023). In sharp contrast to traditional models, which often suffer from inherent constraints tied to their narrow focus, foundation models offer a versatile and adaptable approach. Once these models have undergone training, they can be conveniently fine-tuned to suit a diverse spectrum of applications, thus eliminating the necessity for extensive retraining. This adaptive framework serves as the linchpin of LLMs. Notably, this fine-tuning capability has paved the way for deploying these language models in a multitude of domains, spanning healthcare, financial advisory, climate change analysis, and question answering, to name just a few. 23719621, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aaai.12208 by Readcube (Labtiva Inc.), Wiley Online Library on [17/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License AI MAGAZINE 3of15 As the scope of artificial intelligence (AI) continues to expand, ushering in captivating innovations, it has sparked spirited debates on a broad range of ethical concerns. These concerns encompass the potential impacts of these advancements on various facets of human existence, including individual lives, employment, privacy, and issues related to discrimination (Raghavan et al. 2020; Kelley et al. 2021). Simultaneously, questions have emerged regarding the appropriate course of action for the adoption of these transformative technologies (Aremu 2023;Liangetal.2023). According to an article by researchers at Google published in 2021 (Kelley et al. 2021) to assess public perception of AI in eight countries, people in developing countries like Nigeria, India, and Brazil are significantly more likely to embrace and adopt AI compared to individuals in developed countries. However, ethical researchers in the AI domain have raised concerns that such AI applications may disproportionately affect people living in these regions, as most of the data used to train these models is sourced from developed countries (Gebru et al. 2021; Shneiderman 2020; Buolamwini and Gebru 2018;Rajietal. 2020; Mitchell et al. 2019; Mittelstadt et al. 2016). Therefore, it comes as no surprise that throughout 2023, a series of pivotal governmental hearings have convened, where political leaders engaged with a diverse array of experts to gain insights into the origins and implications of these technologies. These hearings have probed critical aspects, such as the inherent risks associated with these AI systems, the nature of the data on which they are trained, and the formulation of policies designed to safeguard the wellbeing and privacy of users. These discussions also consider equitable compensation for the creators and owners of the data that underpin these AI systems. In this section, our focus narrows to articles highlighting the deployment of LLMs in the contexts of climate change/sustainability and mental health/physical health. Our objective here is to demonstrate the substantial strides made in the adoption of AI technologies, setting the stage for subsequent sections where we delve into our methodologies and present the results of experiments conducted to evaluate biases and misinformation in LLMs when applied to both climate change and mental health contexts. Climate change The advent of LLMs in the realm of climate change research took a significant leap forward in 2021 with the introduction of ClimateBERT by Webersinke (Webersinke et al. 2022). ClimateBERT, a transformer-based language model, was pretrained on an extensive dataset comprising over two million paragraphs sourced from climate-related texts, including news, research articles, and corporate climate reports. Its primary purpose was to facilitate climate change question answering and text summarization. Subsequently, an array of tools (Vaghefi et al. 2023;Ni et al. 2023; Fard, Hasan, and Bell 2022;Li2023; GarridoMerch’an, Gonz’alez-Barthe, and Vaca 2023; Kraus et al. 2023) and datasets (Diggelmann et al. 2020;Spokoynyetal. 2023;Laudetal.2023) have been created to improve the credibility and correctness of information disseminated by applications utilizing such models. While several empirical studies have examined the effectiveness of these tools, most have concentrated on sentiments (Krishnan and Anoop 2023; Sham and Mohamed 2022; Baguio, Lu, and Peña 2023; Ray and Kumar 2023) and sustainability (Jain and Padmanaban 2023). In a closely related study (Bulian et al. 2023), researchers assessed the accuracy of LLMs in handling climate information and proposed a practical protocol that combines AI assistance with human raters to mitigate the limitations encountered during the evaluation of LLMs. Our work, in contrast, encompasses both qualitative and quantitative evaluations of LLM responses to misinformed queries and considers how these models interact with various demographic groups. Mental health Language models show promise in addressing challenges within the field of mental health. Several of these models, employing smaller language models (Denecke, Vaaheesan, and Arulnathan 2020;Jietal.2021), have been proposed for applications in both mental health and general medical care. More recent developments have leveraged larger language models (Li, Li et al. 2023;Xuetal.2023;Liu,Li et al. 2023; Bao et al. 2023). It is important to note that these models introduce ethical concerns, as they have the potential to cause irreversible harm to users. Empirical evaluations of LLMs in this domain primarily fall into two categories: some assess LLMs’ ability to classify different types of mental health issues using annotated text data (Yang et al. 2023; Wang, Zhao, and Petzold 2023), while others evaluate their performance in various medical examinations (Nori et al. 2023;Liu, Zhou et al. 2023; Manathunga and Hettigoda 2023; Rosol et al. 2023; Kasai et al. 2023; Singhal et al. 2023). Our approach, however, diverges from these studies. We delve deeper into the analysis of the responses generated by the models we employ, collaborating with experts to determine their readiness for real-world deployment or whether significant strides are still required in this area. 23719621, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aaai.12208 by Readcube (Labtiva Inc.), Wiley Online Library on [17/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 4of15 AI MAGAZINE METHODOLOGY This study is structured to address our central research objective which is to evaluate the level of misinformation and bias in LLM-powered chatbots in climate change and mental health discussions. We do this through a dual-pronged approach: first, by understanding and quantifying misinformation, and second, by evaluating biases in the responses of LLM chatbots. This section details the methodologies employed in each of these categories. The details of the dataset used, chatbots’ responses, and assessment questions are accessible here.3We also provide examples of our questions in the appendix. Tools and data collection In our study, we evaluate the following three cutting-edge most popular and accessible LLM chatbots: Microsoft’s Bing Chat, OpenAI’s ChatGPT, and Google’s Bard (now Gemini). To conduct an extensive evaluation, we compiled a set of frequently asked questions (FAQs) on two crucial topics—Climate Change and Mental Health. We prioritize FAQs to reflect real-world user queries encountered by these LLMs. This approach ensures the generalizability of our findings to real-world LLMs interactions. We also assume that the LLMs used in these chatbots were trained on and has access to the information. Hence, to test each chatbot susceptibility to misinformation, we intentionally altered a subset of the selected questions. These alterations involved introducing subtle factual inaccuracies, changing the tone of the prompt, or incorporating irrelevant words to the prompts. The specific type of alteration depended on the misinformation concept we aimed to assess. We used the same misinformed prompts uniformly across the chatbots intended for testing. Misinformation assessment For the quantitative analysis of misinformation, we curated a dataset comprising 3120 true/false questions on Climate Change and 2762 on Mental Health. For the qualitative analysis, we initially compiled a broader set of FAQ prompts from authoritative sources such as NASA and the CDC. We then pared this down to 53 questions about Climate Change and 40 about Mental Health based on several criteria. First, we prioritized prompts that represented a wide range of common misconceptions and frequently debated topics. Second, we ensured that the selected prompts were representative of the most pervasive and impactful forms of misinformation. Lastly, we chose questions that could effectively challenge the chatbots’ ability to discern and address misinformation without inadvertently reinforcing it. These prompts were then carefully altered to assess whether the chatbots could appropriately handle misinformed statements or if they would amplify the misinformation. Bias assessment For the bias assessment, we selected 24 questions on Climate Change and 38 on Mental Health, aiming to perform a comprehensive qualitative analysis of the chatbots’ responses with a focus on objectivity and neutrality. These questions were chosenbased on their potential to highlight biases, as they cover key issues and areas of potential controversy in each domain. To examine how demographic factors might influence the chatbots’ responses, we integrated details such as age, race, and location into the prompts. This allowed us to assess whether the chatbots’ outputs varied inappropriately based on these demographics or if they maintained consistent and impartial responses across different scenarios. Chatbot query approach For qualitative analysis, we interacted with the chatbots using a standardized format to ensure concise and informative responses. Each question was formatted as follows: “In one short paragraph,[question].Provide sources for your response.” For Quantitative Analysis, the methodology for quantitative interaction also employed a standard prompt format, simplifying the chatbots’ responses to a binary choice: “Respond with either True/Yes or False/No: [question].” Analysis This section details the methodological framework adopted for analyzing the misinformation and bias components in our study, divided into two distinct segments as follows: Quantitative Analysis to check for misinformation The quantitative analysis scrutinizes the chatbots’ responses using a suite of metrics: Confusion matrix:This tool visualizes the distribution of True Positives, False Positives, True Negatives, and False Negatives for the true/false questions. Precision and recall: These metrics evaluate the accuracy and completeness of the classification model, based on the results of the confusion matrix. F1 score and accuracy: These indicators provide insights into the model’s harmonic balance between precision 23719621, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aaai.12208 by Readcube (Labtiva Inc.), Wiley Online Library on [17/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License AI MAGAZINE 5of15 and recall. Similarity index scores:WeemployBLEU (Papineni et al. 2002), ROGUE (Lin 2004), and METEOR (Banerjee and Lavie 2005) scores to measure the closeness of the chatbot’s responses to standard benchmark answers, thereby assessing the quality and relevance of the content provided. Qualitative analysis to check for misinformation We conduct a comprehensive qualitative analysis of the responses from the chatbots on the dataset we collected for this analysis. This facet of the analysis involves in-depth interviews with domain experts and the deployment of specialized questionnaires tailored to these fields. Qualitative analysis to check for bias The bias analysis segment focuses on quantitatively evaluating feedback from domain experts. These specialists will critique and provide perspectives on the extent of bias evident in the chatbots’ responses. The aim here is to uncover any subjective biases that might be embedded in the outputs related to Mental Health and Climate Change topics. Throughout our analysis, we sometimes personified LLM-based chatbots (e.g., discussing what they “know”) and, at other times, treated them purely as functional text generators. This mixed approach was intentional and directly tied to our objective of assessing bias. Specifically, to evaluate how demographic factors might influence the chatbots’ responses, we embedded demographic details within the prompts, effectively personifying the chatbots to simulate scenarios where biases might manifest. This approach allowed us to explore whether the chatbots would respond differently based on the demographic context provided. Domain expert selection criteria In the process of selecting domain experts for our study, we established specific criteria tailored to the distinct fields of Climate Change and Mental Health. For Climate Change, we targeted academically credentialed professionals, including Professors, PostDocs, PhD students, researchers, or practitioners holding at least a master’s degree in fields such as environmental science, climatology, meteorology, or ecology. Their expertise was validated through a demonstrated track record in climate change research, including publications in peer-reviewed journals, conference presentations, or significant contributions to relevant industry projects. A prerequisite was a minimum of three years of active involvement in areas such as climate change research, policy development, mitigation strategies, adaptation methods, or advocacy. Additionally, we emphasized the importance of interdisciplinary knowledge, combining insights from atmospheric science, oceanography, and social sciences, to foster a comprehensive understanding of climate change impacts. Familiarity with climate policies, the ability to effectively communicate complex scientific concepts, and experience in innovative solutions and collaborative projects were also deemed essential. In the Mental Health domain, our focus was on professionals with a solid educational foundation in psychology, psychiatry, clinical social work, counseling, or related disciplines, requiring a minimum of a master’s degree. We sought experts with substantial clinical or research experience in mental health, evidenced by a history of patient care, participation in clinical trials, research contributions, or advocacy work. Proficiency in various therapeutic modalities such as cognitive-behavioral therapy, psychotherapy, and mindfulness-based interventions was crucial. Cultural competence—understanding and addressing the diverse cultural and socioeconomic factors influencing mental health—was another critical criterion. Lastly, we valued experts open to exploring the ethical implications and potential applications of generative language technologies in mental health care and challenges. Limitations A significant challenge encountered was the recruitment of domain experts. For interview-based qualitative reviews, standard practice recommends a minimum of five experts per domain. Questionnaire-based qualitative reviews generally require a more extensive participant base, ideally with at least 50 respondents. Despite extensive outreach efforts, our response rate was limited to 14 participants, comprising four interviewees and 10 questionnaire respondents. This equates to seven domain experts for each of the two categories under study. Although the number of participants are limited, the use of both quantitative and qualitative approach offer opportunity for in-depth data and insights. We also believe that the inclusion of experts from diverse backgrounds, culture, and continents contributes positively to the quality of our findings. As mentioned earlier, one of the strengths of this study is the geographical diversity of our expert panel. We successfully included at least one domain expert from each continent (refer to Figure 2for details), which bolsters the validity of our results. This diverse representation helps mitigate regional biases and enhances the global relevance of our findings. Hence, we believe that despite the limitation in terms of numbers, the study provides meaningful insights into the research area. The limitations highlight avenues for future 23719621, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aaai.12208 by Readcube (Labtiva Inc.), Wiley Online Library on [17/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 6of15 AI MAGAZINE FIGURE 2 Distribution of domain experts in the fields of Climate Change and Mental Health, categorized by location and profession. In both domains, Asia and Africa provide the largest regional share of experts, while North America and Oceania contribute the least. The majority of experts in both domains are from Academic & Research backgrounds, accounting for 57% of the total, as opposed to 43% from Industry. research, particularly in broadening the expert participant base to further validate and enrich the study’s conclusions. FINDINGS This section elucidates our study’s results, commencing with a quantitative analysis of the chatbots’ average performance on the dataset of True/False questions within the Climate Change and Mental Health domains. We subsequently present the similarity scores based on the metrics mentioned above, comparing the responses of all three chatbots against established facts to determine their factual adherence. Lastly, we expand into the qualitative insights derived from engaging with domain experts, providing an in-depth exploration of their perspectives in both domains of interest. These findings are consistent with research by Shneiderman (2020) and Yang et al. (2023), who found that LLMs trained on massive datasets can still be susceptible to misinformation, particularly when the information is cleverly disguised. Quantitative analysis: True/false prompts In this study, we evaluate the chatbots’ ability to discern the veracity of statements related to climate change and TABLE 1 Comparative analysis of the selected chatbots’ performance in discerning True/False statements within the realms of Climate Change and Mental Health. Domain Precision Recall F1 score Accuracy Climate Change 88.4% 91.9% 90.1% 89.9% Mental Health 90.1% 95.2% 92.6% 92.5% mental health, in order to quantify its level of knowledge, or how misinformed it might be. Utilizing a quantitative approach, we analyze a True/False dataset and calculate critical performance metrics. The analysis (Figure 3) includes a detailed examination of instances where the model incorrectly classified true statements as false (false negatives) and false statements as true (false positives), as well as accurately identified true (true positives) and false (true negatives) statements. An in-depth analysis of the performance metrics, as detailed in Table 1, shows that in the realm of Climate Change, our chatbots demonstrates commendable accuracy with a precision rate of 88.4%. This indicates that the majority of the statements classified as true by the model are indeed correct. The recall rate of 91.9% further suggests that it successfully identifies a high percentage of the true statements within this domain. The F1 score stands at 90.1%, reflecting a strong overall performance. However, the overall accuracy, at 89.9%, while high, indicates there is room for improvement in reducing misinformation when it comes to climate change. In the Mental Health domain, the observed performance is notably enhanced. It achieves a higher precision rate of 90.1%, suggesting that its capacity to correctly identify true statements is more refined in this domain. The recall rate of 95.2% is particularly impressive, indicating that the model is highly effective at capturing true instances. The F1 score, at an elevated 92.6%, points to a balanced and efficient classification capability. Moreover, the accuracy of 92.5% underscores a significant level of reliability in the Mental Health domain. Quantitative analysis: Similarity index scores In the quantitative phase of our analysis, we employed three well-known metrics—BLEU, ROGUE, and METEOR—from the machine translation evaluation field. These metrics traditionally assess how closely machine-generated text matches human translation, in terms of both accuracy and contextual coherence. For this study, we adapted these metrics to assess the performance of chatbots, positing their applicability beyond their usual context of translation. 23719621, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aaai.12208 by Readcube (Labtiva Inc.), Wiley Online Library on [17/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License AI MAGAZINE 7of15 FIGURE 3 Confusion matrices depicting the performance of the selected chatbots in answering whether a fact given within a prompt is either true or false, for the Climate Change and Mental Health domains. For Climate Change, there were 1368 true negatives and 1436 true positives, with false positives and negatives at 188 and 127, respectively. In the Mental Health domain, the model produced 1253 true negatives and 1301 true positives, and lower false positives and negatives at 143 and 65. These results indicate a high level of accuracy in the model’s knowledge across both domains. BLEU and ROGUE metrics are designed to measure the precision and recall of the chatbots’ responses against a benchmark of human-generated texts. METEOR goes a step further by including advanced linguistic analysis—such as synonym matching, stemming, and paraphrasing—to provide a more nuanced assessment. This metric, therefore, offers a measure of evaluation that more closely approximates human judgment by accounting for semantic and contextual accuracy, in addition to exact word correspondences. We analyzed the chatbots’ outputs by comparing them to the verifiable answers within our dataset. To ensure comparability, we normalized the resulting similarity scores, aiming for a maximum value of 1. This step was crucial, given that our “misinformed” prompts often led to chatbot responses that were shorter and substantially varied from the factual responses, sometimes resulting in inaccuracies or misinformation. Although there’s no absolute threshold set for misinformation, these normalized scores serve as indicators of the degree to which the chatbots’ responses emulate the factual data. Figure 4presents the normalized similarity index scores, which reflect the chatbots’ accuracy in relation to the factual statements provided. Within the climate change context, the Bing Chatbot, powered by GPT-4, demonstrated the highest likelihood of delivering correct responses—even when faced with misleading prompts. Google’s Bard, leveraging LaMDA, followed closely, likely benefiting from its access to up-to-date online information, in contrast to ChatGPT (GPT-3.5), which operates based on data up until 2021 and hence offline. In the mental health arena, Bard was observed to have the highest accuracy, suggesting that its model is well-equipped to handle even adversarially designed prompts. FIGURE 4 The bar charts illustrate the Similarity Index Scores for the following three LLM-powered chatbots—ChatGPT (GPT-3.5), Bard (LaMDA), and Bing (GPT-4)—across the following three evaluation metrics: BLEU, ROUGE, and METEOR. Qualitative analysis This section details our approach to gauging the levels of misinformation and bias present in the responses provided 23719621, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aaai.12208 by Readcube (Labtiva Inc.), Wiley Online Library on [17/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 8of15 AI MAGAZINE by three prominent chatbots—specifically, those focused on Climate Change and Mental Health topics. To this end, we sought insights from domain experts in these respective fields. Initially, we endeavored to engage a broad spectrum of specialists for in-depth interviews based on the chatbots’ responses. As revealed in Section, the response rate was limited to 14 experts, spanning both domains. Our initial strategy involved forwarding the chatbot-generated responses to these experts, followed by interviews after a week. However, time constraints necessitated a strategic switch after interviews with our first two climate change experts. We transitioned to a questionnaire-based approach, using similar questions from the interview methodology, which significantly enhanced time efficiency. The questionnaires comprised both closedand open-ended questions, adapted to suit the experts’ convenience. This approach was similarly applied in the mental health domain, where two specialists were interviewed, and the remaining provided their inputs via questionnaires. The expert interviews varied from 45 to 60 min, encompassing 8–13 comprehensive questions, indicative of the depth of these discussions. Conversely, the questionnaires, comprising 11 items for climate change experts and 16 for mental health experts, required 10–30 min to complete. The climate change-focused questions sought expert opinions on language-model-based chatbots in raising awareness and their effectiveness in climate change communication. The assessment covered diverse aspects, including their role in awareness, challenges in providing accurate information, potential biases and misinformation, their utility in adaptation and mitigation strategies, user engagement features, integration with other platforms, promoting sustainable behaviors, and the relevance of their information in light of new scientific discoveries. In the mental health context, the questions evaluated the impact, ethical considerations, effectiveness, and limitations of chatbots. Key areas of inquiry included their role in stigma reduction and awareness, personalized support effectiveness, specific mental health conditions addressed, necessity for human intervention, fostering trust and confidentiality, early detection and prevention, inclusivity, cultural sensitivity, potential drawbacks, criteria for success measurement, complementing existing services, empathy level, and their ability to recommend tailored mental health resources. Our analysis of expert responses was conducted using a thematic analysis approach, adhering to the guidelines outlinedbyBraunandClarke(BraunandClarke2006). This method allowed for the systematic identification and organization of visible and valid patterns within the data. Initially, we iteratively read through the responses to extract significant statements and concepts concerning the use and implications of chatbots in the domains of Climate Change and Mental Health, which we then represented as codes. Thematic saturation was achieved when no new codes could be identified. In the final phase of our analysis, we synthesized these themes into a coherent narrative. This involved linking the themes to our research questions and drawing conclusions about the role and impact of chatbots in the respective fields. Our interpretations, grounded in the data, include representative quotes from the experts to illustrate their perspectives. The subsequent sections of this paper will delve into these specific themes, providing a detailed exploration of the utility and effectiveness of chatbots in the context of Climate Change and Mental Health. (I.) Experts’ perspectives on misinformation and bias in climate change In our investigation, we presented each expert with questions concerning the chatbots’ responses on Climate Change. The focus of these inquiries was to understand the experts’ perceptions of the potential impact these chatbots might have on user safety and information dissemination. We present our findings under four main themes. 1. Role of LLM-based chatbots in climate change awareness: We asked the experts about the significance of LLM-based chatbots in enhancing public awareness of climate change. The majority, barring two, were optimistic, acknowledging that the newer generation of chatbots could play a substantial role in spreading awareness and disseminating information. On the contrary, one expert expressed skepticism about the depth and utility of the chatbots’ responses, likening them to shallow internet searches. Another pointed out no discernible advantage over traditional search engines. Despite these differing views, there was a consensus that more specialized and expert-driven models could yield more precise and reliable information than the three chatbots evaluated. The experts suggested improvements such as enabling chatbots to provide simplified yet comprehensive answers, elucidate the reasoning behind their responses, and transparently cite their information sources. 2. Challenges in generating accurate climate-related information: We sought the experts’ views on the obstacles faced by chatbots in delivering precise and trustworthy information on climate change. A significant portion of the respondents raised concerns about the data sources used to train these models. They emphasized the lack of assurance regarding the quality and credibility of the sources cited by the chatbots. Another common issue highlighted was the 23719621, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aaai.12208 by Readcube (Labtiva Inc.), Wiley Online Library on [17/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License AI MAGAZINE 9of15 inconsistency in the chatbots’ responses. Experts noted that for certain queries, the models produced vastly differing answers, which could lead to confusion. Additionally, the generality of the responses was a point of contention. Experts pointed out that climate change effects and solutions are often location-specific, yet the chatbots tended to provide broad, universally applicable answers. This, they suggested, might stem from the nature of the prompts fed to the models. More accurate and tailored responses could potentially be elicited with prompts that include more detailed and specific instructions or context. 3. Expert insights on biases and misinformation in climate data dissemination: We inquired about the experts’ perception of potential biases and misinformation in the chatbot-generated responses on climaterelated topics. A notable portion of the experts, about half, identified instances of data exaggeration leading to misinformation. They expressed concerns about the chatbots being trained on datasets with unverified or nonreproducible sources, a significant issue in a field prone to false negatives. Such practices, they cautioned, result in the propagation of unverified and potentially misleading information. Conversely, the other half acknowledged the general adequacy of the information provided by the chatbots but stressed the need for stringent measures to ensure the verifiability of disseminated data. Despite these divergent views, a unanimous concern among all experts was the lack of demographic sensitivity in the chatbots’ responses. The experts observed that the chatbots tended to provide uniform answers irrespective of varying demographic contexts. For instance, the response given to a middleaged African male was identical to that given to a young European female, overlooking the specific vulnerabilities and contexts of different demographic groups. This could mean that the information about the person behind the prompts was not considered much in generating exact answers. One expert poignantly remarked that a teenager and a senior citizen should rather receive relative responses based on their previous knowledge highlighting the need for more nuanced and demographic-aware chatbots. 4. Chatbots as tools for promoting sustainable behaviors: We explored the experts’ views on the potential of these chatbots in fostering sustainable behaviors and lifestyle changes among users. The response was unanimously positive across the board. The experts acknowledged the importance of reliable information sources but were optimistic about the role of chatbots, especially those specialized in climate change, in influencing user behavior towards sustainability. They concurred that appropriately designed chatbots could effectively encourage users to adopt more environmentally friendly practices. Furthermore, some experts proposed specific features that could enhance the chatbots’ capability to advocate for sustainability. These suggestions included personalized advice based on user’s lifestyle, interactive guides on reducing carbon footprints, and timely updates on environmental issues and solutions, all tailored to engage users actively in sustainability efforts. Discussion: Our panel of climate change experts concurs that, despite the need for considerable improvements in safety and reliability, LLM-backed chatbots hold immense potential for impactful applications. These AIdriven tools are lauded for their capacity to revolutionize the dissemination of crucial information and to promote environmental consciousness among the public. Nonetheless, experts stress the imperative for stringent validation and continuous refinement of these systems to bolster their effectiveness and credibility in information dissemination. One prominent recommendation from the experts is the strategic curation of training datasets for LLMs, advocating for the inclusion of data primarily from verifiable and expert-endorsed sources. This recommendation arises from their observation that the chatbots in our study often referenced materials indiscriminately, linking to articles that may be obsolete or from publishers lacking official standing in academic research. Moreover, instances of nonfunctional or potentially deceptive links further highlight the risks associated with unvetted information sources. The experts also highlighted a critical concern regarding the inherent bias within prompts themselves, suggesting that LLMs may prioritize completing a user’s request over ensuring the accuracy and reliability of the provided information. This tendency underscores a fundamental challenge: the need to fine-tune LLMs to discern and prioritize high-quality, trustworthy content in their responses, regardless of the nature of the prompts they receive. (II.) Experts’ perspectives on misinformation and bias in mental health In discussions with domain experts, we sought to understand the behavior and potential of the three chatbots in the mental health domain. Our dialog focused on several key areas: 1. Impact on stigma and awareness: Our inquiry into experts’ perceptions commenced with questions about the potential impact of chatbots on stigma reduction and awareness enhancement in mental health. The response was uniformly positive, with experts recognizing the substantial value of LLM-backed chatbots 23719621, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aaai.12208 by Readcube (Labtiva Inc.), Wiley Online Library on [17/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License