A Survey on Automatic Credibility Assessment Using Textual Credibility Signals in the Era of Large Language Models IVAN SRBA,Kempelen Institute of Intelligent Technologies, Slovakia OLESYA RAZUVAYEVSKAYA,The University of Sheffield, UK JOÃO A. LEITE,The University of Sheffield, UK ROBERT MORO,Kempelen Institute of Intelligent Technologies, Slovakia IPEK BARIS SCHLICHT,Deutsche Welle, Germany and Universitat Politècnica de València, Spain SARA TONELLI,Fondazione Bruno Kessler, Italy FRANCISCO MORENO GARCÍA,Universidad Politécnica de Madrid, Spain SANTIAGO BARRIO LOTTMANN,Universidad Politécnica de Madrid, Spain DENIS TEYSSOU,Agence France-Presse, France VALENTIN PORCELLINI,Agence France-Presse, France CAROLINA SCARTON,The University of Sheffield, UK KALINA BONTCHEVA,The University of Sheffield, UK MARIA BIELIKOVA,Kempelen Institute of Intelligent Technologies, Slovakia In the age of social media and generative AI, the ability to automatically assess the credibility of online content has become increasingly critical, complementing traditional approaches to false information detection. Credibility assessment relies on aggregating diverse credibility signals – small units of information, such as content subjectivity, bias, or a presence of persuasion techniques – into a final credibility label/score. However, current research in automatic credibility assessment and credibility signals detection remains highly fragmented, with many signals studied in isolation and lacking integration. Notably, there is a scarcity of approaches that detect and aggregate multiple credibility signals simultaneously. These challenges are further exacerbated by the absence of a comprehensive and up-to-date overview of research works that connects these research efforts under a common framework and identifies shared trends, challenges, and open problems. In this survey, we address this gap by presenting a systematic and comprehensive literature review of 175 research papers, focusing on textual credibility signals within the field of Natural Language Processing (NLP), which undergoes a rapid transformation due to advancements in Large Language Models (LLMs). While positioning the NLP research into the the broader multidisciplinary landscape, we examine both automatic credibility assessment methods as well as the detection of nine categories of credibility signals. We provide an in-depth analysis of three key categories: 1) factuality, subjectivity and bias, 2) persuasion techniques and logical fallacies, and 3) check-worthy and fact-checked claims. In addition to summarising existing methods, Authors’ addresses: Ivan Srba, [email protected], Kempelen Institute of Intelligent Technologies, Bratislava, Slovakia; Olesya Razuvayevskaya, o[email protected], The University of Sheffield, Sheffield, UK; João A. Leite, jaleite1@ sheffield.ac.uk, The University of Sheffield, Sheffield, UK; Robert Moro, [email protected], Kempelen Institute of Intelligent Technologies, Bratislava, Slovakia; Ipek Baris Schlicht,
[email protected], Deutsche Welle, Bonn/Berlin, Germany and Universitat Politècnica de València, Valencia, Spain; Sara Tonelli,
[email protected], Fondazione Bruno Kessler, Trento, Italy; Francisco Moreno García, fran.mor[email protected], Universidad Politécnica de Madrid, Madrid, Spain; Santiago Barrio Lottmann, [email protected], Universidad Politécnica de Madrid, Madrid, Spain; Denis Teyssou, Denis.TEYSSOU@ afp.com, Agence France-Presse, Paris, France; Valentin Porcellini, V[email protected], Agence France-Presse, Paris, France; Carolina Scarton, [email protected], The University of Sheffield, Sheffield, UK; Kalina Bontcheva, [email protected], The University of Sheffield, Sheffield, UK; Maria Bielikova, maria.bieliko[email protected], Kempelen Institute of Intelligent Technologies, Bratislava, Slovakia. ©2025 Copyright held by the owner/author(s). Publication rights licensed to ACM. This is the author’s version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in ACM Transactions on Intelligent Systems and Technology,https://doi.org/10.1145/3770077. ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:2 Srba et al. datasets, and tools, we outline future research direction and emerging opportunities, with particular attention to evolving challenges posed by generative AI. CCS Concepts: •General and reference → Surveys and overviews;•Computing methodologies → Natural language processing;•Human-centered computing →Social media. Additional Key Words and Phrases: Credibility Assessment, Credibility Signals, Natural Language Processing, NLP, Literature Survey, Large Language Models ACM Reference Format: Ivan Srba, Olesya Razuvayevskaya, João A. Leite, Robert Moro, Ipek Baris Schlicht, Sara Tonelli, Francisco Moreno García, Santiago Barrio Lottmann, Denis Teyssou, Valentin Porcellini, Carolina Scarton, Kalina Bontcheva, and Maria Bielikova. 2025. A Survey on Automatic Credibility Assessment Using Textual Credibility Signals in the Era of Large Language Models. ACM Trans. Intell. Syst. Technol. 0, 0, Article 0 ( 2025), 81 pages. https://doi.org/10.1145/3770077 1 INTRODUCTION Tackling false information (i.e., disinformation and misinformation) has attracted a substantial attention in recent years from researchers, media professionals, AI and social media practitioners as well as the general public. Researchers have approached false information detection through a variety of classification tasks, leveraging AI technologies such as machine learning, natural language processing, computer vision, or social network analysis [ 290 ]. The primary stream of research works on false information detection, commonly also denoted as a fake news detection [ 132 , 296 , 371 ], assesses the veracity (i.e., factual accuracy) of information by means of a single-label prediction, typically a binary one (i.e., false/true content) [ 296 ]. Some works go beyond a binary classification and predict multiple classes, such as veracity levels, commonly used by fact-checking organizations (e.g., true, mostly true, mixture, mostly false, false). However, veracity alone offers only a limited perspective on the potential harm of online content. In practice, it should be considered alongside credibility. By proceeding from existing (quite heterogeneous) definitions [ 14 , 127 , 259 , 266 , 285 ], under the credibility we understand the perceived trustworthiness, accuracy, and reliability of information or its source. It reflects the extent to which content is believed to be factual, unbiased, and free from manipulation or deception. False content that is perceived as highly credible by its audience can lead to significantly greater harm than false content that is clearly perceived as lacking credibility. Credibility assessment therefore provides an interesting potential to extend and complement the existing stream of works on false information detection. It typically follows a two-step process. First, granular credibility signals (also referred to as credibility indicators) are detected. Second, these signals (serving as evidence) are aggregated into a single ordinal credibility label or a numerical credibility score. Common examples of credibility signals include subjectivity and bias present in the content, persuasion techniques, logical fallacies or presence of machine-generated content, which has become increasingly relevant with the rise of generative AI. While credibility assessment can be conducted manually by the end users, its full potential lies in automated detection of credibility signals and their consequent aggregation by a credibility assessment algorithm. These automated approaches may range from simple (e.g., rule-based or heuristic) methods to more advanced systems powered by AI, capable of nuanced and context-aware content analysis. Such automation not only improves scalability but also enables real-time evaluation of large volumes of online content. The benefit of credibility signals is their flexibility and wide-applicability. Besides their primary usage, to be aggregated into a credibility label/score, they can also be utilized and interpreted by a user directly (e.g., as an “information nutrition label” [ 98 ]), provide valuable inputs to information retrieval engines or recommender systems in order to prefer content/sources associated with a high ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:3 level of credibility [ 338 ], or even serve as features in subsequent classification tasks (including false information detection itself [ 301 ]). Credibility assessment and credibility signals are, furthermore, especially useful in cases in which we cannot easily ascertain that something is true or false (on a single veracity dimension) – concept of credibility provides a necessary level of granularity to represent a potentially complex information from multiple perspectives. Credibility signals are also more easily interpretable and naturally support semi-automatic human-centered AI approaches that attempt supporting instead of replacing a human expert or a lay person (individual credibility signals can be automatically detected and their interpretation can be subsequently done by a human). Study by Lu et al . [189] showed that even if people are influenced by others when judging the veracity of online content, providing accurate AI-based credibility indicators can effectively improve people’s ability in detecting false information. Last but not least, they are also aligned with journalistic/fact-checking workflows for identifying trustworthy information (to be cited) or potential false information (to be fact-checked). In contrast to other related research areas (including false information detection), research works on automatic credibility assessment are mostly carried out in isolation. As we show in this survey, there is a lack of works addressing multiple categories of credibility signals at once and, moreover, many research works do not explicitly state that the outcome of their solution can actually serve as a credibility signal. Getting familiar with the existing works on automatic credibility assessment and detection of credibility signals is challenging due to an ambiguous and inconsistent terminology used by researchers, a lack of clear definition of credibility and credibility signals as well as due to a missing standardized taxonomy of various credibility signals’ categories. At the same time, automatic credibility assessment can significantly benefit from a holistic approach. For example, due to common underlying similarities, there is a big potential of transfer or multi-task learning. While there is a wide range of possible credibility signals, we specifically focus on such signals that relate to textual content and can be automatically detected by Natural Language Processing (NLP) methods. This area is currently at the spotlight due to the recent advancements of Large Language Models (LLMs). Their rapid development has a revolutionizing effect on many text classification tasks, automatic credibility assessment and detection of credibility signals not being an exception. Despite multiple surveys addressing false information detection (e.g., [ 1 , 132 , 296 , 319 , 371 ]), to the best of our knowledge, there is no survey on automatic credibility assessment and detection of credibility signals from the NLP perspective. In this paper, we therefore provide the first comprehensive and systematic literature overview of credibility signal detection with the focus on NLP and LLMs. The main contributions are as follows: (1) We address the fragmentation of existing works on automatic credibility assessment and credibility signals detection. We connect these research efforts under a common framework and identify shared trends, challenges, and open problems. In this way, we motivate future research in connecting and integrating individual outcomes (i.e., to move from currently isolated individual signals to credibility assessment effectively combining them). (2) We provide a necessary background overview of definitions and dimensions of credibility assessment (including a unified categorization of credibility signals based on their various taxonomies and list used in the existing literature) to setup the common understanding that is currently missing in the existing works. (3) We analyse and describe a total of 175 NLP works tackling the automatic credibility assessment and detection of credibility signals. We provide an in-depth description of 3 key categories of credibility signals, which have been selected due to a considerable NLP research interest and their recognition as important by end users, namely: (i) factuality, subjectivity and bias; (ii) persuasion techniques and logical fallacies; and (iii) check-worthy and fact-checked claims. ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:4 Srba et al. This analysis is complemented with an overview of additional 6 categories of credibility signals. We also thoroughly position such NLP research in the context of other modalities as well as multidisciplinary works. (4) We identify and critically analyse key gaps in the existing literature, highlighting underexplored areas and methodological limitations in the current approaches to credibility assessment. Furthermore, we outline future research challenges and opportunities, with particular emphasis on the rapidly evolving landscape shaped by generative AI, especially Large Language Models (LLMs). We discuss how these advancements both pose new threats, such as the large-scale generation of highly credible yet false content, and offer new capabilities for enhancing the detection and aggregation of credibility signals. Our analysis aims to guide future research toward more comprehensive, robust and explainable credibility assessment approaches. Besides the review of published research works, this survey builds on extensive past research activities and acquired knowledge of its authors in the target area. A unique composition of authors consisting of experts not only on the primary NLP area, but also on other related research domains (such as computer vision, social network analysis) as well as media professionals, provides a novel holistic perspective on the addressed topic. Such diversity contributes, besides others, to identification of open problems and future challenges, in which an application of research in practice plays a prominent role. This survey is structured into 10 sections. Section 2 provides definitions of core concepts related to credibility, describes various dimensions of credibility assessment as well as defines the scope of this survey while situating it into the context of other existing surveys. In Section 3, the methodology employed in the survey process is thoroughly described. Section 4 focuses on research papers that address automatic credibility assessment (i.e., approaches addressing both steps — detection of credibility signals and their aggregation into a credibility label/score). Subsequently, Sections 5-7 address an automatic detection of 3 selected key categories of credibility signals. Section 8 supplements this analysis with a brief overview of additional 6 categories of credibility signals, including the complementary perspective of credibility signals addressed in the non-NLP research. In Section 9, we provide an orthogonal analysis of challenges and open problems that characterize the state of the art in this area. Finally, conclusions are drawn in Section 10. 2 BACKGROUND 2.1 Definitions The construct of credibility has received more scholarly attention than most other communication variables, with foundational research dating back to 1951, when Hovland and Weiss [130] demonstrated that the effectiveness of communication is heavily influenced by the perceived credibility of its source. A systematic meta-analysis of leading international communication journals spanning 1951 to 2011 [ 127 ] examined how source, message, and media credibility have been conceptualized and measured over time. The findings revealed significant inconsistencies across credibility scales, a lack of operational precision, and limited replication and validation of the credibility construct. These discrepancies largely stem from the fact that credibility is a complex, multidimensional concept that lacks a single, unified, and widely accepted definition. Existing definitions vary especially across various disciplines and fields, which study credibility from different perspectives, such as information science, psychology, marketing or human-computer interaction (HCI) [ 266 ]. It is commonly defined as a high-level construct with a help of related concepts, such as believability (which is considered to be roughly a synonym with credibility), trust/trusthworthiness, reliability or their various combinations [ 14 , 259 , 266 , 285 ]. Credibility is ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:5 always tied to a target entity (an object of assessment), that can be either a piece of content or a source spreading such content [ 162 ]. Due to ambiguity of terminology, in this survey we opt for community-contributed definitions [ 126 ] gathered by the Credible Web Community Group 1 . It defines credibility as a degree to which information is credible (believable) and to which information appears non-misleading and useful (for the given audience). Credibility assessment is a process of ascertaining some level/degree of credibility to a target entity. Practically, credibility assessment can be formulated as [ 259 ]: (i) an ordinal classification problem (i.e., assigning a label from pre-defined categories, such as a low/high level of credibility); or as (ii) a regression/scoring problem (i.e., assigning a numerical score). In both cases, there are no standard scales adopted yet. Credible Web Community Group distinguishes two kinds of credibility assessment acts [126]: (1) Making credibility decisions is an act of adhoc deciding for oneself what to believe, which is often informal, immediate, and unconscious. Doing this incorrectly often leads to being misled, although it may not be practical to do it correctly at all times. (2) Credibility analysis is an act of systematic gathering, organizing, and analysing evidence to help people make credibility decisions about a particular information item. A credibility analysis process (performed manually by people or automatically by machines) might produce some kind of report which might itself be called a credibility assessment and might include a credibility score. In the following text, we will focus especially on this act. Consequently, a credibility signal [ 126 ] is a small unit of information used in making a credibility assessment as an evidence. This can be a measurable feature of the information being assessed for credibility, or information about it (metadata), or information about entities which relate to it in various ways, such as the entity who provided it. Credibility indicator [ 126 ] is commonly used interchangeably with credibility signal. In some communities, credibility signal is used for inputs to credibility assessment algorithms and credibility indicator is used for the display features added to the output, to communicate (explain) results to an end user. In this survey, we prefer to use credibility signal, while we would like to emphasize that in practice many signals can be used directly also as indicators communicated to end users. Finally, credibility assessment tools [ 126 ] are software features or applications which perform credibility assessment or help people do so. Such tools can either derive credibility signals, or implement also a credibility assessment algorithm to aggregate such signals into a credibility label/score. Credibility and credibility assessment is inherently very close to false information and the task of false information detection. Nevertheless, credibility (as also defined above) is rather a parallel/complementary concept. While they tend to correlate (low credibility may indicate false content), non-credible content does not necessarily need to be false only and vice versa. We would like to emphasize that credibility signals have been already recognized in false information research as a possible solution towards more accurate and explainable methods. The work by Grieve and Woodfield [105] delve into the complexities of identifying and analysing disinformation deceptive news. The authors study the way language adapts to different contexts and purposes, arguing that the linguistic choices in deceptive news differ systematically from those in genuine news. The difference in intent – to deceive versus to inform – leads to detectable variations in language use. The authors illustrate this with the famous case of fake news involving Jayson Blair of The New York Times, where false news were less informationally dense and less confident than real ones. The pressure to invent news and the lack of factual grounding influenced writing style, making it less precise and authoritative. The authors propose that this type of register 1https://credweb.org/ ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:6 Srba et al. Dimensions of credibility assessment Approaches to credibility assessment Automation-based approaches Human-based approaches Hybrid approaches Levels of credibility assessment Content level User/source level Topic/event level Media level Hybrid level Taxonomies of credibility signals Information Source Content-based signals Context-based signals Polarity Predictors of high credibility Predictors of low credibility Subject Type Text-based subjects: Claim, Text (in general), Article, Title Multimedia-based subjects: Image, Audio, Video Metadata subjects: Web page, Website, Aggregation, Venue, Provider, Creator, Person, Organization Fig. 1. Dimensions of credibility assessment. Bold font highlights such aspects which are within the scope of this survey (see Section 2.4 for more information). Please note that Content level under the dimension of Levels of credibility assessment and Content-based signals under the dimension of Taxonomies of credibility signals refer to two different concepts – the former to the level of analysis, and the latter to the nature of the source data from which the credibility signals are detected. variation could be a broader indicator of fake news. They call for continued research to understand the linguistic patterns of fake news and to develop strategies to combat its spread. Although the above framework is only based on the analysis of one reporter’s articles, this approach may be extended to corpora of non-credible/fake news and credible/real news on various topics to try to identify credibility signals/register variation between the two. To complement this recent case, we can also recall that the idea of analysing the language of disinformation and propaganda has its source in the pioneering work of the German linguist and philologist Viktor Klemperer [ 163 ]. Klemperer showed how hate speech and conspiracy theories were used by the Nazis to justify the persecution of Jews. He also showed how these discourses were repeated and amplified by the media and institutions, until they became the norm. 2.2 Dimensions of credibility assessment In the past research works, a wide range of approaches studying the concept of credibility and credibility assessment was developed. In the following overview (see also Figure 1), we provide various dimensions adopted to study these concepts. Approaches to credibility assessment. First, the credibility assessment methods can be broadly divided into three major groups [14]: (1) A large group of automation-based approaches focuses either on automatically detecting the credibility signals or aggregating them into the credibility indicators. The complexity of such ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:7 automation can be very diverse, ranging from simple ruleand heuristic-based methods, through weighted and information retrieval (IR) approaches to supervised/unsupervised machine learning. This group also covers research works focusing on building datasets appropriate to create such methods/models; as well as systems (web applications, browser extensions, etc.) built on the top of such automated methods. (2) The second group corresponds to human-based approaches, in which various aspects of credibility are studied from the perspective of end-users. The human-based credibility assessment approaches can range from voting methods, cognitive and perception approaches to manual verification approaches. (3) Finally, hybrid approaches combine and utilize the advantages of both the automationand human-based approaches. Levels of credibility assessment. Following a type of a target entity, which is a subject of credibility assessment, we can distinguish five levels of credibility assessment as follows [ 14 , 259 ]: (1) Content level. At the content level, the task is to analyse the content attributes typically of a social media post or a news article. It is the most fundamental and important type of assessment, which is commonly used to determine the credibility at other levels (e.g., if a post is assessed to be credible, also a corresponding user/source or an associated topic/event is considered to be credible). (2) User/source level. This level of credibility assessment depends on features extracted from user accounts (e.g., age, education, profile image) and a history of user-generated content. Besides a specific user, it can relate to a broader source (e.g., a news portal, an organization). (3) Topic/event level. At this level, a credibility is assessed by proceeding from a cluster of posts falling under a specific trending or potentially high-impact topic or event (e.g., elections, societal crisis situations). (4) Media level. At the media level, the medium used to communicate and spread information is a target of credibility analysis (e.g., an online social network). This level typically encompasses credibility analysis of authors, spreaders as well as posts themselves. (5) Hybrid level. To optimally utilize the advantages of individual levels and take advantage of high correlation between them, researchers commonly apply a hybrid approach in which the credibility is assessed at multiple levels at the same time. Taxonomies of credibility signals. Credibility signals, as a core unit in the credibility assessment process, can be categorized by various taxonomies, following different views. First, reflecting the source of information, many works broadly distinguish between contentand context-based credibility signals [ 41 , 59 , 67 , 88 , 93 , 99 , 139 , 145 , 153 , 246 , 258 , 353 , 366 ]. Contentbased signals are derived from the content itself, such as partisanship, emotional appeal, or persuasion techniques. On the other hand, context-based signals, also referred to as provenance- [ 95 ] or meta-information, take into account contextual (e.g., user, source) cues, such as user statistics, location, domain reputation, or the number of shares. In practice, when a content-level credibility (e.g., a credibility of a social media post) is about to be assessed, a combination of both – content-based credibility signals (e.g., a presence of persuasion techniques) as well as context-based credibility signals (e.g., the same claim was already fact-checked) – may be utilized. Second, the credibility signals differ in their polarity. In the existing works, some signals are formulated as predictors of high credibility (e.g., high objectivity), while other signals are formulated as predictors of low credibility (e.g., presence of logical flaws). When interpreting or aggregating individual credibility signals (either by an end user or a credibility assessment algorithm), such polarity must be explicitly taken into consideration (e.g., in [170,251]). ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:8 Srba et al. Third, Credible Web Community Group [ 217 ] structured credibility signals by a subject type (a part of target entity they are derived from): claim, text, image, audio, video, article, title, web page, website, aggregation (e.g., Really Simple Syndication - RSS), venue, provider, creator, person, or organization. Despite many commendable efforts, there is no standardized unified list and categorization of credibility signals. As a result, research works commonly utilize various and very diverse lists of credibility signals. For example, Molina et al . [210] proposed separate lists of credibility signals (or features as they have been denoted by the authors) to distinguish between eight types of online content, such as real news, fabricated news, or satire. Each list grouped signals into 4 main categories roughly corresponding to levels of credibility assessment: (i) message and linguistic, (ii) sources and intentions, (ii) structural, and (iv) network. An alternative approach was used by the Credibility Coalition 2 , which organized weekly remote sessions during which participants drafted about 100 credibility signals specifically aimed at the credibility of web pages [ 366 ]. They have been later coalesced into 12 major categories, including reader behaviour, revenue models, publication metadata, and inbound and outbound references. In 2017, the Credibility Coalition formed the World Wide Web Consortium (W3C) Credible Web Community Group, which continued in this initiative. It crowdsourced from human experts (researchers as well practitioners) an extensive list containing more than 200 credibility signals, commonly referred to as W3C signals [ 217 ]. This list introduces a number of signals categorized by subject types. While such a list still remains an informal and incomplete draft and has not been standardized yet, it has been adopted and served as a foundation to several research works (e.g., [246,251]). We can conclude that no standard or widely used and recognized list or categorization of credibility signals currently exists that could serve as a solid basis for our literature survey. Categorization by means of subject types as proposed by W3C signals [ 217 ] is probably the most comprehensive and elaborated one, nevertheless, it mixes at the top level different modalities (text, video, audio, etc.) and content types (article, title, claim etc.). More specifically, the signals relevant to our survey are not optimally classified under the text modality or across various individual content types. Therefore, we thoroughly analysed the relevant existing works [ 210 , 217 , 218 , 251 , 366 ] and proposed a novel unified categorization of textual credibility signals (see Figure 2). We adopted a bottom-up approach as follows: (1) First, out of mentioned credibility signals, we identified specifically textual credibility signals (i.e., signals that can be derived from textual content by means of diverse NLP methods). (2) We grouped the resulting credibility signals according to their semantic similarity and shared datasets and methods applied for their detection. To this end, we employed three grouping strategies. At first, we considered the existing categorizations that have been proposed by the relevant works, especially individual (sub)categories from the above-mentioned W3C signals [ 217 ]. In this way, we grouped all signals related to a title into a single category named Clickbaits and title representativeness, all signals related to claims into a category of Checkworthy and fact-checked claims as well as reused a subcategory of References and citations. Second, we grouped together signals which refer to the same or highly-similar concepts and just use a different terminology, e.g., grammar and spelling errors, grammar, spelling or punctuation mistakes, punctuation naturally fall under the category of Text quality; or profanity, incivility, impoliteness and hate speech under the category of Offensive language. In some cases, the signals refer to the same concept, but in an opposite meaning, like originality and attribution of non-original content belong to the same category of Originality and content reuse. In some cases, the signals refer to the same concepts but are defined on different 2https://credibilitycoalition.org/ ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:9 Textual Credibility Signals Factuality, subjectivity and bias Factual accuracy [210,218], Event factuality [54,175,272], Subjectivity [251], Selection and presentation bias [218], Hyperpartisanship / political bias [217], Confirmation bias [217], Framing [218], Emotional tone [366], Formal / Informal tone [217], Emotional, sensational or affective language [251], Emotionally charged [210,217], Emotional Valence [217] Use of hyperboles [210], Impartial reporting [210], One-sided reporting [210], Discrepancies or omissions [210] Persuasion techniques and logical fallacies Logical fallacies [366], Logical flaws [210], Inference [366], Ad-hominem attacks [210], Conspirational reasoning [210], Arguments from authority [210], Common man appeals [210], Call to Action (Political) [217] Check-worthy and fact-checked claims Fact-checked [210,366], Fact-check status of claim [217], Article has a central claim [217] Text quality Grammar and spelling errors [251], Grammar, spelling or punctuation mistakes [210,217], Readability [251], Punctuation [251], Vocabulary or reading level [217] References and citations Quotes from outside experts [366], Citation of organizations and studies [366], Representative citations [366], External links [251], No quotes or made-up quotes [210], No source attribution [210,217], Uses standardized references or citations [217], Few to zero references or citations [217] Clickbaits and title representativeness Clickbait title [217,251,366], Misleading and clickbait headlines [210], Title representativeness [217,366] Originality and content reuse Originality [366], Originality types [217], Attribution of Non-Original Content [217] Offensive language Profanity [251], Incivility and impoliteness [217], Hate speech [217] Machine-generated text No appearances in the existing lists/categorizations yet Fig. 2. Unified categorization of textual credibility signals. Light blue nodes consist of credibility signals depicted in the existing taxonomies and categorizations. Bold font highlights categories of credibility signals on which this survey put a focus on (see Section 2.4 for more information). granularity levels, e.g., logical fallacies, logical flaws, ad-hominem attacks all together belong under Persuasion techniques and logical fallacies. Third, we grouped together signals where the borders between them are not clearly defined and in practice they correlate with each other, e.g., subjectivity, political bias, confirmation bias, emotional tone were grouped under Factuality, subjectivity and bias. We also jointly combined these strategies, especially the category of Factuality, subjectivity and bias groups signals that are closely related, commonly refer to the same concepts but from opposite angles, or are at different levels of abstractions. (3) In this way, we finally proposed 8 categories of credibility signals (see also Figure 2; more detailed description of each category is provided in the later sections): (i) Factuality, subjectivity and bias; (ii) Persuasion techniques and logical fallacies; (iii) Check-worthy and fact-checked claims; (iv) Text quality; (v) References and citations; (vi) Clickbaits and title representativeness; (vii) Originality and content reuse; and (viii) Offensive language. (4) In addition, we recognized one more category of signals – (ix) Machine-generated text – that has become very relevant only recently with the emergence of LLMs and, therefore, it did not appear in the existing lists/taxonomies. At the same time, it is a strong predictor of credibility as fully generated text can contain many factual errors either as a result of unintended hallucinations of LLMs or due to their intended misuse to generate false content (the previous works demonstrated significant vulnerabilities of LLMs to generate disinformation [ 328 ], even personalized one [372]). Importance and effects of credibility signals on end users. Due to a broad definition of credibility, as well as rich taxonomies of credibility signals, a number of human-based studies have analysed the importance of individual credibility signals and their effects on the perceived credibility, in general as well as for specific groups of end users (according to their expertise, age, education etc.). ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:16 Srba et al. •year, when the paper was published in, • main outcome (one of the following options: a method/model, a study, a dataset, a tool, a survey). Moreover, we annotated several additional types of information that have been identified as the most useful for the analysis and description of the current state of the art, as well as for recognition of remaining open problems and potential future work, namely: • dataset used (if authors used an existing dataset, a name of such dataset, otherwise information about how own dataset was obtained), • content/context signals (whether the paper tackles with content, context or both types of credibility signals, see Section 2.2 for more information), •list of credibility signals (a specific list of signals addressed in the paper), • automated detection (whether a paper attempts to perform also an automatic credibility assessment and/or automatic detection of credibility signals), • models (if paper attempts to perform automated detection, what kind of models have been used), • metrics (if paper attempts to perform automated detection, what kind of metrics have been used), • human agreement (if paper introduces an annotated dataset, what kind of annotator mutual agreement was achieved, measured by standard metrics, e.g., by Cohen’s Kappa). The full list of papers included in this survey, including all annotations is available as a supplementary material to this article in the ACM Digital Library8. Following this analysis, we describe credibility assessment papers as well as each individual category of credibility signals from three perspectives (that are also utilized to structure the consequent categories): (i) datasets, (ii) methods and models, and (iii) tools and services. For each category we summarised the current state of the art and identified problems and challenges that are specifically present in such credibility signal category (such analysis is further elaborated in Section 9, with an orthogonal discussion across all included papers). 4 AUTOMATIC CREDIBILITY ASSESSMENT This section provides an overview of research efforts that address automatic credibility assessment, encompassing both steps of the process: the detection and subsequent aggregation of credibility signals. In some cases, however, the features used as input to the credibility assessment algorithms are relatively shallow, relying primarily on surface-level linguistic characteristics (e.g., the number of unique words), which only approximate underlying credibility signals rather than explicitly capturing them. This group of approaches explicitly uses the terminology relevant to credibility (such as credibility assessment, credibility signal, credibility score), in contrast to approaches described in the following sections that address detection of a particular category of credibility signals, but typically do not denote them as such. 4.1 Datasets Most of datasets used in the credibility assessment works are annotated at the content level (see Section 2.2 for the overview of levels of credibility assessment). Table 2provides an overview of the selected datasets created or adapted for the purpose of credibility assessment task. Zhang et al. [366] created a thoroughly annotated, however, only very small credibility-related dataset. Out of news articles, being the most shared on social media, 40 articles were selected 8 A full list of papers is available as supplementary material to this article in the ACM Digital Library and at https: //kinit.sk/public/acm-tist-credibility-assessment-survey.html ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:17 (covering multiple topics: public health, climate science, diseases, vaccines). In total, 6 annotators were recruited to annotate the credibility signals. These indicators were adopted from the initial list of signals created by the Credibility Coalition (see Section 2.2). Out of them, 16 content and context indicators were selected to be annotated. Namely, content indicators included: title representativeness, clickbait title, quotes from outside experts, citation of organizations and studies, calibration of confidence, logical fallacies, tone, inference. Context indicators included: originality, fact-checked, representative citations, reputation of citations, number of ads, spammy ads, number of social calls, and finally placement of ads and social calls. Besides that, all articles were evaluated for an overall credibility by domain experts on a 5-point scale. In the following dataset analysis, two backward stepwise multiple regression models were utilized to measure the predictive value provided by both sets of indicators. For content-based signals, after model convergence, two variables remained: clickbait title and logical fallacies (slippery slope). This model was found to significantly predict credibility (F = 13.972, p < 0.001). For context-based signals, 6 variables remained: fact-checked–reported false, fact-checked–reported mixed results, number of social calls, number of mailing list calls, and placement of ads and social calls. Together, they were also found to significantly predict credibility (F = 12.986, p < 0.001). Despite a wide range of annotated credibility signals as well as overall credibility score, the practical value of the resulting dataset, which was created only as a proof of concept, is somehow limited by its size (N = 40) which is insufficient even in few-shot learning scenarios. In the follow-up study, Bhuiyan et al . [41] gathered a dataset of over 4,000 credibility assessments taken from 2 crowd groups (journalism students and Upwork workers) and 2 expert groups (journalists and scientists). Similarly as in the previous case, a varied set of 50 news articles related to climate science were selected. News articles were annotated by both crowd and expert groups on a 5-point Likert scale. The in-depth analysis of the obtained annotations revealed differences in annotation not only between crowd and expert groups, but also within expert groups between journalists and scientists due to differing expert criteria that journalism versus science experts use. Following the observations, authors proposed directions how to better design crowdsourcing of content credibility. El Ballouli et al . [88] collected 17 million tweets in a period of two weeks that were written in Arabic and contained at least one hashtag. After preprocessing and grouping tweets by their hashtags, a topic-independent subset of 9,000 tweets was selected. Each of these tweets was annotated by three annotators on a binary scale (credible, non-credible), while the majority vote was used to determine the final label. Due to a lack of sufficiently large credibility-annotated datasets, some authors (e.g., [ 170 , 251 , 353 ]) decided to use available false information (fake news) datasets, such as LIAR [ 333 ], Weibo [ 353 ], FakeNewsNet [ 295 ], FakeNewsAMT [ 244 ] or Celebrity dataset [ 244 ]. In such a case, various false information labels are mapped to credibility labels by taking an assumption that a false content is considered to be non-credible. Qureshi and Malick [258] also proceeded from the FakeNewsNet dataset, however, only to identify a potential set of tweets that were further manually annotated by 12 experts for credibility on a 5-point scale. The resulting dataset consists of 4,958 tweets, out of them some were discarded due to low annotators’ agreement. Another group of datasets used in the existing works relates to credibility annotated at a source (webpage) level. Microsoft Credibility [ 284 ] dataset consists of top 40 search results on 25 predefined queries on the topics of Health, Politics, Finance, Environmental Science, and Celebrity News, which were manually annotated for credibility on a 5-point scale. Another dataset, Content Credibility Corpus (C3) [ 152 ], consists of 15,750 evaluations of 5,543 pages by more than 2,000 annotators recruited through Amazon Mechanical Turk. The textual credibility signals annotated in this dataset include readability, language quality, informativeness, completeness, and objectivity. ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:18 Srba et al. Finally, at the topic/event level, Mitra and Gilbert [205] introduced a large scale dataset containing 60M tweets collected during a period of more than three months and grouped into 1,049 real-world events, each annotated by 30 human annotators on a 5-point scale. Table 2. Selected datasets used in the works addressing automatic credibility assessment Dataset Language(s) # Instances Content type Classes Used by Zhang et al. [366] English 40 News articles 8 content-based signals, 8 context-based signals, overall credibility on a 5-point scale [251,366] Bhuiyan et al. [41] English 50 News articles overall credibility on a 5-point scale [41] El Ballouli et al. [88] Arabic 9,000 Tweets credible, non-credible [88] LIAR [333] English 12,800 Short statements from politifact.com pants on fire, false barely-true, half-true mostly-true, true [353] Weibo [353] Chinese 18,000 Microblogs same 6 as in LIAR [353] FakeNewsNet [295] English 23,196 News articles, microblogs true news, false news [170,251] FakeNewsAMT [244] English 480 Political news articles fake, legitimate [170] Celebrity [244] English 500 Celebrity news articles fake, legitimate [170] Qureshi and Malick [258] English 4,958 Microblogs overall credibility on a 5-point scale [258] Microsoft Credibility [284] English 1,000 Websites (search results) overall credibility on a 5-point scale [93,251] Content Credibility Corpus (C3) [152] English 5,691 Websites 25 signals grouped into 6 categories [93,251] CREDBANK [205] English 60M Microblogs credibility of 1,049 real-world events annotated on a 5-point scale 4.2 Methods and models The methods and models used in the research works falling into this group typically employ a methodology consisting of three steps. Firstly, a set of credibility signals is selected. Secondly, various techniques are employed for their automatic detection. Finally, detected credibility signals are then used to predict an overall credibility which can be subsequently utilized for example as an input to information retrieval engine. In the following description, we provide an in-depth description of each of these three steps. Credibility signals selection. In the earliest works, credibility signals were rather approximated by various (mostly shallow) linguistic features that can be easily detected by the machine, such as a number of words or exclamation marks. While being potentially helpful for traditional machine learning algorithms, people would not usually use such features for credibility assessment due to the absence of their direct interpretability. Weerkamp and De Rijke [338] adopted a subset of 11 such features for blog posts proposed in [ 273 ] while focusing on those that are text-based and can be easily detected automatically, such as capitalization,shouting,spelling or post length. Similarly, Kang et al . [153] opted for 12 numeric and 7 binary content-based credibility indicators for Twitter, such as a positive sentiment factor or a number of mentions. Approximately half of these indicators were adopted from the previous work by Castillo et al. [55]. El Ballouli et al . [88] selected 48 credibility signals, out of them 26 signals were content-based. Besides shallow linguistic features, as already included in the previous works (e.g., count of hashtags, count of unique words,count of exclamation/question marks) the authors utilized also more advanced sentiment extraction to assess positive sentiment,negative sentiment, and objectivity. A shift towards more complex credibility signals is also present in the work by Qureshi and Malick [258] . Authors initially complied a list of 33 potential credibility signals. By proceeding from the initial experiments, a final set of 11 contentand context-based signals were used. In the case of content-based signals, shallow features (such as a number of hashtags,length of text) were discarded in favor of more complex signals like, post hate,informativeness, or deception. For the purpose of website credibility evaluation, Esteves et al . [93] selected 15 content-based indicators by proceeding from the taxonomy created by Olteanu et al . [224] , such as authority ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:19 (authoritative keywords within the page HTML content), readability metrics,text category or sentiment. In contrast to works mentioned so far, the following works selected signals from the list created by the Credible Web Community Group (see Section 2.2). At first, Podgurski et al . [251] proposed a modular credibility assessment system of web pages using the 23 most frequent contentand context-based signals. The employed signals range from naive numerical and linguistic features, such as a number of external links,linguistic statistics (average text and word count, overall word count, etc.), spelling and grammar,domain endings and author information, to more complex features, such as sentiment, readability and presence of a clickbait title. Similarly, Leite et al . [170] employed 18 credibility signals, including advanced ones that have not been considered in the previous works, such as call to action,impoliteness,sensationalism or explicitly unverified claims. Credibility signal detection. Early shallow linguistic features were detected primarily by simple ruleand heuristic-based techniques [ 153 , 338 ], such as a predefined lexicon of positive and negative words, counting/searching for specific characters/patterns (e.g., capital letters, emoticons). By proceeding from such superficial linguistic features to more advanced credibility signals, also their detection methods became more complex. For sentiment analysis various lexicon and rule-based sentiment analysis were employed, such as Vader for English [ 93 , 251 ], or ArSenL for Arabic [ 88 ]. For text category classification, well-known approaches employed in other NLP tasks were adopted, such binary multinomial Naïve Bayes (NB) classifiers or Latent Semantic Analysis (LSA) [ 93 ]. Various NLP libraries were used for automatic detection of additional credibility signals, such as LanguageTool library for detection of grammar & spelling errors or TextBlob for subjectivity detection [ 251 ]. Even further, Qureshi and Malick [258] detected the selected credibility signals by means of the classification methods proposed in the previous research works. Surprisingly, adoption of LLMs for detection of credibility signals is only very rare despite a great potential of LLMs to target a wide range of various credibility signals categories. To this end, Leite et al . [170] used 3 instruction-tuned LLMs (GPT-3.5-Turbo, Alpaca-LoRA-30B and OpenAssistantLLaMa-30B) with a specific prompt designed for each credibility signals (e.g., “Does the article make use of sensationalist claims?”). Credibility label/score prediction. To estimate overall credibility from the credibility signals, Weerkamp and De Rijke [338] opted for simple weighting schemata. The resulting credibility score was subsequently incorporated into a blog post retrieval model. Kang et al . [153] proposed three computational models (based on a weighted combination, and on a probabilistic language-based approach) for assessing tweet’s credibility, using not only content-based but also social and hybrid strategies. Experiments on the test set consisting of 1,023 instances (while using 10-fold crossvalidation) revealed that social model was able to outperform content-based and hybrid model. This can be explained by a short textual content of tweets, as well as very shallow linguistic and possibly noisy credibility signals, such as a presence of exclamation mark. Podgurski et al . [251] combined individual signals’ subscores with signal weights through a linear combination function. Signal weights were determined by authors empirically by considering prior research results, signal measurement accuracy, and experimental calculations on test data. By advancing from simple weighting and heuristics, consequent works started to employ feature engineering with traditional machine learning approaches. El Ballouli et al . [88] trained a random forest decision tree classifier on the top of detected credibility signals – each serving as a feature. To predict website credibility score, Esteves et al . [93] used Gradient Boosting and AdaBoost classification algorithms on the top of Microsoft Credibility and Content Credibility Corpus (C3). For the purpose of feature selection, Qureshi and Malick [258] utilized the Kolmogorov–Smirnov (KS) test to identify discriminating features. Subsequently, a wide range of machine learning algorithms (12 regression and 10 classification ones) were used to predict original 5-scale as well as ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:20 Srba et al. simplified binary credibility label. Results revealed that context-bases signals, such as high user influence and medium-to-high spread count, are very indicative of credible tweets. Credible tweets are normally spread by topic experts (the topic of the tweet is in the top 3 topics for that user), typically do not contain deceptive information and are not posted by malicious accounts. In terms of content-based signals, credible tweets have high informativeness and tent to not contain hate speech or links to low-credible sources. Analogically to other ML/NLP areas, also credibility assessment underwent a shift towards deep learning methods. In this direction, Wu et al . [353] introduced an ANSP model based on adversarial networks and multi-task learning to capture differential credibility features for information credibility evaluation. The input to the model consists of the concatenations of word embeddings (Word2Vec) and meta-data embeddings (various credibility signals provided by the datasets). Evaluation on English and Chinese datasets, LIAR and Weibo respectively, showed that the combination of all features performs best for both datasets, while content-based features were consistently outperformed by contextual ones. Among these, some signals are particularly powerful on their own, such as the speaker, state info, party affiliation, or credit history. Finally, Leite et al . [170] used Prompted Weak Supervision (PWS) on the top of LLM-predicted credibility signals. This approach was compared with unsupervised and supervised baselines. In the supervised finetuning scenario, a RoBERTa-Base model was finetuned with the ground-truth labels. In the zero-shot scenario, an instruction-tuned LLM based on LLaMa2 was prompted without any signals or in-context examples. Experimental results on four datasets (FakeNewsNet - GossipCop, FakeNewsNet - PolitiFact, FakeNewsAMT, and Celebrity) showed that Prompted Weak Supervision outperformed the zero-shot baseline by 38 . 3%, and achieved 86 . 7% of the performance of the state-of-the-art supervised baseline. Furthermore, in cross-domain settings, where the domain of the train set differs from the domain of the test set (e.g., Politics and Gossip), Prompted Weak Supervision outperformed the supervised baseline by 63%. Lastly, the authors study the association between credibility signals and veracity through (i) a statistical test, where 12 out of the 19 signals were shown to have an association with veracity, and (ii) an ablation study, in which signals were individually removed from the train set in order to inspect their contribution to the model’s performance. This showed that the method’s strength is in the combination of a wide range of signals rather than relying on a small set of signals that could be strong predictors of veracity. Through these analyses, the authors verified that some credibility signals are domain-specific. For example, signals such as Source Credibility and Misleading about Content improve performance mostly for the political domain, while others such as Expert Citation and Call to Action show benefits in entertainment news. In addition to the efficiency of this approach in terms of being independent of the costly annotation of long news articles with each signal, the main advantage of this method is generalisation capability. The authors found that their approach outperforms the fine-tuned methods on the out-of-domain test data. This result is particularly important in the constantly-evolving online environment, when new non-credible and false information emerge every day, making it impossible to have up-to-date annotated data. 4.3 Tools and services While several services dedicated to credibility assessment exist, many of them are on a border with false information detection. From 2019, the Credibility Coalition maintain a CredCatalog 9 – a catalogue of initiatives that have a stated aim to improve information quality. Besides various organizations (like fact-checking or academic institutions), it provides overview of relevant tools. 9https://credibilitycoalition.org/credcatalog/ ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:21 Besides that, there are commercial projects supporting end-users in evaluating the credibility of the online content directly in their browsers. NewsGuard 10 provides News Reliability Ratings for news outlets based on nonpartisan journalistic criteria. Similarly, GroundNews 11 rate news stories according to their bias distribution (left-/right-leaning bias) or factuality (in the meaning of source reliability and general factual accuracy). The labelling is, however, done rather indirectly through sources writing about such news stories, instead of analysing a content of individual news articles themselves. Another relevant tool is the Tanbih [ 369 ] – a news aggregator that performs media-level and article-level analyses, aiming to help users better understand the content they consume. It organizes news articles into event-based clusters and generates media outlet profiles, which include indicators such as the general factuality (in the meaning of factual accuracy) of reporting, the degree of propagandistic content, hyper-partisanship, leading political ideology, overall framing, and the stance of the outlet toward various claims and topics. Finally, an automatic analysis of several credibility signals (persuasion techniques, subjectivity, or machine-generated text among others) is available as a part of the Assistant tool in the wellknown Verification plugin 12 developed and further enhanced as a part of EU-funded projects (InVID, WeVerify and vera.ai). The Assistant tool allows users to provide a URL or a local file as input, from which the text is extracted and various text analysis are performed, credibility signals being a recent addition to them. 4.4 Discussion Limited research focus with absence of LLM-based solutions. As already shown in Figure 4, the number of works explicitly addressing automatic credibility assessment is considerably lower in comparison with other individual signal categories. This is in contrast to false information detection, which attracted a plethora of research (and also practitioners’ and public) attention. In parallel, our literature survey pointed out a lack of approaches utilizing LLMs. At the same time, LLMs provide a great opportunity to automatize detection of multiple credibility signals even in a zero-shot settings, as the work by Leite et al . [170] clearly demonstrated. Even further improvement in the terms of accuracy can be achieved by employing in-context learning, or instruction-tuning LLMs. While such LLM-based approaches would inquire higher computational costs, existing Parameter-efficient Fine Tuning (PEFT) techniques [355] may be employed if necessary. Lack of multilingual datasets with annotated credibility signals. Similarly to a limited research focus, also the situation with credibility-annotated datasets is falling behind false information research. There are very few datasets with manual (human) annotations at the level of individual credibility signals as well as overall credibility [ 41 , 366 ]. Unfortunately, these datasets are very small, which prevents the use of language models even in few-shot settings. Despite the declared original intentions to extend them in future, to the best of our knowledge, no follow-up datasets have been created so far. The remaining datasets are larger (containing hundreds or thousands of instances), nevertheless, annotated only at the level of overall credibility (see Table 2). Absence of datasets providing annotations of multiple categories of credibility signals for the same set of instances, prevents multi-task training of models. Multi-task approaches hold considerable potential, particularly because many credibility signals inherently share underlying similarities that can enhance model performance, as already demonstrated by the results of Prompted Weak Supervision (PWS) in [ 170 ]. 10https://www.newsguardtech.com/ 11https://ground.news/ 12https://www.veraai.eu/category/verification-plugin ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:22 Srba et al. Robustness of credibility signals. For reliable credibility assessment, it is necessary that credibility signals and their relation to overall credibility remains the same in time, or at least is not influenced by significant domain and data drifts. To shed more light on the robustness of credibility signals, Ji et al . [145] evaluated how the signals and their prediction capability change over time. By utilizing the posts from the Weibo platform, two topic-specific datasets were created, one for climate change and second one for genetically modified organisms (GMO). The authors analysed how credibility signals evolve over time by splitting these datasets year-wise. They found that certain signals remain more stable over time, however, this stability is topic-dependent. For example, post sentiment performed better for the climate change dataset in most of the years compared with other features, while the post topic was the best predictor for the GMO dataset. In general, results showed that content-based features remained effective across time. On the other hand, contextual features, such as activeness and gender, were not correlated with veracity in the climate change dataset, while popularity showed no significant correlation with GMO misinformation in any given year from 2010 to 2020. In summary, content-based signals (especially sentiment and topic) seem to be more robust and less prone to the temporal data drift compared to context-based signals. 5 FACTUALITY, SUBJECTIVITY AND BIAS Within false information research, the term factuality has been used to refer to a variety of related problems, with no unified ontology to define its levels. For example, Nakov et al . [218] equate factuality with factual accuracy (i.e., veracity) and propose a four-level ontology encompassing the claim, article, user, and source factuality levels. In this interpretation, factuality is seen not as as a credibility signal, but as the final verdict regarding truthfulness of information. In the context of large language models (LLMs), factuality is typically defined as a model’s ability to generate outputs grounded in established real-world facts [ 331 ]. This aligns with the Cambridge English Dictionary’s definition of factual as “based on or containing facts” 13 . From an NLP perspective, however, factuality is often approached differently. Rather than relying solely on external fact-checking, NLP literature tends to treat factuality as a contextual signal – reflecting the degree of certainty a speaker conveys about the occurrence of events [ 54 , 174 , 175 , 255 , 272 ]. As identified through the systematic literature review conducted as part of this work, factuality in this view functions as a credibility signal rather than an overall veracity assessment, expressing how confidently an event is presented as having occurred or not occurred. This more fine-grained perspective follows earlier work by Saurí and Pustejovsky [279] , which defines factuality as the extent to which events are portrayed as corresponding to real-world situations, hypothetical scenarios, or uncertain interpretations. Under this definition, the task of factuality detection can be used as the first step towards identifying check-worthy claims (discussed in Section 7), as only the events that the speaker is confident about typically form claims. In the rest of this section, we use the terms factuality and event factuality interchangeably. Where the meaning is not obvious from the context or could lead to ambiguity, we distinguish whether event factuality, factual accuracy or a different notion of factuality is meant there. The literature on event factuality traditionally distinguishes between individual event factuality detection (EFD) [ 175 ] and document-level event factuality identification (DEFI) [ 54 ] tasks. The main distinction between these two sub-tasks is in the scope of information concerned – while the EFD task considers individual event mentions, DEFI aims to aggregate several mentions of a certain event within a document. Subjectivity represents an indicator of the overall subjectivity (or objectivity) of information. The task of identifying subjective information is modelled as either a binary or a more fine-grained 13https://dictionary.cambridge.org/dictionary/english/factual ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:23 Table 3. Selected datasets used for event factuality detection. Dataset Language(s) # Instances Content type Classes Qian et al. [255] English, Chinese 1,948 (English) 4,649 (Chinese) News articles from China Daily, Sina Bilingual News, and Sina News. Sentence-level event factuality CT-: negated events PS+: speculative events PS-: speculative negative events Uu: events appear in question CT+: factual events Qian et al. [256] English, Chinese 1,948 (English) 4,649 (Chinese) Extension of Qian et al. [255] with individual and document-level event annotations CT-: negated events PS+: speculative events PS-: speculative negative events Uu: events appear in question CT+: factual events Li et al. [174] English 112,276 events Documents and events CT-: negated events PS+: speculative events PS-: speculative negative events Uu: events appear in question CT+: factual events problem that distinguishes between various degrees of subjectivity. Over the past years, there have been many shared tasks that focus on specific types of subjectivity. For example, Piskorski et al . [249] discriminate between objective reporting, opinionated news and satire. Derczynski et al . [83] and Gorrell et al . [103] organised the challenges for rumour detection. Irony [ 83 ] and sarcasm [ 103 ] is another common type of non-factual information that received close attention during the last years, with shared tasks and challenges dedicated to this problem. Finally, bias represents a similar but more complex phenomenon of imbalance in terms of opinions or facts. As highlighted by [ 218 ], there is no one single concept of bias among scholars. However, in many cases bias is seen as a systematic favouring of a certain ideology when covering information [ 329 ]. This can be expressed in deliberately withholding certain part of information that contradict the favourable point of view [ 302 ] or vice versa, specifically searching for the facts that cover information from a certain political or ideological viewpoint [ 125 ]. In some cases, even if the choice of information sources is unbiased, the presentation of information in those sources may be performed in a biased manner that highlights the importance of only certain facts. Some instances of biased presentation include framing [ 92 , 249 ], opinionated reporting style [ 249 , 305 ] and even visual clues [32] used to distort the perception of information. 5.1 Datasets Tables 3,4, and 5represent a summary of the papers that introduce datasets annotated for detecting event factuality, subjectivity, and bias respectively. As can be seen, the prevalent majority of the datasets are only available in English. Besides English, there are datasets in Urdu [ 216 ], Arabic [ 213 ] and German [ 8 , 17 ]. Furthermore, multilingual subjectivity detection was a part of shared tasks at CheckThat! Labs of CLEF 2023 [ 33 ] and CLEF 2024 [ 34 ], covering Arabic, German, English, Italian, and Turkish languages. The ontology proposed by Qian et al . [255] is the most widely accepted classification for detecting event factuality, with subsequent extension of the dataset with document-level event factuality annotations [ 256 ]. More recently, Li et al . [174] built the first large-scale annotated dataset based on this classification, by using LLM predictions with human judgments as a final step. MPQA opinion corpus [ 344 ] is a particularly widespread benchmarking dataset for subjectivity detection, appearing in the majority of studies covered by the systematic review [ 43 , 179 , 332 , 345 ]. The dataset is manually annotated with three frames, objective,expressive subjective and direct subjective. The dataset contains span-level annotations along with the annotation of the source ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:24 Srba et al. Table 4. Selected datasets used for subjectivity detection. Dataset Language(s) # Instances Content type Classes Piskorski et al. [249] English, German, French, Italian, Russian, Polish 1,592 news articles News articles Objective, Opinionated, Satire Spinde et al. [310] English 2,800 news articles 175,807 comments and retweets referring to these articles News articles, tweets Hateful vs Neutral Biyani et al. [43] English 700 Threads from two popular online forums, Trip Advisor–New York and Ubuntu. Subjective vs Non-subjective Wiebe et al. [344] English 10,657 sentences (535 documents) 187 different news sources Objective Expressive subjective Direct subjective Spinde et al. [309] English 3,700 Sentences collected from news organizations with different political leaning. Opinionated, Factual, or Mixed Banea et al. [28] English, Arabic, French, German, Romanian, Spanish 9,700 in each language This is an extension of MPQA dataset [342] created by translating sentence-level data into other languages. All information is parallel. Objective vs Subjective Atalla et al. [17] German 6,848 Sentence-level annotation of news articles. Created to be compatible with MPQA. Objective vs Subjective Wiebe and Riloff [342] Urdu 500 articles: 700 sentences annotated with emotion and 4,000 unbiased sentences News articles from BBC Urdu Objective vs Subjective Mourad and Darwish [213] Arabic 2,300 Tweets published 2012, randomly sampled Neutral, Positive, Negative, Both, Sarcastic Jeronimo et al. [144] Portuguese 450 words Discourse markers for subjectivity in Portuguese Argumentation, Presupposition Modalization, Sentiment Valuation Maks and Vossen [199] Dutch 11,000–56,000 tokens Lexicon for subjectivity in Dutch based on Wikipedia articles and user comments Actor and Speaker/Writer subjectivity Pang and Lee [230] English 1,000 Movie reviews Subjective vs Non-subjective Table 5. Selected datasets used for bias detection. Dataset Language(s) # Instances Content type Classes Piskorski et al. [249] English, German, French, Italian, Russian, Polish 1,592 news articles News articles 14 presentation frames Spinde et al. [312] English 1,700 Short statements Biased vs non-biased Spinde et al. [309] English 3,700 Sentences collected from news organizations with different political leaning Biased vs non-biased Word-level bias annotation. Liu et al. [185] English 2,990 US news articles from 2018 annotated in terms of frames based on headlines only 4 general frames: Politics; Public opinion; Society/Culture Economic consequences 5 issue-specific frames: Race/Ethnicity 2nd Amendment (Gun Rights); Gun control; Mental health; School/Public space safety Fan et al. [94] English 300 news articles News articles from FOX, NYT and HPO Informational and Lexical Bias (Sentence, token level) Chen et al. [64] English 6,964 news articles Articles from 41 publishers Labels derived from AllSides Bias Detection Unfairness Detection (different levels of text granularity) Aksenov et al. [8] German 47,362 news articles Articles from 36 publishers Fine-grained Bias Detection (different levels of text granularity) ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:25 (author, specific person, etc.) who expresses the subjective frame and the intensity of subjectivity. Direct subjective expressions are typically more explicit than expressive subjective. Originally created in English for detecting a phrase-level subjectivity, it is widely used in multilingual tasks by utilising parallel translations into other languages, such as Arabic, French, German, Romanian and Spanish [ 28 , 208 ]. In addition, some of the datasets adopted the ontology of MPQA to create comparable corpora in other languages [17]. In addition to the datasets for bias detection provided in Table 5, Wessel et al . [341] introduced Media Bias Identification Benchmark (MBIB) collection, which is the most comprehensive set of the benchmark corpora for media bias detection consisting of 22 datasets. The types of biases covered include hate speech,lexical,contextual,linguistic,gender,cognitive,racial and political biases. 3 out of 22 datasets, however, represent a more high-level annotation of information into fake and trustworthy rather than biases. 5.2 Methods and models The methods discussed in this section can be categorized into three groups – event factuality, subjectivity and bias/framing detection techniques. The methods in the first group can be further divided into two subcategories: (i) those targeting the factuality of individual event mentions and (ii) those assessing the document-level factuality of events. Event factuality detection (EFD). The methods falling into this category aim to identify the author’s level of certainty regarding the possibility of the individual mentioned event [ 53 , 169 , 175 , 254 , 360 ]. The most common architectures for event factuality detection are LSTM and bi-LSTM models trained on BERT representations [ 53 , 175 , 360 ]. A few studies also employ traditional machine learning approaches, such as SVM with LOSSO regression [ 169 ] and a combination of rule-based and maximum entropy methods [ 254 ]. All the reviewed methods for event factuality detection are trained on either English [ 169 , 254 ] or Chinese [ 175 , 360 ], or a combination of both languages [53]. Document-level event factuality identification (DEFI). Sentence-level event factuality often results in conflicts within a document, as different mentions of the same event may reflect varying degrees of factuality. Therefore, the methods in this sub-group aim to conclude the overall event factuality in a document based on the various sentence-level factuality values within that document. Cao et al . [54] propose an Uncertain Local-to-Global Network (ULGN) that makes use of two important characteristics of event factuality, local uncertainty and global structure. Similarly, Qian et al . [255] address the challenge of multiple event factuality values within a document by employing an LSTM model trained with both intraand inter-sequence attention mechanisms to assess document-level event factuality. The model incorporates two types of input features, syntactic and semantic. Syntactic features are based on dependency paths from negative or speculative references to the event, while semantic features are derived from the sentences containing the event. Another approach to estimating document-level factuality is the Sentence-to-Document Inference Network (SDIN) proposed by Zhang et al . [368] . This architecture features a multilayer interaction network that aggregates individual event mentions into a global prediction. The last step employs gated aggregation that uses a sigmoid function to generate a mask vector that captures the most critical semantic and factual features of the event mentions. The training process applies a multi-task learning approach, where individual and document-level event factuality prediction tasks share the same pretrained model and interaction network. As an input, the model receives all the event mentions encoded into BERT representations at the final hidden state of [CLS] token. This approach makes it possible to significantly outperform the models by Cao et al . [54] and Zhang et al . [368] described above on Chinese and English DLEF corpora [ 255 ]. Qian et al . [256] propose an approach called Document-level Event Factuality identification via Machine Reading ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:32 Srba et al. Macagno [195] introduced the first multilingual dataset for detection of persuasion techniques, containing tweets in English, Italian, and Portuguese. However, the dataset is small ( < 3 , 000 instances in total) for the purpose of training deep learning models, and similarly to [ 23 ], the set of persuasion techniques considered differs substantially from other datasets. SemEval-2023 [ 249 ] introduced a large-scale multilingual dataset covering 9languages (English, French, German, Georgian, Greek, Italian, Polish, Russian, and Spanish), with almost 50 , 000 news articles labelled with 23 different persuasion techniques, therefore being the largest dataset currently available in terms of number of instances, languages, and persuasion techniques. Three languages were considered “surprise languages” (Georgian, Greek, and Spanish), for which training sets were not available during the competition, thus encouraging systems to deal with out-of-domain data. Furthermore, the set of 23 persuasion techniques were grouped into 6coarse-grained categories. For example, the techniques of Loaded Language,Obfuscation, Intentional Vagueness, Confusion, Exaggeration or Minimisation, and Repetition, were grouped into the umbrella of Manipulative Wording. 6.2 Methods and models The majority of methods and models developed for the task of automatic detection of persuasion techniques were introduced in the shared tasks discussed in the previous section. The joint effort of multiple different attempts at producing the best-scoring system allows to identify which methodological decisions are key to producing accurate models to detect persuasion. Table 7 summarises the highest-scoring 19 systems across the shared tasks aimed at automatic detection of persuasion techniques. Transformer-based models were employed in all top-scoring systems for automatic detection of persuasion techniques due to their ability to capture nuanced contextual information [ 322 ]. In the NLP4AI-2019 shared task [ 77 ], 5out of the 6submissions included transformer-based models, while more recently in SemEval-2023 [ 249 ], all 16 submissions were comprised of transformer-based models. In earlier shared tasks, BERT was initially preferred over other architectures such as RoBERTa, ALBERT, and DeBERTa, which were adopted more frequently later, specially RoBERTa. In SemEval-2019 BERT was used in all transformer-based submissions. In SemEval-2020 and SemEval-2021, BERT was used in a total of 25 systems, while RoBERTa is used in 14 systems. Nevertheless, systems using RoBERTa achieved the top 2submissions in both shared tasks [ 108 , 151 , 211 , 317 ]. For tasks in Arabic (WANLP-2022 and ArAIEval-2022), monolingual models (AraBERT and variations) outperformed multilingual models (e.g., XLM-RoBERTa and mBERT) [ 122 ]. In SemEval2023 where 9languages were available, multilingual models (e.g., mBERT and XLM-RoBERTa) outperformed monolingual models trained with each language separately [ 350 ]. Unsurprisingly, the larger variations of the models generally outperformed the smaller versions. For reference, RoBERTa large has almost triple the size of RoBERTa base (355 and 125 million parameters, respectively). Also, model ensembles combining different pretrained models were widely and effectively employed, although at the cost of requiring more computational resources [66,108,129,211,317]. In addition to the contextual embeddings generated by transformer-based models, some works have experimented with supplementary input features such as part-of-speech (POS) tags [ 159 , 211 , 270 ], named entity recognition encodings [ 159 , 211 ], word-level n-grams [ 101 , 159 , 161 ], characterlevel n-grams [ 66 , 317 ], and sentiment scores [ 159 , 270 ]. However, apart from the contextual embeddings, it is not clear if other input representation methods play a significant role in improving performance. In fact, apart from Morio et al . [211] who used named entity recognition encodings, 19Shared task systems that did not publish a description of their approach are not considered in our analysis. ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:33 Table 7. Top-scoring systems for automatic detection of persuasion techniques across different datasets. Dataset System Model Approach F1 Micro SemEval-2019 [81]newspeak [361] BERT base uncased - Token-level classification with 20 classes: No PTs, one of the 18 PTs, and an auxiliary class to handle word-level tokenisation. - Oversampling and class weighting. 0.2488 stalin [87] GROVER large [365] - Linear projection of contextual embeddings - SMOTE oversampling - BiLSTM classifier 0.1453 SemEval-2020 [78]Hitachi [211] Ensemble (BERT, GTP-2, RoBERTa, XLM XLM-RoBERTa, XLNet) - BIO encodings - Contextual embeddings with POS tags and named entities. - Three training objectives: (i) BIO tag, (ii) token-level, and (iii) sentence-level classification. - Two BiLSTMs, one for objectives (i) and (ii), and another for (iii). - Class weighting 0.5155 ApplicaAI [151] RoBERTa large - Self-training using additional data (500k sentences from OpenWebText). - Added a continional random field (CRF) layer. 0.4915 SemEval-2021 [85]MinD [317] Ensemble (BERT, RoBERTa, XLNet, DeBERTa, ALBERT) - Uses additional data from Da San Martino et al. [81]. - Model ensemble - Custom rules for the Repetition technique. - Character-level n-grams 0.593 Volta [110] RoBERTa - Used backtranslation as data augmentation 0.57 WANLP-2022 [12]NGU CNLP [136] AraBERT - Translated the PTC dataset [81] to Arabic and used as additional training data - Stacking-based model ensemble 0.649 IITD [206] XLM-RoBERTa large - Simply fine-tuned the model using the task dataset 0.609 ArAIEval-2022 [122]UL & UM6P [166] AraBERT-Twitter-v2 - Used an asymmetric multi-label loss objective [264]0.5666 rematchka [2] AraBERT-v2 - Class weighting - Balanced data sampler 0.5658 SemEval-2023 [249] Razuvayevskaya et al. [262] XLM-RoBERTa large - Multilingual joint fine-tuning. - LoRA - Class weighting 0.429 KInITVeraAI [131] XLM-RoBERTa large - Multilingual joint fine-tuning. - Carefully chosen classification threshold for each language (around 0.2). 0.42 Ampa [235] XLM-RoBERTa large - Oversampling - Ensemble with models trained with one and multiple languages. 0.395 all other 1st and 2nd placing systems in shared tasks have not used other supplementary features apart from the contextual embeddings. Arguably the most relevant methods employed in top-performing systems are aimed towards dealing with data skewness. Most datasets for this task suffer from class imbalance, meaning the majority of persuasion techniques are underrepresented, while a small set of persuasion techniques comprise the majority of the dataset. To deal with this issue, several different methods are employed, such as oversampling underrepresented techniques [ 87 , 235 , 361 ], scaling the contribution of the techniques to the loss function according to their proportion (i.e., class weighting) [ 2 , 211 , 262 , 361 ], and using supplementary data obtained with (i) data augmentation [ 110 ], (ii) semi-supervised methods [151], or (iii) similar datasets [3,136,206,317]. ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:34 Srba et al. 6.3 Tools and services In terms of production grade tools to detect persuasion techniques in texts, to the best of our knowledge, two services are currently available. The Propaganda Persuasion Techniques Analyzer (PRTA) [ 80 ] is a discontinued tool that is no longer available online. It was designed to detect instances of propaganda in texts by highlighting the spans where specific techniques were used. It provided users with the ability to compare texts based on their use of propaganda techniques, offering detailed statistics on the prevalence of these techniques, both overall and over time. PRTA used a BERT model trained on the SemEval-2019 dataset, fine-tuned for both fragment-level and sentence-level classification tasks. To gather data, PRTA crawled a growing list of over 250 RSS feeds, Twitter accounts, and websites, extracting text via the Newspaper3k library and performing deduplication using a hash function. The system then identified sentences containing propaganda, organised the articles into topics (such as COVID-19 or Brexit), and allowed users to compare the use of propaganda techniques across various media sources. The tool also allowed for filtering based on time intervals, keywords, or political orientation of the media. Although the user interface of the PRTA tool is no longer available, the underlying API for propaganda detection (along with other credibility signals) remains accessible20. The GATE Cloud infrastructure [ 76 ] operated by the University of Sheffield provides a service to detect 23 different persuasion techniques across multiple languages using a multilingual BERT (mBERT) model 21 . The model was fine-tuned using data from SemEval-2023 in English, French, German, Italian, Polish, and Russian. Class weighting was applied during fine-tuning to account for the label skewness issue. The tool is capable of processing up to 1 , 200 documents per day free of charge through its API, with an average processing rate of 2 documents per second. Researchers can request higher quotas if needed for larger-scale analysis. 6.4 Discussion Most commonly adopted persuasion techniques. Table 6shows that 7out of the 10 datasets present a similar set of persuasion techniques [ 12 , 13 , 78 , 81 , 85 , 122 , 249 ], while [ 23 , 167 , 195 ] largely diverge from them. Considering these seven similar datasets, we observe a set of 12 persuasion techniques that are consistent among them: Appeal to authority,Appeal to fear/prejudice,Causal oversimplification,Doubt,Exaggeration or minimization,Flag-waving,Loaded language,Name calling or labelling,Slogans, and Whataboutism appear in all seven datasets, while Repetition appears in six datasets (except [ 13 ]), and Reductio ad Hitlerum appears in five datasets (except [ 122 , 249 ]). Combining existing datasets with overlapping persuasion techniques could be a promising research direction to develop more robust models and benchmarking resources. For example, Tian et al . [317] achieved 1st place in SemEval-2021 Task 6 by leveraging the dataset from [ 81 ] as additional training data. Label skewness. A common characteristic inherent to this label scheme is the significant skewness in distribution of the persuasion techniques. For example, in SemEval-2023 [ 249 ] (the largest multilingual dataset, with 50 , 000 instances), a small set of 6techniques represent 71 . 8% of the entire dataset - Loaded Language (18 . 5%), Name Calling-Labelling (23 . 7%), Casting Doubt (12 . 5%), Questioning the Reputation (7 . 6%), Appeal to Fear-Prejudice (4 . 8%), and Exageration/Minimisation (4 . 7%) - while the remaining 17 techniques account for only 28 . 2% of the dataset. The same trend can be seen in [ 81 ], with 6techniques representing 68 . 2% of the dataset - Loaded language (34%), Name calling, labelling (17 . 3%), Repetition (10 . 2%), Exaggeration, minimization (7 . 6%), Doubt (7 . 5%), and Appeal to fear/prejudice (4 . 9%) - while the remaining 12 techniques account for 31 . 2% of the dataset. 20https://apihub.tanbih.org/propaganda/ 21https://cloud.gate.ac.uk/shopfront/displayItem/persuasion-classifier ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:35 This sparse label distribution poses several challenges for training and evaluation. In particular, the overrepresentation of a few dominant techniques may lead models to bias their predictions toward these frequent categories, while underrepresented techniques suffer from insufficient training examples. This data imbalance increases the risk of overfitting to the few available samples of rare classes, potentially resulting in poor generalization. Evaluation. Shared tasks on persuasion techniques detection typically employ the F1-Micro measure as the official evaluation metric [ 12 , 78 , 85 , 122 , 249 ], however, the F1-Micro does not account for skewed class distributions, as opposed to F1-Macro 22 . This evaluation strategy encourages models that excel at predicting a few densely distributed persuasion techniques. For example, in ArAIEval-2023 [ 122 ], the baseline model (majority vote classifier that always predicts the most common persuasion technique), achieves an F1-Micro of 0 . 3599, and an F1-Macro of 0 . 0279, representing a difference of 92% between the two scores. Similarly, the best submission by Lamsiyah et al. [166] achieves F1-Micro and F1-Macro scores of 0.5666 and 0.2156, a difference of 62%. Similar figures are seen in other shared tasks such as SemEval-2023 [ 249 ], with the average F1-Micro score for first place submissions across all languages resulting in 0 . 454, and the F1-Macro in 0 . 211, a difference of 53%. These figures highlight how current state-of-the-art models trained for this task still lack the ability to predict underrepresented persuasion techniques effectively. Multilinguality. Despite efforts to create multilingual datasets such as SemEval-2023 [ 249 ] and Macagno [195] , most datasets focus on English or Arabic. For languages without annotated datasets, machine-translation and zero-shot classification approaches may be viable alternatives (if some level of noise introduced by machine translation is acceptable). Nevertheless, joint multilingual training often result in more accurate models. For instance, Razuvayevskaya et al . [262] experimented training monolingual models with translated data from SemEval-2023, which improved performance only for English, but not for other languages. Therefore, multilingual datasets with wider variety of languages are still required to produce state-of-the-art models and evaluation benchmarks. Adoption of LLMs. Recent research has increasingly explored the capabilities of large language models in generating persuasive and deceptive content, particularly in the context of disinformation and propaganda [ 48 , 236 , 271 , 298 , 356 ]. LLMs can manipulate existing factual content into misleading or false narratives at scale, often outperforming human-written disinformation in believability and fluency [ 62 , 202 ] and can even personalize the content to specific target groups [ 372 ]. However, while the generation side has been widely studied, significantly less attention has been paid to using LLMs for detecting or countering persuasive techniques. Bridging this gap by harnessing LLMs for persuasion detection presents a promising but underexplored research direction. 7 CHECK-WORTHY AND FACT-CHECKED CLAIMS Given the impossibility to verify the veracity of every single piece of information posted online, two credibility signals and corresponding tasks are particularly crucial in supporting fact-checkers and other actors involved in the fight against disinformation and misinformation [ 106 ]: check-worthiness detection and the retrieval of previously fact-checked claims. The former is aimed at identifying which claims are check-worthy because of their relevance and interest to the general public [ 124 ]. The second, which ideally takes place after the first one, is meant to ease claim verification by checking whether the given claim was already verified before by searching in existing repositories of fact-checked claims [287], both monolingual and multilingual. While these tasks are usually studied in separation as a part of the fact-checking pipelines, their outputs can be and also are used (usually in combination) as credibility signals to indicate whether 22 Generally F1-Macro is also reported as a secondary metric, but for the purpose of the competition, the F1-Micro determines which system is best. ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:36 Srba et al. the check-worthy (central) claims contained in a piece of content have been fact-checked and if so, what was the given veracity value or values. In fact, they are recognized as such both in the credibility signals list created by the Credible Web Community Group (signals ‘article has a central claim’ and ‘fact-check status of a claim’; see [ 217 ]) as well as in the related works [ 366 ]. The check-worthiness of a claim can be considered a content-based signal, while fact-check status of a claim is a context-based one, since it requires additional external sources to be detected. On the other hand, fact-checking itself is not considered a credibility signal per se – its aim is to directly ascertain the veracity value of a piece of information, while the former are only signals that the information might be check-worthy (and thus more attention is needed before the users make credibility judgment) or that it contains a claim that have already been fact-checked (but it does not by itself analyse the stance towards that claim, only its presence). It is also worth noting that these signals differ from factuality signal as it is understood in the surveyed works as discussed in Section 5). 7.1 Datasets Check-worthy claim detection is a popular task within the NLP community thanks to the series of CheckThat! shared tasks organised at CLEF, which led to the release of related datasets. Indeed, check-worthy claim detection is the only subtask that has been proposed at all seven CheckThat! editions [ 10 , 18 , 19 , 34 , 36 , 219 , 221 ]. Through those editions, datasets for training check-worthy claim detection models have been constantly extended to cover additional languages, starting from English and Arabic at CheckThat! 2018 [ 18 ] to Arabic, Dutch, English and Spanish at CheckThat! 2024 [ 34 ] and even multimodal and multi-genre data in the 2023 edition [ 10 ]. Beside CheckThat!, additional datasets have been created and made available for research, as shown in Table 8. The sources covered by such datasets are typically social media, news and political debates, i.e. three relevant areas where the presence of disinformation may have detrimental effects on a large audience. Particular relevance was given to COVID-related content during the pandemic [ 112 , 275 ], when the continuous flow of misleading information could negatively affect public health. Although English is the most represented language, datasets for check-worthiness detection have also been developed in several other languages. This can be attributed to the fact that public relevance is both timeand geographically bounded, with claims typically referring to events specific to individual countries or regions. Concerning the retrieval of previously fact-checked claims, there are several datasets containing verified (fact-checked) claims collected from professional fact-checking organizations, such as X-Fact [ 107 ], MultiFC [ 21 ] or ClaimsKG [ 316 ]. Being usually collected for the task of automatic fact-checking, these are not directly usable for the task of retrieval of previously fact-checked claims by themselves, as they lack the input claims (e.g., social media posts) and the pairs of the input and the verified claims. The first datasets designed specifically for the task appeared in 2020 in the works by Shaar et al. [ 287 ] and Vo and Lee [ 324 ] who independently prepared datasets based on both Snopes and PolitiFact. The former became the basis for a series of CheckThat! Lab shared tasks (Task 2) organised at CLEF in 2020 [ 36 ], 2021 [ 289 ] and most recently 2022 [ 220 ]. The tasks gradually expanded the original dataset in size and also by adding an additional language (Arabic). The political debates part of the dataset was additionally expanded in [288]. The datasets for retrieval of previously fact-checked claims are usually collected by using one of the following approaches: (i) looking into the fact-checking articles for links to the original content making the claim that is being verified, e.g., [ 36 , 220 , 247 , 287 , 289 , 300 ] or (ii) searching for (social media) content (such as discussion threads) that contains URL of or semantic links to the fact-checking articles, e.g. [ 120 , 223 , 324 ]. The former typically achieves high precision at the cost of lower number of pairs and possible missing connections (i.e., many false negatives), while ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:37 Table 8. Selected datasets used for check-worthiness detection. Note that the datasets used in the different CLEF CheckThat! editions often overlap. Dataset Language(s) # Instances Content type Classes TR-Claim19 [156] Turkish 2,287 Tweets Check-worthiness 26 rationale categories CW-USPD-2016 [100] English 5,415 Political debates Check-worthiness Dhar and Das [84] Bengali, Hindi 2,402 Political news, Twitter Check-worthiness MM-Claims [61] English 3,400 tweets (image + text) Claim detection, check-worthiness, visual relevance Sheikhi et al. [292] Norwegian 4,885 News Check-worthiness AraCOVID19-MFH [112] Arabic 10,828 COVID-related tweets Check-worthiness, factual, hateful Faramarzi et al. [275] English 7,017 COVID-related tweets Check-worthiness, claim extraction ClaimBuster dataset [16] English 22,281 Presidential debates Check-worthiness, factual CLEF-2018 CheckThat! Lab Task 1 [18]English, Arabic 17,300 Political debates Check-worthiness CLEF-2019 CheckThat! Lab Task 1 [19]English 24,000 Debates, speeches, press conferences Check-worthiness CLEF-2020 CheckThat! Lab Task 1 [36]English, Arabic 962 (EN) 7,500 (A), Tweets, political debates, speeches Check-worthiness CLEF-2021 CheckThat! Lab Task 1 [221] Arabic, Bulgarian, English, Spanish, Turkish 18,014 (tweets), 50,123 (sentences) Debates, speeches, tweets Check-worthiness CLEF-2022 CheckThat! Lab Task 1 [219] Arabic, Bulgarian, Dutch, English, Spanish, Turkish 30,363 Tweets Check-worthiness, verifiable, harmful CLEF-2023 CheckThat! Lab Task 1 [10]Arabic, English, Spanish 70,806 Tweets, political debates, speeches Multimodal and multigenre check-worthiness CLEF-2024 CheckThat! Lab Task 1 [34] Arabic, Dutch, English, Spanish 64,700 Tweets, political debates, speeches Check-worthiness the latter can maximise recall at the cost of introducing noise unless manual checking is applied. Typical example of a high noise are CrowdChecked [ 120 ] or MuMiN [ 223 ] which both contain a large number of social media posts which distinguishes them from other datasets. While many available datasets focus on English only, there are several newer ones supporting other languages (e.g., Arabic [ 220 , 289 ] or Spanish [ 201 ]) or even a range of languages [ 158 , 223 , 247 , 300 ]. The most notable among these are MultiClaim [ 247 ] and MMTweets [ 300 ] due to the number of included languages, high precision of identified pairs and their amount as well as due to the fact that they both introduced a task of crosslingual retrieval, in which input claims are in different language than that of the verified claims. A full list of selected relevant datasets is reported in Table 9. Besides these, it is also worth mentioning two datasets focusing on COVID-related claims, namely CoAID [ 74 ] and MM-COVID [ 176 ]; however, these focus on a different task of fake news/disinformation detection. A dataset of COVIDrelated tweets and fact-checks is also presented in [ 119 , 148 ]. In this case, although the dataset was introduced for a different task, it could also be useful for previously fact-checked claim retrieval. 7.2 Methods and models To propose methods for check-worthy claim detection, existing works first define how checkworthiness is operationalized. In a recent survey, Panchendrarajan and Zubiaga [229] identify two main aspects making a claim check-worthy: its verifiability and its priority. The first item refers to the possibility of determining the veracity of a claim, which can be likely supported by evidence. A claim is verifiable when it contains a factual statement that can be checked, which means that personal opinions or events presented as uncertain are excluded. Second, given that verifying all statements about the world is impossible, it is important to prioritize claims which are considered ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:38 Srba et al. Table 9. Selected datasets used for retrieval of previously fact-checked claims. Dataset Language(s) # Instances Content type # Input claims # Verified claims # Pairs That is a Known Lie – Snopes [287] English 1,000 10,396 1,000 social media posts (Twitter) That is a Known Lie – PolitiFact [287] English 768 16,636 768 political debates Snopes (Vo and Lee [324]) English 11,167 1,703 11,202 social media posts (Twitter) PolitiFact (Vo and Lee [324]) English 2,026 467 2,037 social media posts (Twitter) Kazemi et al. [158] English, Hindi, Bengali, Malayalam, Tamil NA NA 2,343 instant messages (WhatsApp) CLEF-2022 CheckThat! Lab Task 2A [220] English, Arabic 2,518 44,214 2,699 social media posts (Twitter) CLEF-2022 CheckThat! Lab Task 2B [220] English 752 20,771 869 political debates and speeches CrowdChecked [120] English 316,564 10,340 332,660 social media posts (Twitter) MuMiN [223] 41 languages 21,565,018 12,914 NA social media posts (Twitter) NLI19-SP (FacTeR-Check) [201] Spanish 40,000 61 NA social media posts (Twitter) MultiClaim [247] posts in 27 languages, factchecks in 39 languages 28,092 205,751 31,305 social media posts (Twitter, Facebook, Instagram) MMTweets [300] posts in 4 languages, factchecks in 11 languages 1,600 30,452 4,258 social media posts (Twitter) timely, interesting to the general public and whose verification might have a broader impact [ 82 ]. In this latter aspect it differs from a related, but a distinct task of claim detection which solely aims to identify what constitutes a claim in a text (either using a binary classification or by identification of spans) [109,207,354]. Given this need for prioritization, check-worthiness detection has been cast as a ranking problem since the first editions of the CheckThat! Lab within the CLEF Evaluation initiative 23 , which made the task well-known within the NLP community, although some related works had already been presented before [ 124 ]. In particular, given a political debate, the first CheckThat! task was aimed at predicting which claims should be prioritized for fact checking. This is reflected in the evaluation approach proposed for the CheckThat! series, which has become the de facto standard in checkworthiness detection: systems should output the list of input claims ranked by check-worthiness, which is usually evaluated using Mean Average Precision (MAP), reciprocal rank and 𝑃 @ 𝑘 for 𝑘 ∈ { 1 , 3 , 5 , 10 , 20 , 30 } . More recently, check-worthiness has been, however, evaluated using F1 as a binary classification task rather than ranking [10,34]. Regarding methods, in the CheckThat! editions up to 2023 state-of-the-art results for checkworthy claim detection were reached by methods that rely on fine-tuned transformer-based methods such as BERT, RoBERTa, DistilBERT, [ 97 , 282 , 348 ] and language-specific variants, [ 10 ], often combined with data manipulation or model ensembling strategies. For example, the best results on the English portion of the CheckThat 2022 dataset were achieved by a RoBERTa model that leveraged a back-translation-driven data augmentation process [ 280 ]. The 2024 edition, instead, has seen an increased interest in using also generative LLMs for the task such as LLama 2 and 3, GPT3.4 and 5, Mixtral and Mistral [ 34 ]. For instance, the best performing system on the English language dataset relied on fine-tuned Llama2 7b on the original training data, using prompts generated by ChatGPT [ 178 ]. Nevertheless, despite the notable progress in data and methods for check-worthy claim detection, there is currently a lack of an extended coverage across languages and topics. Concerning previously fact-checked claims, the task is also formulated as a retrieval one, although the name of the task may vary across the literature; fact-checking URL recommendation [ 323 ], 23https://www.clef-initiative.eu/ ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:39 detection of previously fact-checked claims [ 287 ], verified claim retrieval [ 36 ], searching for factchecked information [ 324 ], claim matching [ 158 ] or retrieval of previously fact-checked claims [ 247 ] have all been previously used to denote it. Being a retrieval task, the existing methods apply one or a series of (re-)rankers and mean average precision (MAP),mean reciprocal rank (MRR), 𝑃 @ 𝑘 or a 𝑆𝑢𝑐𝑐𝑒𝑠𝑠 @ 𝑘 ( 𝐻𝑖𝑡 @ 𝑘 ) are used as evaluation metrics. In case of a series of rankers, the works tend to use a baseline ranker that is easy to compute and has a good recall and then one or more subsequent rerankers that work over a progressively smaller subset of results retrieved by a previous ranker in the pipeline, see, e.g., [ 120 , 287 ]. The rerankers’ task is to improve precision by moving the relevant results to the top; since they are working with a smaller set of results, they can be more computationally demanding. Alternatively, the rankers could be used in combination as an ensemble to improve the precision at the cost of higher computational demands, but this is not observed in the surveyed works. In most of the works, BM25 [ 268 ] or similar information retrieval algorithms are used as a baseline. Various neural text embedding models are used as either sole rankers (e.g., in [ 247 ]), rerankers (e.g., in [ 287 ]) or as ensembles [ 201 ], especially sentence transformers [ 263 ], which use Siamese networks to pre-train text representations, usually on various sentence similarity datasets. The approaches also use several other techniques to improve the retrieval performance, such as text embedding models fine-tuning [ 158 , 247 ], distance supervision to work with noisy data [ 120 ], key sentences extraction [ 293 ], extended context of the input and verified claims (especially for political debates [ 286 ]), extraction of text from images [ 247 , 324 ], using multimodal representation combining text and images [ 324 ], or query (input claim) modification or rewriting to be more easily matched with the fact-checked claims [ 40 , 157 , 313 ]. The approaches may further differ by the use of loss, selection of negative examples and other (hyper-)parameters when fine-tuning the neural models. All surveyed approaches (including solutions submitted to the CheckThat! Lab challenge [ 36 , 220 , 289 ]) rely on smaller languages models, such as BERT, XLMRoBERTa, etc. One exception is the work of Sundriyal et al . [313] , where the authors employ LLMs to normalize the claims for the purpose of query rewriting for the retrieval task; however, they do not perform experiments on the retrieval task, thus focusing only on the first step (claim detection) in the pipeline. 7.3 Tools and services Given that automatizing check-worthiness detection can greatly support and speed up fact-checking activities, a number of systems has been already developed for the task, some of which are based on insights and databases actually used by fact-checking organisations. ClaimBuster [ 123 ] was the first end-to-end system for computer-assisted fact-checking trained on a human-labeled dataset of check-worthy factual claims from the U.S. general election debate transcripts. The first component of the pipeline is a detector of check-worthy factual claims which, given a sentence, first labels it as being ‘non-factual’, ‘factual and unimportant’ or ‘factual and check-worthy’. In case of the latter, a ranking score is assigned based on SVM decision function. Patwari et al . [234] present Tathya, a tool focusing only on check-worthiness detection, which compared to ClaimBuster can yield a significant performance improvement, particularly on recall. It is based on a multi-classifier system using features such as topics, entity history and PoS tuples. ClaimRank [ 141 ] performs check-worthy claim detection and supports English and Arabic texts. Its strength is that it was trained on actual annotations from nine reputable fact-checking organizations, therefore mimicking their real strategy for claim selection. The ranking is based on a number of lexical, structural and semantic features, used to train a neural network with two hidden layers as proposed by Gencheva et al . [100] . Another system, focusing specifically on tweets in ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:40 Srba et al. Arabic, is Tahaqqaq [ 291 ], which includes the possibility to identify check-worthy claims, estimate the user credibility in terms of spreading fake news, and find authoritative accounts. dEFEND [ 75 ] is another end-to-end system that, given a link to a post or a news, detects checkworthy sentences by assigning them a score, with the goal to distinguish between check-worthy factual claims from subjective ones. The system displays also the propagation network of the text as well as an analysis of the news comments and the textual evidence supporting the classifier decision. Another similar end-to-end platform, providing evidence snippets to credibility classification and check-worthiness decisions, is BRENDA [ 46 ], which provides also the possibility to collect users’ feedback about wrong predictions. The tool is released as Google Chrome extension. Among the few systems dealing with languages other than English, FactRank [ 39 ] was the first system able to process check-worthy texts in Dutch. The classification algorithm was developed iteratively, combining expert fact-checker input, a codebook to support reliable human labelling, and active-learning. Check-worthiness classification performance is comparable to results obtained on English with ClaimBuster. Concerning retrieval of previously fact-checked claims, it is supported by some of the end-to-end verification systems mentioned above, namely ClaimBuster [ 123 ], Tahaqqaq [ 291 ] or BRENDA [ 46 ]. Besides these, Google Fact-Check Explorer 24 is often used to perform the task since it indexes a large corpus of fact-checks. Other specialized tools include Fact-Check Finder 25 built on models developed in [247]. 7.4 Discussion Multilinguality. We observe in both check-worthiness detection as well as in retrieval of previously fact-checked claims stronger shift towards multilinguality. This is, on one hand, reflected in newer datasets containing also languages other than English (either multilingual ones or focused on a specific language), on the other by the more prevalent use of multilingual models. Since the amount of data in other languages is often limited, approaches for transfer learning [ 158 , 190 ], lowresource fine-tuning [ 58 ] or adapter fusion [ 283 ] are explored. However, using translation to English in combination with an English language model can still sometimes outperform a multilingual approach, as was observed, e.g., in the previously fact-checked claim retrieval task [ 247 ]. This is likely to change in the future with the employment and/or development of larger and better balanced multilingual models. Multimodality. Although most available datasets are mostly textual, if they contain links to the original content where the claim was made, it is sometime possible to get to other modalities, such as images, videos or audio, which can be contained in a piece of content (e.g., a social media post). These can be important, because in many cases it is there where the actual claim is being made or the claim requires multiple modalities to be properly interpreted. Multimodal content can also be perceived as more credible by users, has a higher engagement and is increasingly easier to produce [ 7 ]. Thus, multimodal approaches continue to grow in both prevalence and importance. At the moment, most approaches transcribe the modality to text by either using OCR or image description approaches [ 200 , 247 ], however, there are already some approaches that process the other modalities directly, be it images [324] or speech [138]. The future advances will likely lie in advancements of the latter category of approaches. Availability of datasets. Although there are available resources for both check-worthiness detection and retrieval of previously fact-checked claims, both tasks have their own (sometimes overlapping) sets of challenges. In case of check-worthiness detection, it is relatively easy to collect 24https://toolbox.google.com/factcheck/explorer 25https://fact-check-finder.kinit.sk ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:41 check-worthy claims – these are all claims verified by fact-checkers. However, collecting non-checkworthy ones is much more challenging. In case of retrieval of previously fact-checked claims, the challenge lies in collecting input claims (e.g., in the form of social media posts) and in identification of pairs between the input and the fact-checked claims. As discussed above, existing methods either lead to too strict matching with many unidentified (false negative) pairs or to too much noise in the data. Another issue for both tasks is that many datasets were previously built using Twitter. If only the IDs of the tweets have been published, it is now very expensive for researchers to use them due to the X’s current API limitations and pricing, thus making their use impractical or completely unfeasible. Combination of check-worthy claim detection and retrieval of previously fact-checked claims. As can be seen from the surveyed works, most of them approach the tasks in separation. This is reasonable from the scientific perspective, but more end-to-end (combined) approaches capable of first detecting a check-worthy claim and then retrieving previously fact-checked claims are needed for practitioners to use. Adoption of LLMs. Although we observe some approaches using LLMs in check-worthiness detection [ 34 ], their potential for the retrieval of fact-checked claims begins only now to be more systematically explored. Most recently, LLMs have been used and evaluated as text embedding models as well as rerankers of retrieved fact-checked claims in multilingual and crosslingual settings in [ 260 ], outperforming the fine-tuned smaller models. They have been also employed in zeroand few-shot settings as binary classifiers filtering out irrelevant retrieved results using a range of prompting strategies [ 248 , 327 ]; these recent works showed that while useful also in this setup, no single prompting strategy proved as the best overall and the performance of the current LLMs is lower for low-resource languages compared to high-resource ones. The use of LLMs (as text embedders, filters or rerankers) have been further explored in system papers submitted to the recent SemEval-2025 Task 7 as summarised in [ 242 ]. Finally, the benefit of LLMs for both check-worthiness detection or retrieval of previously fact-checked claims lies in input claim normalization [ 313 ], using retrieval augmented generation [ 134 ], or in providing summaries of the check-worthy content or of the retrieved fact-checks [326]. 8 ADDITIONAL CREDIBILITY SIGNALS 8.1 Text quality Text quality is a broad category of credibility signals measuring text’s linguistic accuracy, such as readability, grammatical correctness, or spelling mistakes. It is strongly related to perceived credibility, since high-quality, more professional, content is often seen as more trustworthy. Research on statin-related websites found that more readable and accurate information significantly improves users’ perceptions of credibility [ 184 ]. A similar study by Kiili et al . [160] highlighted that professionalism in text, such as proper grammar and clear structure, plays a crucial role in how credibility is judged. The similar relation to credibility can be observed for a low-quality content. Harris [121] showed that poor grammar or frequent spelling errors can be a signal of lack of credibility, prompting readers to question the reliability of the information presented. Greškovičová et al . [104] investigated how various editorial elements such as superlatives, clickbaits, boldface and poor grammar affect the quality and credibility of online messages. Thus, the quality of the text directly enhances trust in it. Various NLP techniques have been developed to rate text quality and, by extension, its credibility. Mosquera and Moreda [212] evaluated text by extracting features like contractions, slang, misspellings, emoticons and readability (using the Readability Index). They also measured information content through entropy and emotional tone using emotional distance and, finally, applied the ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:48 Srba et al. Think Twice Before Getting That FLU SHOT! I personally know three people who got the flu right after being vaccinated—coincidence? I don’t think so. Big Pharma just wants your money and doesn’t care if you get sick. They keep pushing these so-called “safe” vaccines every year, but how come flu cases always go up right after the campaigns? I am convinced that the shots are spreading the virus. Don’t be a sheep. Protect yourself naturally— boost your immune system with vitamins, not chemicals! #WakeUp #FluShotScam #NaturalImmunity Think Twice Before Getting That FLU SHOT! I personally know three people who got the flu right after being vaccinated—coincidence? I don’t think so. Big Pharma just wants your money and doesn’t care if you get sick. They keep pushing these so-called “safe” vaccines every year, but how come flu cases always go up right after the campaigns? I am convinced that the shots are spreading the virus. Don’t be a sheep. Protect yourself naturally— boost your immune system with vitamins, not chemicals! #WakeUp #FluShotScam #NaturalImmunity Event Factuality Subjectivity Bias Persuasion Techniques Logical Fallacies Fact-checked status Text Quality Offensive Language Machine-generated Text Clickbait Title Highly subjective: Relies on personal anecdotes ("I personally know...") and emotionally loaded language ("Big Pharma", "Don’t be a sheep"). Strong anti-vaccine bias: Positions vaccination as harmful and implies malicious intent by pharmaceutical companies. Highly present: Text is using persuasion techniques like Fear appeal, Bandwagoning appeal or Appeal to nature. Medium present: Text is using logical fallacies like Ad hominem or Straw man. Medium-low quality: Text is using informal style, all-caps emphasis, emojis, hashtags, and exaggeration. Mildly offensive: “Don’t be a sheep” is a derogatory phrase implying stupidity or blind obedience. Likely machine generated: Text appears to use the selection of words corresponding to OpenAI models. Uses clickbait elements: Urgent command (“Think Twice”) and controversial implication (flu shots are harmful). High event factuality: Author is very sure that the events happened (three people got flu after being vaccinated, the shots are spreading the virus). Previously fact-checked: Contains claim that has been previously factchecked as false (https://factcheck.afp.com/doc.afp.com.36JN78V). Overall Credibility Low (23%) Fig. 7. An illustrative example of diverse credibility signals determining the social media post as being of a low credibility. Each signal is associated with predicted values and a short explanation. The overall credibility label and score reflects the credibility predicted by individual signals. Third, researchers working on individual credibility signals frequently appear unaware of related efforts in neighbouring categories, despite these signals often being identified under the shared conceptual umbrella of credibility. Stronger interconnection between these lines of research could unlock mutual benefits – for instance, the creation of shared, curated datasets annotated for both individual signals and overall credibility (also building on the first efforts in this direction by [ 41 , 366 ], as discussed in Section 4.4). This would not only increase annotation quality and sample size but also enable the study of correlations between different signals and support the development of multitask learning models – techniques currently underutilized due to a lack of such integrated resources. Interestingly, although research efforts remain isolated in their own silos built around the single NLP task, some end-user-oriented tools already incorporate multiple credibility signals in practice. Notable examples include the Tanbih system27 [369] and the Verification Plugin28. To illustrate the potential benefits of detecting and aggregating multiple credibility signals (representing individual categories systematically covered by this survey) into a unified credibility label/score, we provide two illustrative examples: one depicting a low-credibility social media post (Figure 7) and another representing a high-credibility example (Figure 8). These examples clearly showcase an untapped potential of synergy between detection of diverse advanced credibility signals, their explanations, as well as aggregation. To achieve such desired state, we advocate for significantly stronger integration of current research efforts. 9.2 Adoption and potential of generative LLMs Generative Large Language Models (LLMs) have demonstrated substantial improvements in complex tasks that require reasoning abilities [ 257 ]. Brown et al . [50] showed that pretrained LLMs are 27https://tanbih.org/ 28https://www.veraai.eu/category/verification-plugin ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:49 Why Getting a Flu Shot Matters Flu vaccines do not cause the flu. In fact, they are designed to help your immune system recognize and fight the virus before you get exposed. While it’s possible to catch the flu after vaccination (especially if you were already exposed), the shot significantly reduces your chances of severe illness and hospitalization. According to the CDC, flu vaccination prevented an estimated 7.5 million illnesses and 6,300 deaths in the U.S. during the 2019–2020 flu season. https://www.cdc.gov/flu/vaccines-work If you have concerns, talk to your doctor—not the internet. #PublicHealth #VaccinesWork #FluPrevention Why Getting a Flu Shot Matters Flu vaccines do not cause the flu. In fact, they are designed to help your immune system recognize and fight the virus before you get exposed. While it’s possible to catch the flu after vaccination (especially if you were already exposed), the shot significantly reduces your chances of severe illness and hospitalization. According to the CDC, flu vaccination prevented an estimated 7.5 million illnesses and 6,300 deaths in the U.S. during the 2019–2020 flu season. https://www.cdc.gov/flu/vaccines-work If you have concerns, talk to your doctor—not the internet. #PublicHealth #VaccinesWork #FluPrevention Event Factuality Subjectivity Bias Persuasion Techniques Logical Fallacies Fact-checked status Text Quality Offensive Language Machine-generated Text Clickbait Title Objective tone: Neutral and informative, avoids emotional or personal wording. Slight institutional bias: Text contains bias toward health promotion, but aligned with scientific consensus and transparently sourced. Medium present: Text contains Evidence-based persuasion, Call to action, but uses facts and statistics, not fear or emotion. None present: Reasoning is logically sound; does not overclaim or generalize. High quality: Clear grammar, complete sentences, link to source, and no excessive punctuation or slang. None: Language is inclusive and respectful. Likely human written: Text appears to use selection of words that correspond to human-written texts. No clickbait: Straightforward and non-sensationalized title. Medium-high event factuality: Author is quite sure that the event happened (flu vaccination prevented illnesses and deaths). High (95%) No existing fact-checks: No existing fact-checked claim was detected. Overall Credibility Fig. 8. An illustrative example of diverse credibility signals determining the social media post as highly credible. Each signal is associated with predicted values and a short explanation. The overall credibility label and score reflects the credibility predicted by individual signals. capable of few-shot learning, meaning they can learn to perform new tasks with only a few training examples. Similarly, Petroni et al . [245] highlighted the strong ability of LLMs to recall relational knowledge acquired during pretraining to perform various tasks without further annotated labels or human supervision (i.e., zero-shot learning). Additionally, several key advancements, such as Retrieval-Augmented Generation (RAG) [ 172 ], Reinforcement Learning with Human Feedback (RLHF) [ 225 ], and robust prompting techniques [ 257 , 339 ], have further enhanced the capabilities of LLMs. In this context, recent generative LLMs operate as dialogue systems, where the model is prompted by the user with instructions to perform specific tasks. As we have discussed in the sections on individual credibility signals (see Sections 5,6,7, but also 8), the uptake of LLMs differs across the signals. In some cases, first such uses appeared only in late 2024 and beginning of 2025, such as in the case of previously fact-checked claims (Section 7). Nevertheless, their prevalence gradually increases. However, approaches using LLMs for credibility assessment (see Section 4) are still rare with one such notable exception being the work of Leite et al . [170] . Consequently, there are still several unexplored or underexplored opportunities how such models could address challenges related to the automatic detection of credibility signals. One of the key advantages of using prompting with LLMs is the flexibility in adapting a single foundational model to handle multiple subtasks associated with credibility assessment. With a carefully designed framework, LLMs can be guided to focus on different aspects of content analysis, such as detecting persuasion techniques, evaluating the veracity of claims, identifying potential bias, and recognizing patterns of misinformation. This flexibility reduces the need to develop and fine-tune separate models for each task, allowing practitioners to use the same model across various credibility-related tasks. In fact, a promising research direction is to explore multi-task learning, as in verifying if the capacity of performing certain credibility-assessment tasks can aid in other related ones (e.g., persuasion and bias). ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:50 Srba et al. Moreover, the capability of learning with zero/few examples is particularly valuable for tasks where domain-specific data is scarce or constantly changing, as is the case with credibility assessment. An enormous amount of human effort is required to curate high-quality annotated datasets for the different subtasks involved in assessing credibility. Specially since labelling most of credibility signals (such as biased content) often demand the expertise of domain specialists such as fact-checkers, journalists and social scientists. Adding to this challenge, credibility indicators can be highly context-dependent, varying across cultural and temporal dimensions. In this context, LLMs offer a more scalable approach by drawing on vast amounts of unsupervised pretraining data, and by adjusting to specific end-tasks through careful prompting strategies, which require far less human effort than manual data labelling. As an example, in Leite et al . [170] , a generative LLM was employed to predict 19 different credibility signals present in textual content without using any training data (i.e., in a zero-shot setting). Finally, the generative capabilities of large language models can be leveraged to produce more explainable and interpretable 29 outputs, which is crucial for subject-matter experts that may leverage the model’s predictions. Instead of providing only binary or scalar outputs (e.g., true/false, misinformation/non-misinformation, credible/non-credible) as in usual classification tasks, generative LLMs can produce detailed explanations or summaries that can highlight the reasoning behind their predictions. This transparency allows human experts to critically assess the model’s outputs, cross-check them with external information, and ensure that (i) incorrect predictions (in this context, known as model hallucinations [ 146 ]) are properly mitigated, and (ii) any credibility assessment aligns with the context of the content being investigated. This property of interpretable outputs can significantly increase trust in model-driven decisions and reduce the likelihood of over-reliance on machine predictions, ensuring that humans remain in the loop for final judgments. We acknowledge that LLM adoption (despite providing a lot of potential) is also accompanied with several challenges. Fine-tuning as well as deployment of LLMs require a considerably higher computing power which directly translates to higher costs. Moreover, learning techniques commonly used in limited labelled data scenarios (prompting, in-context learning, fine-tuning, meta-learning, or few-shot learning) are known to be sensitive to various randomness factors [ 241 ], what can result in undesired performance instability. Beside randomness factors, also systematic choices, such as the format of the prompt or how many (in-context) samples are used, have significant effect on the overall performance and the stability of these approaches [ 240 , 325 ]. Nevertheless, ongoing research in the area of LLMs is already providing suitable solutions to such problems, such as Parameter-efficient Fine Tuning (PEFT) techniques [ 355 ], which have been demonstrated to perform well also on credibility signal detection tasks [ 262 ]; or instability mitigation techniques, such as ensembling, noise regularisation and model interpolation [ 238 ]. At the same time, we would like to stress that employing LLMs does not necessarily provide benefit in all cases. When a sufficient amount of labelled samples is available (100-1000 depending on the specific task and model), fine-tuned smaller language models can outperform larger general ones [239]. Finally, as highlighted in Section 5, implicit biases present in various LLMs can make them unreliable in when assessing external biases or other credibility signals [ 29 , 192 ]. Efficient debiasing strategies, evaluation metrics for bias detection and alignment strategies must be implemented before such models can be applied to real-world tasks [ 180 , 209 ]. Factual correctness of LLMs is also vulnerable to slight contextual shifts and hallucinations [ 20 ], and improvements in source attribution, domain-specific adaptation and reasoning capabilities are cornerstone tasks in this direction that need to be addressed. 29 Here, the concept of interpretability differs from explainability, which is often used in the field of machine learning to refer to specific methods to analyse how intermediate states of the model lead to certain outcomes [52]. ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:51 9.3 Dataset availability and multilinguality Dataset availability heavily differs across various categories of credibility signals. Firstly, we observed a significant lack of large-enough datasets providing overall (expertly-determined) annotation of credibility that can be used for training and evaluating solutions on credibility assessment task (Section 4.1). Furthermore, datasets containing annotations for overall credibility and (at least some) credibility signals at the same time are even more scarce. As a result, many researchers opted to use (more available) fake news datasets as a replacement. At this place, we would like to highlight again that fake news annotations cannot reliably replace credibility annotations (non-credible content is not necessarily only false content and vice versa, since credibility is rather a parallel and complementary dimension to veracity; see Section 2). On the other hand, the situation with dataset availability for individual categories of credibility signals is much better. This is especially thanks to data challenges (particularly SemEval tasks and CLEF CheckThat! Labs as evidenced in Sections 5,6, or 7) in which either new datasets were introduced or existing datasets were extended (with more data, new languages, or additional types of annotations). Unfortunately, in some cases (e.g., persuasion techniques dataset by Piskorski et al . [249] introduced in Section 6), the datasets from data challenges are not published completely and a hidden/testing set is not shared with researchers not even when the competition is over (and thus only the training/validation sets remain available for further research). Especially for international news or the investigation of global claims, journalists and factcheckers need to verify the credibility of information by cross-checking sources in different languages. Furthermore, emerging disinformation in one country could be spread to other countries, especially when its topic is global (e.g., pandemic, wars, international relations). Therefore, it is important to implement credibility analysis tools that support multiple languages. However, due to the scarcity of multilingual datasets, trained models can exhibit biases towards some languages and cultures and hence can underperform on content in low-resourced languages (as is for example the case for fact-checked claims retrieval as evidenced in [327]; see Section 7). In Table 11, we present a summarised overview of language-specific datasets available for different categories of credibility signals. As in many other areas of NLP, English remains the dominant language, with the majority of available resources focused on English-language content. Nevertheless, notable datasets also exist for other widely spoken languages, particularly Arabic and Spanish. In contrast, there is a clear lack of annotated resources for many lowand midresource languages, such as Czech or Tamil. Among the covered categories, fact-checked claims demonstrate the broadest language coverage. This can be attributed to the global presence of fact-checking organizations, which produce multilingual artifacts that serve as the foundation for constructing such diverse and multilingual datasets. On the other hand, categories such as factuality and bias show limited availability of resources beyond English, highlighting a significant gap in the development of multilingual tools for credibility signals detection. To overcome the scarcity of multilingual datasets, global collaborations could be initiated for creating multilingual resources. In this direction, we can already observe a positive trend within data challenges. Many of them introduced multilingual datasets, commonly considering some languages as surprise ones (i.e., languages that are present in a test set, but missing in a train/validation set). Such approach motivates participants to develop multilingual solutions that are capable to make a prediction in a zero-shot setting (considering a language a predicted content is written in). Besides datasets availability and multilinguality, a quality of annotations remain another challenge. Annotation of overall credibility as well as individual credibility signals is many times highly subjective (as also shown in [ 41 ]), especially in cases such as persuasion techniques where presence of a signal and borders between various signals can be blurred. The situation is getting ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:52 Srba et al. Table 11. Summary of language-specific resources per individual categories of credibility signals. In a case of [ 223 , 247 ] datasets – which are highly multilingual, comprising 42 and 39 languages respectively – we directly report those languages that have an overlap with other datasets, while the number of remaining languages is provided in the Additional languages row. For a full list, please, refer to the original papers. Language Credibility Signals Subjectivity Factuality Bias Persuasion techniques Checkworthiness Fact-checked claims English [43,249,310] [28,230,309] [344] [174,255,256][249,309,312] [64,94,185] [77,167,249] [78,85,195] [61,100,275] [16,18,19] [36,219,221] [10,34] [220,287] [158,324] [120,223,247,300] Chinese – [255,256] – – – [223,247] German [17,28,249] – [8,249] [249] – [223,247] Urdu [342] – – – – [223,247] Arabic [28,213] – – [12,13,122] [18,36,112] [10,219,221] [34] [220,223,247] French [28,249] – [249] [249] – [223,247] Polish [249] – [249] [249] – [223,247] Italian [249] – [249] [195,249] – [223,247] Russian [249] – [249] [249] – [223,247] Spanish [28] – – – [10,219,221] [34][201,223,247,300] Romanian [28] – – – – [247] Portuguese [144] – – [195] – [223,247,300] Dutch [199] – – – [34,219] [223,247] Czech – – – [23] – [223,247] Turkish – – – – [156,219,221] [223,247] Hindi – – – – [84] [158,223,247,300] Bengali – – – – [84] [158,223,247] Norwegian – – – – [292] [223,247] Bulgarian – – – – [219,221] [247] Tamil – – – – – [158,223,247] Malayalam – – – – – [158,223,247] Additional languages – – – – – + 22 [223], + 18 [247] even more challenging in multilingual settings, where typically native speakers are needed to annotate data. Firstly, acquiring human experts fluent in several languages is challenging itself. Secondly, organizing and consolidating annotation process (including post-annotation verification) is a complex task. Considering also our own experience working with such datasets, the provided labels cannot be easily verified and many times we identified (inevitable) incorrect labels. Last but not least, credibility assessment and automatic detection of credibility signals naturally happen in very dynamically evolving (online) environment. New topics and global events constantly emerge, causing significant data and concept drifts. In some cases, such drifts can cause that the existing datasets may become obsolete and non-representative. Secondly, the list of credibility signals itself evolves. We can take a machine-generated text as an illustrative example. This kind of signal become highly relevant only recently with generative LLMs becoming easily available for a large end user base (and unfortunately also bad actors). Such new/redefined signals thus naturally result into a demand for new datasets. Finally, the dynamics of this area also lies in its adversarial character. Bad actors (e.g., the ones who are spreading propaganda or disinformation) will always try to get their content undetected as low-credible one by employing various obfuscation techniques. This must be taken into consideration when introducing new datasets. By continuing with an illustrative example of machine-generated text, there are already datasets (e.g., [ 198 ]) providing ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:53 besides machine-generated texts also their alternative versions after applying several authorship obfuscation techniques, which allows to train and evaluate more robust detection models. 9.4 Ethical and legal issues Credibility assessment of online (primarily social media) content is from its nature an area that must address several ethical and legal issues. First, such ethical and legal issues are especially prominent when the researched outcomes get deployed and used in practice by end users. The challenge for tool makers is to be fully transparent about the limitations of their tools, to provide guidelines to avoid misleading their users (e.g., by false positives and hallucinations) and to support human control (in line with the human-in-the-loop approach), for more details see Section 9.5. Second, similarly to other related research areas (e.g., false information detection), credibility assessment methods may be potentially misused by bad actors in order to create content that appears to be more credible. Also by following open-research spirit, publishing credibility assessment systems can theoretically result into misusing such system in the adversarial manner to tune disinformation/propaganda generation systems and allowing them to stay undetected. This kind of potential threat is, however, an analogical issue to the security domain and the principle of security by obscurity. Nothing prevents bad actors to develop their own credibility assessments systems and apply then in adversarial training scenario. Moreover, positive outcomes of credibility assessment research are more tangible, with many practical tools (as also showed across this survey) already put into the hands of media professionals or general public. Besides potential misuse, additional ethical considerations must be addressed thoroughly during the research activities. First of all, training various classification/detection systems is inherently a subject of potential biases. Such biases can come directly from the datasets used for the training purposes (in terms of data selection, data pre-/post-processing, or data annotation itself; see Section 9.3) or can be introduced during training the models, especially when fine-tuning pre-trained LLMs that have incorporated biases by themselves (including biases between highand low-resource languages; see Section 9.2). Secondly, the authors should always clearly formulate intended use and failure cases of the trained models/deployed tools (e.g., by means of model cards). In this way, we may prevent media professionals to over-rely on the predicted (potentially incorrect) values. Fortunately, the two-step approach to credibility assessment (i.e., detect more granular credibility signals and then aggregate them) makes the whole process more transparent. Last but not least, explainability and interpretability of the models’ predictions plays a crucial role in this area, since credibility assessment must be credible itself, otherwise it would not provide expected level of trust to its end users. Besides ethical issues, the research in this area must address also several legal issues. Many of them are shared with other research works on social media. Firstly, as we have also witnessed recently, social media platforms can change their data processing policies and restrict access to their data, which may delay progress in research and development of credibility tools. Additionally, social media data must be anonymized to protect user privacy before being used as training data or for the model inference. Media organizations also impose limitations due to copyright laws, with some not permitting their content to be used in AI tool development. LLMs, especially closed LLMs such as ChatGPT and GPT-4, lack transparency regarding pre-training data and LLMs can memorize content in their pretraining data [ 155 , 214 ]. Therefore, anonymization and removal of copyrighted content is crucial to credibility tools, even when they serve as foundational or backbone models. Paradigms such as unlearning [ 63 ] or LLM editing could be potential research directions to tackle these issues. ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:54 Srba et al. 9.5 Practical deployment Deploying credibility assessment and credibility signal detection models in real-world settings presents significant challenges. For media professionals, editorial guidelines in most newsrooms and fact-checking organizations recommend treating AI-based credibility assessment or fact-checking tools as exploratory aids rather than definitive sources of evidence. Fact-checkers are generally advised to use AI-generated outputs as complementary cues – only after accumulating sufficient independent evidence pointing to the falsity of the content in question. A similar critical challenge concerns the general public, who often lacks the expertise to critically assess AI-detected credibility signals [ 189 ]. This poses serious risks – if such indicators are incorrect or poorly communicated, users may be misled, potentially reinforcing belief in or further disseminating false information. These practical concerns stem from multiple underlying (not only purely technological) challenges. First, explainability, interpretability, and transparent communication of model confidence and limitations are essential components of any credibility assessment system intended for real-world use for both types of end users – media professionals as well as the general public. Many classifiers (e.g., on subjectivity or persuasion techniques) are providing clues that the editor can proofread, maintaining editorial control on the analysis. At the same time, the epistemic shift of the AIgenerated content is that humans are now struggling to determine if the content is synthetic or not, and therefore tend to rely more on automated tools, on which exercising editorial control is much more complicated. How can an editor trust an AI-based detector that has an unexplained and non-negligible known rate of false positives? How can she take the reputation risk of writing that the content is synthetic, especially in the context of political life, elections, and public debate, if it proves to be authentic later on? To this end, end users must be able to understand the strengths and weaknesses of models, and model providers must disclose the constraints and limits of the tools they provide (including information on the provenance of the datasets and the training of the models). The previous research showed that appropriate visual explanations foster end users’ trust in AI-predicted classification labels [ 358 ]. In this direction, Przybyła and Soto [253] proposed a set of interactive visualizations designed to explain the rationale behind automatic credibility assessments, with the aim of increasing users’ confidence in the underlying methods. In a user study involving 14 participants, the authors found that the stylometric classifier was perceived as more interpretable than the neural classifier, although the latter achieved higher predictive performance. Participants were also significantly more accurate and more confident in their credibility judgments after interacting with the visual explanations. These findings underscore that providing meaningful explanations and interpretability – beyond a simple black-box credibility label or score – remains an open challenge. Moreover, the study highlights that addressing this issue is non-trivial, particularly given that state-of-the-art models such as deep neural networks often suffer from limited inherent explainability, despite their high performance. Another critical challenge lies in the communication gap between technical practitioners and end users, particularly media professionals and the general public. These groups often operate using distinct terminologies and conceptual frameworks, which can hinder the interpretation of AI-generated outputs or the understanding of AI system limitations. Furthermore, many existing datasets and benchmarks have been developed primarily by computational linguists or computer scientists, without active involvement of end users. This can lead to divergent perspectives regarding what constitutes credible content or when a credibility signal is present. As our survey shows, certain credibility signal detection tasks, such as bias detection (Section 5) and check-worthiness assessment (Section 7), are inherently subjective, which may further exacerbate these discrepancies. ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:55 Consequently, model predictions may align poorly with end users’ expectations or only succeed on simple unambiguous examples. Second, detection models making use of credibility signals should minimize false positive rates as much as possible for integration in media organisations and journalism toolkits. However, based on both literature and our own practical experience, we observe that model performance often deteriorates when transitioning from offline evaluations to real-world applications. This discrepancy largely stems from the out-of-distribution (OOD) nature of deployment data, which may differ from training and testing datasets in various aspects, such as topic, format or language. These shifts highlight the need for more robust and generalizable models. However, achieving such robustness is complicated by the already-discussed limited availability of diverse datasets (Section 9.3) and the dynamic nature of online environments. Furthermore, deployed methods and tools for automatic credibility assessment are inherently exposed to potential adversarial threats. Malicious actors who intentionally disseminate non-credible information may attempt to evade detection by slightly altering the content – creating adversarial examples that exploit vulnerabilities in detection models and trigger incorrect classifications [ 252 ]. To be effective in real-world applications, automatic classifiers must therefore be robust to such adversarial manipulations. Recent studies have demonstrated that LLMs, when combined with diversity incentives techniques, can be effectively leveraged for data augmentation [ 57 ]. By generating diverse paraphrases and content variations, such approaches not only improve model generalization but also increase resilience against potential adversarial attacks – including those generated using similar LLM-based strategies. From a technical deployment perspective, resource constraints present additional barriers. While from the research perspective, open-source models are desirable for their transparency and reproducibility, hosting and running multiple models can be financially burdensome for media organizations or academic institutions. Researchers and practitioners often face a trade-off between model size and deployment feasibility: larger language models generally yield higher classification performance but require costly GPU infrastructure to support real-time inference. This necessitates compromises between computational cost and acceptable performance degradation. Techniques such as model distillation, which transfer knowledge from large models into smaller more efficient versions, may alleviate some of these challenges, but often at the cost of reduced performance. 9.6 Multimodal approaches While text is still one of the most common methods to spread information, in practice, it is commonly combined with other media types, like images, videos and audios. The demand for multimodal approaches also grows with a continuous shift of social media platforms towards short multimedia formats, such as short videos, reels or slideshows. While Natural Language Processing (NLP) allows to detect linguistic patterns that may signal low credibility content (as shown throughout this survey), it can overlook cues that are often found in such non-textual elements. For example, low-credibility content can back-up its misleading text with an altered or fully generated image or video to enhance its perceived credibility. As previously discussed in Section 8.7, there are promising results on multimodal approaches that involve processing multiple data streams simultaneously, offering a richer understanding of the content and its credibility. The main potential of these methods lies in their ability to overcome limitations inherent in text-only analysis. To this end, they combine traditional NLP methods with analysis of visual, auditive, contextual or interactive content. Even in such multimodal systems, textual analysis still remains an important component as text provides rich semantic information, often containing the explicit claims or arguments being made. Without appropriate text analysis, it would be difficult to discern the specific intent behind multimedia content. ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:56 Srba et al. In the future, there is potential to enhance multimodal models by integrating additional modalities, such as biometric signals (e.g., eye-tracking data [ 118 ]), interactive content quality, or real-time user engagement metrics. Research on multimodal credibility assessment is therefore still ongoing and provides several promising directions for future research. 10 CONCLUSIONS With the rapid development of generative AI, the potential misuse of LLMs for hybrid operations has become one of the most frequently cited risks [ 51 , 102 ]. Recent research has highlighted the particularly high capacity of LLMs to generate (personalized) disinformation across both global and local narratives [ 328 , 372 ]. In this context, the large-scale capability to automatically assess the credibility of online social media content becomes even more crucial as it was in the pre-LLM era. Potential negative impact on society and democratic values was recognized by the research community and previously led to emergence of many false information detection approaches. Unfortunately, automatic credibility assessment and the detection of credibility signals have not yet received the same level of attention. At this point, we emphasize that credibility assessment, while closely related to false information detection, provides complementary insights. In particular, false information that appears credible can have far more detrimental effects than content that is easily recognized as non-credible. Thanks to inherent two-step nature of credibility assessment (to detect credibility signals and then to aggregate them into a credibility label/score), credibility assessment encompasses a high level of explainability and wide opportunities for application in research as well as in practice. Detected high-credible content can be promoted by recommender systems or search engines, or highlighted in a user interface of social media platforms, while low-credible content can be accompanied with a low-credibility warning. With a breakdown of predicted credibility score to individual credibility indicators, end users (either media professionals or the general public) are able to explore the evidence leading to the credibility assessment or manually evaluate individual detected credibility signals and make a final assessment by themselves. Such level of explainability and pro-active human involvement in the decision process is vital and unfortunately lacks in many current false information detection approaches that commonly result into a single predicted value (commonly a binary one) with challenging or even impossible explanation caused by a black-box nature of the employed techniques (e.g., deep learning approaches). Despite these considerable advantages of credibility assessment, the current state-of-the-art research works suffer from multiple drawbacks. Most crucially, credibility assessment field can be characterized as highly fragmented. On one side, there are credibility assessment approaches that automatically detect credibility signals and aggregate them to make a final prediction about the content credibility. The utilized credibility signals are, however, mostly shallow linguistic ones (such as a number of hashtags), their automatic detection relies mostly on outdated and simple methods (like ruleor heuristic-based techniques), and also prediction utilizes mostly basic weighting schemata. There are only very few approaches that are in line with the current state of the art (deep learning, LLMs), such as [ 170 ]; or building upon the previous research results to detect more complex credibility signals, such as [258]. On the other side, there are automatic approaches detecting various categories of credibility signals, like event factuality, biases, persuasion techniques, or previously fact-checked claims. The prevalence of state-of-the-art techniques (including LLMs and various fine-tuning approaches, including PEFTs) is much higher in these works. However, such approaches remain isolated from credibility assessment, many times even not mentioning that their prediction can be considered as one of more advanced and more reliable credibility signals. ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:57 To address this undesired gap between research works, lack of interconnection of research results, as well as hindering application of the outcomes in practice, we conducted this systematic survey on automatic credibility assessment and detection of credibility signals from the NLP perspective. By collecting and describing 175 research papers, we not only systematically summarised the current state of the research in (currently fragmented) research areas, but also identified shared challenges and potential for future research. Our thorough analyses and discussions aim to point to an interesting avenues for future research – particularly, we would like to highlight three most promising research streams: • Adoption of advanced multilingual and multimodal LLMs. The emerging multilingual and multimodal LLMs offer substantial – yet still largely untapped – potential for advancing credibility assessment. Their capabilities can significantly improve both detection performance and language coverage. Moreover, these models open new avenues for deploying credibility assessment tools to support both media professionals and the general public. Such tools can serve both purposes – to detect non-credible content as well as to identify credible one that is worth further reading and sharing in social media environment. • Development of multilingual and multi-category benchmark datasets. The current lack of standardized, multilingual benchmark datasets hampers direct comparison between methods and slows progress in the field. We advocate for future work to focus on unifying existing – but often fragmented – efforts into a shared benchmark that would include diverse credibility signals annotated over the same content. Such a resource would enable novel research opportunities – such as multi-task learning or novel credibility assessment algorithms – and at the same time, would potentially reduce current redundant efforts in dataset creation. Especially existing highly multilingual datasets used for previously factchecked claim retrieval could serve as a foundation for building such a benchmark. • Addressing practical deployment challenges. Finally, bridging the gap between academic research and real-world deployment requires greater attention to practical considerations – ranging from ethical and legal implications to the robustness of models against out-ofdistribution and even adversarial inputs. Additionally, issues of computational efficiency and explainability must be addressed to ensure credibility assessment systems are both effective, trustful and reliable. Despite these challenges, existing tools already demonstrate the positive real-world impact such systems can have, both in supporting media professionals and enhancing the information ecosystem for everyday users. ACKNOWLEDGMENTS This work was partially supported by the European Union under the Horizon Europe projects: vera.ai (GA No. 101070093), AI-CODE (GA No. 101135437), and by AI4TRUST (GA No. 101070190); by the UK’s innovation agency (Innovate UK) grant 10039055; by EU NextGenerationEU through the Recovery and Resilience Plan for Slovakia under the project No. 09I03-03-V03-00020. REFERENCES [1] Athira A.b., S. D. Madhu Kumar, and Anu Mary Chacko. 2023. A Systematic Survey on Explainable AI Applied to Fake News Detection. Engineering Applications of Artificial Intelligence 122 (June 2023), 106087. https://doi.org/10. 1016/j.engappai.2023.106087 [2] Reem Abdel-Salam. 2023. rematchka at ArAIEval Shared Task: Prefix-Tuning & Prompt-tuning for Improved Detection of Propaganda and Disinformation in Arabic Social Media Content. In Proceedings of ArabicNLP 2023, Hassan Sawaf, Samhaa El-Beltagy, Wajdi Zaghouani, Walid Magdy, Ahmed Abdelali, Nadi Tomeh, Ibrahim Abu Farha, Nizar Habash, Salam Khalifa, Amr Keleg, Hatem Haddad, Imed Zitouni, Khalil Mrini, and Rawan Almatham (Eds.). Association for Computational Linguistics, Singapore (Hybrid), 536–542. https://doi.org/10.18653/v1/2023.arabicnlp-1.52 ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:64 Srba et al. Machinery, New York, NY, USA, 19–24. https://doi.org/10.1145/3543895.3543928 [92] Robert M Entman. 2007. Framing bias: Media in the distribution of power. Journal of communication 57, 1 (2007), 163–173. [93] Diego Esteves, Aniketh Janardhan Reddy, Piyush Chawla, and Jens Lehmann. 2018. Belittling the Source: Trustworthiness Indicators to Obfuscate Fake News on the Web. In Proceedings of the First Workshop on Fact Extraction and VERification (FEVER). 50–59. [94] Lisa Fan, Marshall White, Eva Sharma, Ruisi Su, Prafulla Kumar Choubey, Ruihong Huang, and Lu Wang. 2019. In Plain Sight: Media Bias Through the Lens of Factual Reporting. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (Eds.). Association for Computational Linguistics, Hong Kong, China, 6343–6349. https://doi.org/10.18653/v1/D19-1664 [95] KJ Kevin Feng, Nick Ritchie, Pia Blumenthal, Andy Parsons, and Amy X Zhang. 2023. Examining the Impact of Provenance-Enabled Media on Trust and Accuracy Perceptions. Proceedings of the ACM on Human-Computer Interaction 7, CSCW2 (2023), 1–42. [96] Elisabetta Fersini, Paolo Rosso, and Maria Anzovino. 2018. Overview of the Task on Automatic Misogyny Identification at IberEval 2018. In IberEval@SEPLN (CEUR Workshop Proceedings, Vol. 2150). CEUR-WS.org, 214–228. [97] Raphael Antonius Frick, Inna Vogel, and Jeong-Eun Choi. 2023. Fraunhofer SIT at CheckThat!-2023: Enhancing the Detection of Multimodal and Multigenre Check-Worthiness Using Optical Character Recognition and Model Souping.. In CLEF (Working Notes). 337–350. [98] Norbert Fuhr, Anastasia Giachanou, Gregory Grefenstette, Iryna Gurevych, Andreas Hanselowski, Kalervo Jarvelin, Rosie Jones, YiquN Liu, Josiane Mothe, Wolfgang Nejdl, Isabella Peters, and Benno Stein. 2018. An Information Nutritional Label for Online Documents. SIGIR Forum 51, 3 (Feb. 2018), 46–66. https://doi.org/10.1145/3190580.3190588 [99] Qin Gao, Ye Tian, and Mengyuan Tu. 2015. Exploring factors influencing Chinese user’s perceived credibility of health and safety information on Weibo. Computers in Human Behavior 45 (2015), 21–31. https://doi.org/10.1016/j. chb.2014.11.071 [100] Pepa Gencheva, Preslav Nakov, Lluís Màrquez, Alberto Barrón-Cedeño, and Ivan Koychev. 2017. A Context-Aware Approach for Detecting Worth-Checking Claims in Political Debates. In Proceedings of the International Conference Recent Advances in Natural Language Processing, RANLP 2017, Ruslan Mitkov and Galia Angelova (Eds.). INCOMA Ltd., Varna, Bulgaria, 267–276. https://doi.org/10.26615/978-954-452-049-6_037 [101] Erfan Ghadery, Damien Sileo, and Marie-Francine Moens. 2021. LIIR at SemEval-2021 task 6: Detection of Persuasion Techniques In Texts and Images using CLIP features. In Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021), Alexis Palmer, Nathan Schneider, Natalie Schluter, Guy Emerson, Aurelie Herbelot, and Xiaodan Zhu (Eds.). Association for Computational Linguistics, Online, 1015–1019. https://doi.org/10.18653/v1/2021. semeval-1.139 [102] Josh A. Goldstein, Girish Sastry, Micah Musser, Renee DiResta, Matthew Gentzel, and Katerina Sedova. 2023. Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations. https://doi.org/10.48550/arXiv.2301.04246 arXiv:2301.04246 [cs]. [103] Genevieve Gorrell, Elena Kochkina, Maria Liakata, Ahmet Aker, Arkaitz Zubiaga, Kalina Bontcheva, and Leon Derczynski. 2019. Semeval-2019 task 7: Rumoureval 2019: Determining rumour veracity and support for rumours. In Proceedings of the 13th International Workshop on Semantic Evaluation: NAACL HLT 2019. Association for Computational Linguistics, 845–854. [104] Katarína Greškovičová, Radomír Masaryk, Nikola Synak, and Vladimíra Čavojová. 2022. Superlatives, clickbaits, appeals to authority, poor grammar, or boldface: Is editorial style related to the credibility of online health messages? Frontiers in Psychology 13 (2022). https://doi.org/10.3389/fpsyg.2022.940903 [105] Jack Grieve and Helena Woodfield. 2023. The Language of Fake News. Cambridge University Press. [106] Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A survey on automated fact-checking. Transactions of the Association for Computational Linguistics 10 (2022), 178–206. [107] Ashim Gupta and Vivek Srikumar. 2021. X-Fact: A New Benchmark Dataset for Multilingual Fact Checking. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). Association for Computational Linguistics, Online, 675–682. https://doi.org/10.18653/v1/2021.aclshort.86 [108] Kshitij Gupta, Devansh Gautam, and Radhika Mamidi. 2021. Volta at SemEval-2021 Task 6: Towards Detecting Persuasive Texts and Images using Textual and Multimodal Ensemble. In Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021), Alexis Palmer, Nathan Schneider, Natalie Schluter, Guy Emerson, Aurelie Herbelot, and Xiaodan Zhu (Eds.). Association for Computational Linguistics, Online, 1075–1081. https: //doi.org/10.18653/v1/2021.semeval-1.149 ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:65 [109] Shreya Gupta, Parantak Singh, Megha Sundriyal, Md. Shad Akhtar, and Tanmoy Chakraborty. 2021. LESA: Linguistic Encapsulation and Semantic Amalgamation Based Generalised Claim Detection from Online Content. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, Paola Merlo, Jorg Tiedemann, and Reut Tsarfaty (Eds.). Association for Computational Linguistics, Online, 3178–3188. https://doi.org/10.18653/v1/2021.eacl-main.277 [110] Vansh Gupta and Raksha Sharma. 2021. NLPIITR at SemEval-2021 Task 6: RoBERTa Model with Data Augmentation for Persuasion Techniques Detection. In Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021), Alexis Palmer, Nathan Schneider, Natalie Schluter, Guy Emerson, Aurelie Herbelot, and Xiaodan Zhu (Eds.). Association for Computational Linguistics, Online, 1061–1067. https://doi.org/10.18653/v1/2021.semeval-1.147 [111] Neal Haddaway, Matthew Page, Chris Pritchard, and Luke McGuinness. 2022. PRISMA2020: An R package and Shiny app for producing PRISMA 2020-compliant flow diagrams, with interactivity for optimised digital transparency and Open Synthesis. Campbell Systematic Reviews 18 (03 2022). https://doi.org/10.1002/cl2.1230 [112] Mohamed Seghir Hadj Ameur and Hassina Aliane. 2021. AraCOVID19-MFH: Arabic COVID-19 Multi-label Fake News & Hate Speech Detection Dataset. Procedia Computer Science 189 (2021), 232–241. https://doi.org/10.1016/j. procs.2021.05.086 AI in Computational Linguistics. [113] Felix Hamborg, Norman Meuschke, and Bela Gipp. 2017. Matrix-based news aggregation: exploring different news perspectives. In Proceedings of the 17th ACM/IEEE Joint Conference on Digital Libraries (Toronto, Ontario, Canada) (JCDL ’17). IEEE Press, 69–78. https://doi.org/10.1109/JCDL.2017.7991561 [114] Felix Hamborg, Anastasia Zhukova, and Bela Gipp. 2020. Automated identification of media bias by word choice and labeling in news articles. In Proceedings of the 18th Joint Conference on Digital Libraries (Champaign, Illinois) (JCDL ’19). IEEE Press, 196–205. https://doi.org/10.1109/JCDL.2019.00036 [115] Hugo Lewi Hammer, Per Erik Solberg, and Lilja Øvrelid. 2014. Sentiment classification of online political discussions: a comparison of a word-based and dependency-based method. In Proceedings of the 5th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, Alexandra Balahur, Erik van der Goot, Ralf Steinberger, and Andres Montoyo (Eds.). Association for Computational Linguistics, Baltimore, Maryland, 90–96. https://doi.org/ 10.3115/v1/W14-2616 [116] Sakshini Hangloo and Bhavna Arora. 2022. Combating multimodal fake news on social media: methods, datasets, and future perspective. Multimedia Systems 28, 8 (2022), 2391–2422. https://doi.org/10.1007/s00530-022-00966-y [117] Abhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Goldstein. 2024. Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text. arXiv:2401.12070 [cs.CL] https://arxiv.org/abs/2401.12070 [118] Christian Hansen, Casper Hansen, Jakob Grue Simonsen, Birger Larsen, Stephen Alstrup, and Christina Lioma. 2020. Factuality Checking in News Headlines with Eye Tracking. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (Virtual Event, China) (SIGIR ’20). Association for Computing Machinery, New York, NY, USA, 2013–2016. https://doi.org/10.1145/3397271.3401221 [119] Fatima Haouari, Maram Hasanain, Reem Suwaileh, and Tamer Elsayed. 2021. ArCOV19-Rumors: Arabic COVID-19 Twitter Dataset for Misinformation Detection. In Proceedings of the Sixth Arabic Natural Language Processing Workshop, Nizar Habash, Houda Bouamor, Hazem Hajj, Walid Magdy, Wajdi Zaghouani, Fethi Bougares, Nadi Tomeh, Ibrahim Abu Farha, and Samia Touileb (Eds.). Association for Computational Linguistics, Kyiv, Ukraine (Virtual), 72–81. https://aclanthology.org/2021.wanlp-1.8 [120] Momchil Hardalov, Anton Chernyavskiy, Ivan Koychev, Dmitry Ilvovsky, and Preslav Nakov. 2022. CrowdChecked: Detecting Previously Fact-Checked Claims in Social Media. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Yulan He, Heng Ji, Sujian Li, Yang Liu, and Chua-Hui Chang (Eds.). Association for Computational Linguistics, Online only, 266–285. https://aclanthology.org/2022.aacl-main.22 [121] Robert Harris. 1997. Evaluating Internet Research Sources. http://www.virtualsalt.com/evaluating-internet-researchsources/. Accessed: 2024-10-01. [122] Maram Hasanain, Firoj Alam, Hamdy Mubarak, Samir Abdaljalil, Wajdi Zaghouani, Preslav Nakov, Giovanni Da San Martino, and Abed Freihat. 2023. ArAIEval Shared Task: Persuasion Techniques and Disinformation Detection in Arabic Text. In Proceedings of ArabicNLP 2023, Hassan Sawaf, Samhaa El-Beltagy, Wajdi Zaghouani, Walid Magdy, Ahmed Abdelali, Nadi Tomeh, Ibrahim Abu Farha, Nizar Habash, Salam Khalifa, Amr Keleg, Hatem Haddad, Imed Zitouni, Khalil Mrini, and Rawan Almatham (Eds.). Association for Computational Linguistics, Singapore (Hybrid), 483–493. https://doi.org/10.18653/v1/2023.arabicnlp-1.44 [123] Naeemul Hassan, Fatma Arslan, Chengkai Li, and Mark Tremayne. 2017. Toward Automated Fact-Checking: Detecting Check-worthy Factual Claims by ClaimBuster. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (Halifax, NS, Canada) (KDD ’17). Association for Computing Machinery, New York, NY, USA, 1803–1812. https://doi.org/10.1145/3097983.3098131 ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:66 Srba et al. [124] Naeemul Hassan, Chengkai Li, and Mark Tremayne. 2015. Detecting Check-worthy Factual Claims in Presidential Debates. In Proceedings of the 24th ACM International on Conference on Information and Knowledge Management (Melbourne, Australia) (CIKM ’15). Association for Computing Machinery, New York, NY, USA, 1835–1838. https: //doi.org/10.1145/2806416.2806652 [125] Hans J. G. Hassell, John B. Holbein, and Matthew R. Miles. 2020. There is no liberal media bias in which news stories political journalists choose to cover. Science Advances 6 (2020). https://api.semanticscholar.org/CorpusID:215516592 [126] Sandro Hawke. [n.d.]. Technological Approaches to Improving Credibility Assessment on the Web: Final Community Group Report 11 October 2018. https://www.w3.org/2018/10/credibility-tech/. [127] Lea Hellmueller and Damian Trilling. 2012. The Credibility of Credibility Measures: A Meta-Analysis in Leading Communication Journals, 1951 to 2011. In World Association for Public Opinion Research (WAPOR). Hongkong. [128] Hendrik Heuer and Elena L. Glassman. 2024. Reliability Criteria for News Websites. ACM Trans. Comput.-Hum. Interact. 31, 2, Article 21 (jan 2024), 33 pages. https://doi.org/10.1145/3635147 [129] Tashin Hossain, Jannatun Naim, Fareen Tasneem, Radiathun Tasnia, and Abu Nowshed Chy. 2021. CSECU-DSG at SemEval-2021 Task 6: Orchestrating Multimodal Neural Architectures for Identifying Persuasion Techniques in Texts and Images. In Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021), Alexis Palmer, Nathan Schneider, Natalie Schluter, Guy Emerson, Aurelie Herbelot, and Xiaodan Zhu (Eds.). Association for Computational Linguistics, Online, 1088–1095. https://doi.org/10.18653/v1/2021.semeval-1.151 [130] Carl I. Hovland and Walter Weiss. 1951. The Influence of Source Credibility on Communication Effectiveness*. Public Opinion Quarterly 15, 4 (01 1951), 635–650. https://doi.org/10.1086/266350 arXiv:https://academic.oup.com/poq/articlepdf/15/4/635/5422100/15-4-635.pdf [131] Timo Hromadka, Timotej Smolen, Tomas Remis, Branislav Pecher, and Ivan Srba. 2023. KInITVeraAI at SemEval2023 Task 3: Simple yet Powerful Multilingual Fine-Tuning for Persuasion Techniques Detection. In Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), Atul Kr. Ojha, A. Seza Doğruöz, Giovanni Da San Martino, Harish Tayyar Madabushi, Ritesh Kumar, and Elisa Sartori (Eds.). Association for Computational Linguistics, Toronto, Canada, 629–637. https://doi.org/10.18653/v1/2023.semeval-1.86 [132] Bo Hu, Zhendong Mao, and Yongdong Zhang. 2024. An Overview of Fake News Detection: From a New Perspective. Fundamental Research (Feb. 2024). https://doi.org/10.1016/j.fmre.2024.01.017 [133] Dongchen Huang, Yige Zhu, and Eni Mustafaraj. 2019. How Dependable are "First Impressions" to Distinguish between Real and Fake NewsWebsites?. In Proceedings of the 30th ACM Conference on Hypertext and Social Media (Hof, Germany) (HT ’19). Association for Computing Machinery, New York, NY, USA, 201–210. https://doi.org/10. 1145/3342220.3343670 [134] Kung-Hsiang Huang, ChengXiang Zhai, and Heng Ji. 2022. CONCRETE: Improving Cross-lingual Fact-checking with Cross-lingual Retrieval. In Proceedings of the 29th International Conference on Computational Linguistics, Nicoletta Calzolari, Chu-Ren Huang, Hansaem Kim, James Pustejovsky, Leo Wanner, Key-Sun Choi, Pum-Mo Ryu, Hsin-Hsi Chen, Lucia Donatelli, Heng Ji, Sadao Kurohashi, Patrizia Paggio, Nianwen Xue, Seokhwan Kim, Younggyun Hahm, Zhong He, Tony Kyungil Lee, Enrico Santus, Francis Bond, and Seung-Hoon Na (Eds.). International Committee on Computational Linguistics, Gyeongju, Republic of Korea, 1024–1035. https://aclanthology.org/2022.coling-1.86 [135] Emelia May Hughes, Renee Wang, Prerna Juneja, Tony W Li, Tanushree Mitra, and Amy X. Zhang. 2024. Viblio: Introducing Credibility Signals and Citations to Video-Sharing Platforms. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 807, 20 pages. https://doi.org/10.1145/3613904.3642490 [136] Ahmed Samir Hussein, Abu Bakr Soliman Mohammad, Mohamed Ibrahim, Laila Hesham Afify, and Samhaa R. El-Beltagy. 2022. NGU CNLP atWANLP 2022 Shared Task: Propaganda Detection in Arabic. In Proceedings of the Seventh Arabic Natural Language Processing Workshop (WANLP), Houda Bouamor, Hend Al-Khalifa, Kareem Darwish, Owen Rambow, Fethi Bougares, Ahmed Abdelali, Nadi Tomeh, Salam Khalifa, and Wajdi Zaghouani (Eds.). Association for Computational Linguistics, Abu Dhabi, United Arab Emirates (Hybrid), 545–550. https: //doi.org/10.18653/v1/2022.wanlp-1.66 [137] Hasan Iqbal, Yuxia Wang, Minghan Wang, Georgi Georgiev, Jiahui Geng, Iryna Gurevych, and Preslav Nakov. 2024. OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs. arXiv preprint arXiv:2408.11832 (2024). [138] Petar Ivanov, Ivan Koychev, Momchil Hardalov, and Preslav Nakov. 2024. Detecting Check-Worthy Claims in Political Debates, Speeches, and Interviews Using Audio Data. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 12011–12015. https://doi.org/10.1109/ICASSP48485.2024.10447064 [139] Farnaz Jahanbakhsh and David R Karger. 2024. A Browser Extension for in-place Signaling and Assessment of Misinformation. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 946, 21 pages. https://doi.org/10.1145/ 3613904.3642473 ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:67 [140] Farnaz Jahanbakhsh, Amy X. Zhang, and David R. Karger. 2022. Leveraging Structured Trusted-Peer Assessments to Combat Misinformation. Proc. ACM Hum.-Comput. Interact. 6, CSCW2, Article 524 (nov 2022), 40 pages. https: //doi.org/10.1145/3555637 [141] Israa Jaradat, Pepa Gencheva, Alberto Barrón-Cedeño, Lluís Màrquez, and Preslav Nakov. 2018. ClaimRank: Detecting Check-Worthy Claims in Arabic and English. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations, Yang Liu, Tim Paek, and Manasi Patwardhan (Eds.). Association for Computational Linguistics, New Orleans, Louisiana, 26–30. https://doi.org/10.18653/v1/N18-5006 [142] Israa Jaradat, Haiqi Zhang, and Chengkai Li. 2024. On Detecting Cherry-picking in News Coverage Using Large Language Models. arXiv e-prints (2024), arXiv–2401. [143] Ganesh Jawahar, Muhammad Abdul-Mageed, and Laks Lakshmanan, V.S. 2020. Automatic Detection of Machine Generated Text: A Critical Survey. In Proceedings of the 28th International Conference on Computational Linguistics, Donia Scott, Nuria Bel, and Chengqing Zong (Eds.). International Committee on Computational Linguistics, Barcelona, Spain (Online), 2296–2309. https://doi.org/10.18653/v1/2020.coling-main.208 [144] Caio LM Jeronimo, Claudio EC Campelo, Leandro Balby Marinho, Allan Sales, Adriano Veloso, and Roberta Viola. 2020. Computing with subjectivity lexicons. In Proceedings of the Twelfth Language Resources and Evaluation Conference. 3272–3280. [145] Jiaojiao Ji, Yuqi Zhu, and Naipeng Chao. 2023. A comparison of misinformation feature effectiveness across issues and time on Chinese social media. Information Processing & Management 60, 2 (2023), 103210. https://doi.org/10. 1016/j.ipm.2022.103210 [146] Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of Hallucination in Natural Language Generation. ACM Comput. Surv. 55, 12, Article 248 (March 2023), 38 pages. https://doi.org/10.1145/3571730 [147] Chenyan Jia, Alexander Boltz, Angie Zhang, Anqing Chen, and Min Kyung Lee. 2022. Understanding Effects of Algorithmic vs. Community Label on Perceived Accuracy of Hyper-partisan Misinformation. Proc. ACM Hum.-Comput. Interact. 6, CSCW2, Article 371 (nov 2022), 27 pages. https://doi.org/10.1145/3555096 [148] Ye Jiang, Xingyi Song, Carolina Scarton, Iknoor Singh, Ahmet Aker, and Kalina Bontcheva. 2023. Categorising Fine-to-Coarse Grained Misinformation: An Empirical Study of the COVID-19 Infodemic. In Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing, Ruslan Mitkov and Galia Angelova (Eds.). INCOMA Ltd., Shoumen, Bulgaria, Varna, Bulgaria, 556–567. https://aclanthology.org/2023.ranlp-1.61 [149] Yiping Jin, Leo Wanner, and Alexander Shvets. 2024. GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue (Eds.). ELRA and ICCL, Torino, Italia, 7867–7885. https://aclanthology.org/2024.lrecmain.694/ [150] Prerna Juneja, Wenjuan Zhang, Alison Marie Smith-Renner, Hemank Lamba, Joel Tetreault, and Alex Jaimes. 2024. Dissecting users’ needs for search result explanations. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 841, 17 pages. https://doi.org/10.1145/3613904.3642059 [151] Dawid Jurkiewicz, Łukasz Borchmann, Izabela Kosmala, and Filip Graliński. 2020. ApplicaAI at SemEval-2020 Task 11: On RoBERTa-CRF, Span CLS and Whether Self-Training Helps Them. In Proceedings of the Fourteenth Workshop on Semantic Evaluation, Aurelie Herbelot, Xiaodan Zhu, Alexis Palmer, Nathan Schneider, Jonathan May, and Ekaterina Shutova (Eds.). International Committee for Computational Linguistics, Barcelona (online), 1415–1424. https://doi.org/10.18653/v1/2020.semeval-1.187 [152] Michal Kakol, Radoslaw Nielek, and Adam Wierzbicki. 2017. Understanding and predicting web content credibility using the content credibility corpus. Information Processing & Management 53, 5 (2017), 1043–1061. [153] Byungkyu Kang, John O’Donovan, and Tobias Höllerer. 2012. Modeling topic specific credibility on twitter. In Proceedings of the 2012 ACM International Conference on Intelligent User Interfaces (Lisbon, Portugal) (IUI ’12). Association for Computing Machinery, New York, NY, USA, 179–188. https://doi.org/10.1145/2166966.2166998 [154] Uku Kangur, Roshni Chakraborty, and Rajesh Sharma. 2024. Who Checks the Checkers? Exploring Source Credibility in Twitter’s Community Notes. arXiv preprint arXiv:2406.12444 (2024). [155] Antonia Karamolegkou, Jiaang Li, Li Zhou, and Anders Søgaard. 2023. Copyright Violations and Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 7403–7412. [156] Yavuz Selim Kartal and Mucahid Kutlu. 2020. TrClaim-19: The First Collection for Turkish Check-Worthy Claim Detection with Annotator Rationales. In Proceedings of the 24th Conference on Computational Natural Language Learning, Raquel Fernández and Tal Linzen (Eds.). Association for Computational Linguistics, Online, 386–395. https://doi.org/10.18653/v1/2020.conll-1.31 ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:68 Srba et al. [157] Ashkan Kazemi, Artem Abzaliev, Naihao Deng, Rui Hou, Scott Hale, Veronica Perez-Rosas, and Rada Mihalcea. 2023. Query Rewriting for Effective Misinformation Discovery. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), Jong C. Park, Yuki Arase, Baotian Hu, Wei Lu, Derry Wijaya, Ayu Purwarianti, and Adila Alfa Krisnadhi (Eds.). Association for Computational Linguistics, Nusa Dua, Bali, 398–407. https://doi.org/ 10.18653/v1/2023.ijcnlp-main.26 [158] Ashkan Kazemi, Kiran Garimella, Devin Gaffney, and Scott Hale. 2021. Claim Matching Beyond English to Scale Global Fact-Checking. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). Association for Computational Linguistics, Online, 4504–4517. https://doi.org/10.18653/v1/2021.acl-long.347 [159] Sopan Khosla, Rishabh Joshi, Ritam Dutt, Alan W Black, and Yulia Tsvetkov. 2020. LTIatCMU at SemEval-2020 Task 11: Incorporating Multi-Level Features for Multi-Granular Propaganda Span Identification. In Proceedings of the Fourteenth Workshop on Semantic Evaluation, Aurelie Herbelot, Xiaodan Zhu, Alexis Palmer, Nathan Schneider, Jonathan May, and Ekaterina Shutova (Eds.). International Committee for Computational Linguistics, Barcelona (online), 1756–1763. https://doi.org/10.18653/v1/2020.semeval-1.230 [160] Carita Kiili, Leena Laurinen, and Miika Marttunen. 2007. How students evaluate credibility and relevance of information on the internet? IADIS International Conference on Cognition and Exploratory Learning in Digital Age, CELDA 2007 (01 2007). [161] Moonsung Kim and Steven Bethard. 2020. TTUI at SemEval-2020 Task 11: Propaganda Detection with Transfer Learning and Ensembles. In Proceedings of the Fourteenth Workshop on Semantic Evaluation, Aurelie Herbelot, Xiaodan Zhu, Alexis Palmer, Nathan Schneider, Jonathan May, and Ekaterina Shutova (Eds.). International Committee for Computational Linguistics, Barcelona (online), 1829–1834. https://doi.org/10.18653/v1/2020.semeval-1.240 [162] Spiro Kiousis. 2001. Public Trust or Mistrust? Perceptions of Media Credibility in the Information Age. Mass Communication and Society 4, 4 (Nov. 2001), 381–403. https://doi.org/10.1207/S15327825MCS0404_4 [163] V. Klemperer and M. Brady. 2006. Language of the Third Reich: LTI: Lingua Tertii Imperii. Bloomsbury Academic. https://books.google.sk/books?id=kwsleqxx_SMC [164] Akihiro Kondo and Hironobu Abe. 2023. Travel blogger credibility scoring method for recommendation system. In 2023 17th International Conference on Ubiquitous Information Management and Communication (IMCOM). 1–4. https://doi.org/10.1109/IMCOM56909.2023.10035643 [165] Jan-David Krieger, Timo Spinde, Terry Ruas, Juhi Kulshrestha, and Bela Gipp. 2022. A domain-adaptive pre-training approach for language bias detection in news. In Proceedings of the 22nd ACM/IEEE Joint Conference on Digital Libraries (Cologne, Germany) (JCDL ’22). Association for Computing Machinery, New York, NY, USA, Article 3, 7 pages. https://doi.org/10.1145/3529372.3530932 [166] Salima Lamsiyah, Abdelkader El Mahdaouy, Hamza Alami, Ismail Berrada, and Christoph Schommer. 2023. UL & UM6P at ArAIEval Shared Task: Transformer-based model for Persuasion Techniques and Disinformation detection in Arabic. In Proceedings of ArabicNLP 2023, Hassan Sawaf, Samhaa El-Beltagy, Wajdi Zaghouani, Walid Magdy, Ahmed Abdelali, Nadi Tomeh, Ibrahim Abu Farha, Nizar Habash, Salam Khalifa, Amr Keleg, Hatem Haddad, Imed Zitouni, Khalil Mrini, and Rawan Almatham (Eds.). Association for Computational Linguistics, Singapore (Hybrid), 558–564. https://doi.org/10.18653/v1/2023.arabicnlp-1.55 [167] Patrick Lawson, Carl J. Pearson, Aaron Crowson, and Christopher B. Mayhorn. 2020. Email phishing and signal detection: How persuasion principles and personality influence response patterns and accuracy. Applied Ergonomics 86 (2020), 103084. https://doi.org/10.1016/j.apergo.2020.103084 [168] Konstantina Lazaridou, Ralf Krestel, and Felix Naumann. 2017. Identifying Media Bias by Analyzing Reported Speech. In 2017 IEEE International Conference on Data Mining (ICDM). 943–948. https://doi.org/10.1109/ICDM.2017.119 [169] Kenton Lee, Yoav Artzi, Yejin Choi, and Luke Zettlemoyer. 2015. Event Detection and Factuality Assessment with Non-Expert Supervision. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Lluís Màrquez, Chris Callison-Burch, and Jian Su (Eds.). Association for Computational Linguistics, Lisbon, Portugal, 1643–1648. https://doi.org/10.18653/v1/D15-1189 [170] João A. Leite, Olesya Razuvayevskaya, Kalina Bontcheva, and Carolina Scarton. 2024. Weakly Supervised Veracity Classification with LLM-Predicted Credibility Signals. arXiv:2309.07601 [cs.CL] https://arxiv.org/abs/2309.07601 [171] Elisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini, and Sara Tonelli. 2021. Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators’ Disagreement. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 10528–10539. [172] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:69 R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 9459–9474. https://proceedings.neurips.cc/ paper_files/paper/2020/file/6b493230205f780e1bc26945df7481e5-Paper.pdf [173] Chang Li and Dan Goldwasser. 2021. MEAN: Multi-head Entity Aware Attention Networkfor Political Perspective Detection in News Media. In Proceedings of the Fourth Workshop on NLP for Internet Freedom: Censorship, Disinformation, and Propaganda, Anna Feldman, Giovanni Da San Martino, Chris Leberknight, and Preslav Nakov (Eds.). Association for Computational Linguistics, Online, 66–75. https://doi.org/10.18653/v1/2021.nlp4if-1.10 [174] Chunyang Li, Hao Peng, Xiaozhi Wang, Yunjia Qi, Lei Hou, Bin Xu, and Juanzi Li. 2024. MAVEN-Fact: A Large-scale Event Factuality Detection Dataset. arXiv preprint arXiv:2407.15352 (2024). [175] Xiaojia Li, Yun Zhang, and Zhong Qian. 2020. An End-to-End Approach for Document-level Event Factuality Identification in Chinese. In 2020 International Conference on Asian Language Processing (IALP). 215–220. https: //doi.org/10.1109/IALP51396.2020.9310484 [176] Yichuan Li, Bohan Jiang, Kai Shu, and Huan Liu. 2020. MM-COVID: A Multilingual and Multimodal Data Repository for Combating COVID-19 Disinformation. arXiv:2011.04088 [cs.SI] https://arxiv.org/abs/2011.04088 [177] Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi, and Yue Zhang. 2024. MAGE: Machine-generated Text Detection in the Wild. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 36–53. https://doi.org/10.18653/v1/2024.acl-long.3 [178] Yufeng Li, Rrubaa Panchendrarajan, and Arkaitz Zubiaga. 2024. FactFinders at CheckThat! 2024: Refining Checkworthy Statement Detection with LLMs through Data Pruning. In Working Notes of the Conference and Labs of the Evaluation Forum (CLEF 2024), Grenoble, France, 9-12 September, 2024 (CEUR Workshop Proceedings, Vol. 3740), Guglielmo Faggioli, Nicola Ferro, Petra Galuscáková, and Alba García Seco de Herrera (Eds.). CEUR-WS.org, 520–537. https://ceur-ws.org/Vol-3740/paper-47.pdf [179] Chenghua Lin, Yulan He, and Richard Everson. 2011. Sentence Subjectivity Detection with Weakly-Supervised Learning. In Proceedings of 5th International Joint Conference on Natural Language Processing, Haifeng Wang and David Yarowsky (Eds.). Asian Federation of Natural Language Processing, Chiang Mai, Thailand, 1153–1161. https: //aclanthology.org/I11-1129 [180] Luyang Lin, Lingzhi Wang, Jinsong Guo, and Kam-Fai Wong. 2024. Investigating bias in llm-based bias detection: Disparities between llms and human perception. arXiv preprint arXiv:2403.14896 (2024). [181] Luyang Lin, Lingzhi Wang, Xiaoyan Zhao, Jing Li, and Kam-Fai Wong. 2024. IndiVec: An Exploration of Leveraging Large Language Models for Media Bias Detection with Fine-Grained Bias Indicators. In Findings of the Association for Computational Linguistics: EACL 2024, Yvette Graham and Matthew Purver (Eds.). Association for Computational Linguistics, St. Julian’s, Malta, 1038–1050. https://aclanthology.org/2024.findings-eacl.70 [182] Szu-Yin Lin, Yun-Ching Kung, and Fang-Yie Leu. 2022. Predictive intelligence in harmful news identification by BERT-based ensemble learning model with text sentiment analysis. Information Processing & Management 59, 2 (2022), 102872. https://doi.org/10.1016/j.ipm.2022.102872 [183] Xialing Lin, Patric R. Spence, and Kenneth A. Lachlan. 2016. Social media and credibility indicators: The effect of influence cues. Computers in Human Behavior 63 (2016), 264–271. https://doi.org/10.1016/j.chb.2016.05.002 [184] Eunice Ling, Domenico de Pieri, Evenne Loh, Karen M Scott, Stephen C H Li, and Heather J Medbury. 2024. Evaluation of the Accuracy, Credibility, and Readability of Statin-Related Websites: Cross-Sectional Study. Interact J Med Res 13 (14 Mar 2024), e42849. https://doi.org/10.2196/42849 [185] Siyi Liu, Lei Guo, Kate Mays, Margrit Betke, and Derry Tanti Wijaya. 2019. Detecting Frames in News Headlines and Its Application to Analyzing News Framing Trends Surrounding U.S. Gun Violence. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), Mohit Bansal and Aline Villavicencio (Eds.). Association for Computational Linguistics, Hong Kong, China, 504–514. https://doi.org/10.18653/v1/K19-1047 [186] Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR abs/1907.11692 (2019). arXiv:1907.11692 http://arxiv.org/abs/1907.11692 [187] Hao Lu, Yanwu Yang, Runsheng Gan, and Nan Zhang. 2012. The research on micro-blog public opinion index and the application of prototype system. In Proceedings of 2012 9th IEEE International Conference on Networking, Sensing and Control. 405–410. https://doi.org/10.1109/ICNSC.2012.6204953 [188] Qiang Lu, Yunfei Long, Xia Sun, Jun Feng, and Hao Zhang. 2024. Fact-sentiment incongruity combination network for multimodal sarcasm detection. Information Fusion 104 (2024), 102203. https://doi.org/10.1016/j.inffus.2023.102203 [189] Zhuoran Lu, Patrick Li, Weilong Wang, and Ming Yin. 2022. The effects of ai-based credibility indicators on the detection and spread of misinformation under social influence. Proceedings of the ACM on Human-Computer Interaction 6, CSCW2 (2022), 1–27. [190] Jason Lucas, Limeng Cui, Thai Le, and Dongwon Lee. 2022. Detecting False Claims in Low-Resource Regions: A Case Study of Caribbean Islands. In Proceedings of the Workshop on Combating Online Hostile Posts in Regional Languages ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:70 Srba et al. during Emergency Situations, Tanmoy Chakraborty, Md. Shad Akhtar, Kai Shu, H. Russell Bernard, Maria Liakata, Preslav Nakov, and Aseem Srivastava (Eds.). Association for Computational Linguistics, Dublin, Ireland, 95–102. https://doi.org/10.18653/v1/2022.constraint-1.11 [191] Jason Lucas, Adaku Uchendu, Michiharu Yamashita, Jooyoung Lee, Shaurya Rohatgi, and Dongwon Lee. 2023. Fighting Fire with Fire: The Dual Role of LLMs in Crafting and Detecting Elusive Disinformation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 14279–14305. https://doi.org/10.18653/v1/2023.emnlp-main.883 [192] Riccardo Lunardi, David La Barbera, and Kevin Roitero. 2024. The Elusiveness of Detecting Political Bias in Language Models. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management. 3922– 3926. [193] Emma Lurie and Eni Mustafaraj. 2018. Investigating the Effects of Google’s Search Engine Result Page in Evaluating the Credibility of Online News Sources. In Proceedings of the 10th ACM Conference on Web Science (Amsterdam, Netherlands) (WebSci ’18). Association for Computing Machinery, New York, NY, USA, 107–116. https://doi.org/10. 1145/3201064.3201095 [194] Iffat Maab, Edison Marrese-Taylor, and Yutaka Matsuo. 2023. An Effective Approach for Informational and Lexical Bias Detection. In Proceedings of the Sixth Fact Extraction and VERification Workshop (FEVER), Mubashara Akhtar, Rami Aly, Christos Christodoulopoulos, Oana Cocarascu, Zhijiang Guo, Arpit Mittal, Michael Schlichtkrull, James Thorne, and Andreas Vlachos (Eds.). Association for Computational Linguistics, Dubrovnik, Croatia, 66–77. https: //doi.org/10.18653/v1/2023.fever-1.7 [195] Fabrizio Macagno. 2022. Argumentation profiles and the manipulation of common ground. The arguments of populist leaders on Twitter. Journal of Pragmatics 191 (2022), 67–82. https://doi.org/10.1016/j.pragma.2022.01.022 [196] Dominik Macko, Jakub Kopál, Robert Moro, and Ivan Srba. 2025. MultiSocial: Multilingual Benchmark of MachineGenerated Text Detection of Social-Media Texts. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 727–752. https: //doi.org/10.18653/v1/2025.acl-long.36 [197] Dominik Macko, Robert Moro, Adaku Uchendu, Jason Lucas, Michiharu Yamashita, Matúš Pikuliak, Ivan Srba, Thai Le, Dongwon Lee, Jakub Simko, and Maria Bielikova. 2023. MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 9960–9987. https://doi.org/10.18653/v1/2023.emnlp-main.616 [198] Dominik Macko, Robert Moro, Adaku Uchendu, Ivan Srba, Jason S Lucas, Michiharu Yamashita, Nafis Irtiza Tripto, Dongwon Lee, Jakub Simko, and Maria Bielikova. 2024. Authorship Obfuscation in Multilingual Machine-Generated Text Detection. In Findings of the Association for Computational Linguistics: EMNLP 2024, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 6348–6368. https://doi.org/10.18653/v1/2024.findings-emnlp.369 [199] Isa Maks and Piek Vossen. 2012. Building a fine-grained subjectivity lexicon from a web corpus. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12), Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, and Stelios Piperidis (Eds.). European Language Resources Association (ELRA), Istanbul, Turkey, 3070–3076. http://www.lrecconf.org/proceedings/lrec2012/pdf/1018_Paper.pdf [200] Watheq Mansour, Tamer Elsayed, and Abdulaziz Al-Ali. 2022. Did I See It Before? Detecting Previously-Checked Claims over Twitter. In Advances in Information Retrieval, Matthias Hagen, Suzan Verberne, Craig Macdonald, Christin Seifert, Krisztian Balog, Kjetil Nørvåg, and Vinay Setty (Eds.). Springer International Publishing, Cham, 367–381. [201] Alejandro Martín, Javier Huertas-Tato, Álvaro Huertas-García, Guillermo Villar-Rodríguez, and David Camacho. 2022. FacTeR-Check: Semi-automated fact-checking through semantic similarity and natural language inference. Knowledge-Based Systems 251 (2022), 109265. https://doi.org/10.1016/j.knosys.2022.109265 [202] Sandra C Matz, Jacob D Teeny, Sumer S Vaid, Heinrich Peters, Gabriella M Harari, and Moran Cerf. 2024. The potential of generative AI for personalized persuasion at scale. Scientific Reports 14, 1 (2024), 4692. [203] Mohsen Mesgar and Michael Strube. 2018. A Neural Local Coherence Model for Text Quality Assessment. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). Association for Computational Linguistics, Brussels, Belgium, 4328–4339. https://doi.org/10.18653/v1/D18-1464 [204] Clyde R Miller. 1939. The techniques of propaganda. From “how to detect and analyze propaganda,” an address given at town hall. The Center for learning (1939). [205] Tanushree Mitra and Eric Gilbert. 2021. CREDBANK: A Large-Scale Social Media Corpus With Associated Credibility Annotations. Proceedings of the International AAAI Conference on Web and Social Media 9, 1 (Aug. 2021), 258–267. ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:71 https://doi.org/10.1609/icwsm.v9i1.14625 [206] Shubham Mittal and Preslav Nakov. 2022. IITD at WANLP 2022 Shared Task: Multilingual Multi-Granularity Network for Propaganda Detection. In Proceedings of the Seventh Arabic Natural Language Processing Workshop (WANLP), Houda Bouamor, Hend Al-Khalifa, Kareem Darwish, Owen Rambow, Fethi Bougares, Ahmed Abdelali, Nadi Tomeh, Salam Khalifa, and Wajdi Zaghouani (Eds.). Association for Computational Linguistics, Abu Dhabi, United Arab Emirates (Hybrid), 529–533. https://doi.org/10.18653/v1/2022.wanlp-1.63 [207] Shubham Mittal, Megha Sundriyal, and Preslav Nakov. 2023. Lost in Translation, Found in Spans: Identifying Claims in Multilingual Social Media. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 3887–3902. https://doi.org/10.18653/v1/2023.emnlp-main.236 [208] Aditya Mogadala and Vasudeva Varma. 2012. Language independent sentence-level subjectivity analysis with feature selection. In Proceedings of the 26th Pacific Asia Conference on Language, Information, and Computation. 171–180. [209] Suvendu Mohanty. 2025. Fine-Grained Bias Detection in LLM: Enhancing detection mechanisms for nuanced biases. arXiv preprint arXiv:2503.06054 (2025). [210] Maria D Molina, S Shyam Sundar, Thai Le, and Dongwon Lee. 2021. "Fake News" Is Not Simply False Information: A Concept Explication and Taxonomy of Online Content. American Behavioral Scientist 65, 2 (2021), 180–212. https://doi.org/10.1177/0002764219878224 [211] Gaku Morio, Terufumi Morishita, Hiroaki Ozaki, and Toshinori Miyoshi. 2020. Hitachi at SemEval-2020 Task 11: An Empirical Study of Pre-Trained Transformer Family for Propaganda Detection. In Proceedings of the Fourteenth Workshop on Semantic Evaluation, Aurelie Herbelot, Xiaodan Zhu, Alexis Palmer, Nathan Schneider, Jonathan May, and Ekaterina Shutova (Eds.). International Committee for Computational Linguistics, Barcelona (online), 1739–1748. https://doi.org/10.18653/v1/2020.semeval-1.228 [212] Alejandro Mosquera and Paloma Moreda. 2021. SMILE: An Informality Classification Tool for Helping to Assess Quality and Credibility in Web 2.0 Texts. Proceedings of the International AAAI Conference on Web and Social Media 6, 3 (Aug. 2021), 2–7. https://doi.org/10.1609/icwsm.v6i3.14347 [213] Ahmed Mourad and Kareem Darwish. 2013. Subjectivity and sentiment analysis of modern standard Arabic and Arabic microblogs. In Proceedings of the 4th workshop on computational approaches to subjectivity, sentiment and social media analysis. 55–64. [214] Felix B Mueller, Rebekka Görge, Anna K Bernzen, Janna C Pirk, and Maximilian Poretschkin. 2024. LLMs and Memorization: On Quality and Specificity of Copyright Compliance. arXiv preprint arXiv:2405.18492 (2024). [215] Zain Muhammad Mujahid, Dilshod Azizov, Maha Tufail Agro, and Preslav Nakov. 2025. Profiling News Media for Factuality and Bias Using LLMs and the Fact-Checking Methodology of Human Experts. arXiv preprint arXiv:2506.12552 (2025). [216] Smruthi Mukund and Rohini K Srihari. 2010. A vector space model for subjectivity classification in Urdu aided by co-training. In Coling 2010: Posters. 860–868. [217] Multiple authors and Sandro Hawke. 2019. Credibility Signals: Unofficial Draft 26 November 2019. https://credweb.org/signals-20191126. [218] Preslav Nakov, Jisun An, Haewoon Kwak, Muhammad Arslan Manzoor, Zain Muhammad Mujahid, and Husrev Taha Sencar. 2024. A Survey on Predicting the Factuality and the Bias of News Media. (Aug. 2024), 15947–15962. https://doi.org/10.18653/v1/2024.findings-acl.944 [219] Preslav Nakov, Alberto Barrón-Cedeño, Giovanni Da San Martino, Firoj Alam, Rubén Míguez, Tommaso Caselli, Mücahid Kutlu, Wajdi Zaghouani, Chengkai Li, Shaden Shaar, Hamdy Mubarak, Alex Nikolov, and Yavuz Selim Kartal. 2022. Overview of the CLEF-2022 CheckThat! Lab Task 1 on Identifying Relevant Claims in Tweets. In Proceedings of the Working Notes of CLEF 2022 - Conference and Labs of the Evaluation Forum, Bologna, Italy, September 5th - to - 8th, 2022 (CEUR Workshop Proceedings, Vol. 3180), Guglielmo Faggioli, Nicola Ferro, Allan Hanbury, and Martin Potthast (Eds.). CEUR-WS.org, 368–392. https://ceur-ws.org/Vol-3180/paper-28.pdf [220] Preslav Nakov, Giovanni Da San Martino, Firoj Alam, Shaden Shaar, Hamdy Mubarak, and Nikolay Babulkov. 2022. Overview of the CLEF-2022 CheckThat! Lab Task 2 on Detecting Previously Fact-Checked Claims. In Proceedings of the Working Notes of CLEF 2022 - Conference and Labs of the Evaluation Forum, Vol. 3180. CEUR-WS, 11. https://ceurws.org/Vol-3180/paper-29.pdf [221] Preslav Nakov, Giovanni Da San Martino, Tamer Elsayed, Alberto Barrón-Cedeño, Rubén Míguez, Shaden Shaar, Firoj Alam, Fatima Haouari, Maram Hasanain, Watheq Mansour, Bayan Hamdan, Zien Sheikh Ali, Nikolay Babulkov, Alex Nikolov, Gautam Kishore Shahi, Julia Maria Struß, Thomas Mandl, Mücahid Kutlu, and Yavuz Selim Kartal. 2021. Overview of the CLEF-2021 CheckThat! Lab on Detecting Check-Worthy Claims, Previously Fact-Checked Claims, and Fake News. In Experimental IR Meets Multilinguality, Multimodality, and Interaction - 12th International Conference of the CLEF Association, CLEF 2021, Virtual Event, September 21-24, 2021, Proceedings (Lecture Notes in Computer Science, Vol. 12880), K. Selçuk Candan, Bogdan Ionescu, Lorraine Goeuriot, Birger Larsen, Henning ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:72 Srba et al. Müller, Alexis Joly, Maria Maistro, Florina Piroi, Guglielmo Faggioli, and Nicola Ferro (Eds.). Springer, 264–291. https://doi.org/10.1007/978-3-030-85251-1_19 [222] Eric Ness, Arooj Fatima, and Mahdi Maktabdar Oghaz. 2023. Data Driven Model to Investigate Political Bias in Mainstream Media. IEEE Access 11 (2023), 41880–41893. https://doi.org/10.1109/ACCESS.2023.3270630 [223] Dan S. Nielsen and Ryan McConville. 2022. MuMiN: A Large-Scale Multilingual Multimodal Fact-Checked Misinformation Social Network Dataset. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (Madrid, Spain) (SIGIR ’22). Association for Computing Machinery, New York, NY, USA, 3141–3153. https://doi.org/10.1145/3477495.3531744 [224] Alexandra Olteanu, Stanislav Peshterliev, Xin Liu, and Karl Aberer. 2013. Web Credibility: Features Exploration and Credibility Prediction. In Advances in Information Retrieval, Pavel Serdyukov, Pavel Braslavski, Sergei O. Kuznetsov, Jaap Kamps, Stefan Rüger, Eugene Agichtein, Ilya Segalovich, and Emine Yilmaz (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 557–568. [225] Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 27730–27744. https://proceedings. neurips.cc/paper_files/paper/2022/file/b1efde53be364a73914f58805a001731-Paper-Conference.pdf [226] Laura O’Grady. 2006. Future directions for depicting credibility in health care web sites. International Journal of Medical Informatics 75, 1 (2006), 58–65. https://doi.org/10.1016/j.ijmedinf.2005.07.035 Health and the Internet for All. [227] Matthew Page, David Moher, Patrick Bossuyt, Isabelle Boutron, Tammy Hoffmann, Cynthia Mulrow, Larissa Shamseer, Jennifer Tetzlaff, Elie Akl, Sue Brennan, Roger Chou, Julie Glanville, Jeremy Grimshaw, Asbjørn Hróbjartsson, Manoj Lalu, Tianjing Li, Elizabeth Loder, Evan Mayo-Wilson, Steve Mcdonald, and Joanne Mckenzie. 2021. PRISMA 2020 explanation and elaboration: Updated guidance and exemplars for reporting systematic reviews. BMJ 372 (03 2021), n160. https://doi.org/10.1136/bmj.n160 [228] Surya Kant Pal, Omar Junior Raffik, Rita Roy, Vishwakarma Bipin Lalman, Shristee Srivastava, and Bhoomi Sharma. 2023. Automatic Plagiarism Detection Using Natural Language Processing. In 2023 10th International Conference on Computing for Sustainable Global Development (INDIACom). 218–222. [229] Rrubaa Panchendrarajan and Arkaitz Zubiaga. 2024. Claim detection for automated fact-checking: A survey on monolingual, multilingual and cross-lingual research. Natural Language Processing Journal 7 (2024), 100066. https: //doi.org/10.1016/j.nlp.2024.100066 [230] Bo Pang and Lillian Lee. 2004. A Sentimental Education: Sentiment Analysis Using Subjectivity Summarization Based on Minimum Cuts. In Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics (ACL-04). Barcelona, Spain, 271–278. https://doi.org/10.3115/1218955.1218990 [231] Jana Papcunová, Marcel Martončik, Denisa Fedáková, Michal Kentoš, Miroslava Bozogáňová, Ivan Srba, Robert Moro, Matúš Pikuliak, Marián Šimko, and Matúš Adamkovič. 2023. Hate speech operationalization: a preliminary examination of hate speech indicators and their structure. Complex & Intelligent Systems 9, 3 (01 Jun 2023), 2827–2842. https://doi.org/10.1007/s40747-021-00561-0 [232] Valeria Pastorino, Jasivan A Sivakumar, and Nafise Sadat Moosavi. 2024. Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection. arXiv preprint arXiv:2402.11621 (2024). [233] Victoria Patricia Aires, Fabiola G. Nakamura, and Eduardo F. Nakamura. 2019. A Link-based Approach to Detect Media Bias in News Websites. In Companion Proceedings of The 2019 World Wide Web Conference (San Francisco, USA) (WWW ’19). Association for Computing Machinery, New York, NY, USA, 742–745. https://doi.org/10.1145/3308560.3316460 [234] Ayush Patwari, Dan Goldwasser, and Saurabh Bagchi. 2017. TATHYA: A Multi-Classifier System for Detecting Check-Worthy Statements in Political Debates. In Proceedings of the 2017 ACM on Conference on Information and Knowledge Management (Singapore, Singapore) (CIKM ’17). Association for Computing Machinery, New York, NY, USA, 2259–2262. https://doi.org/10.1145/3132847.3133150 [235] Amalie Pauli, Rafael Sarabia, Leon Derczynski, and Ira Assent. 2023. TeamAmpa at SemEval-2023 Task 3: Exploring Multilabel and Multilingual RoBERTa Models for Persuasion and Framing Detection. In Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), Atul Kr. Ojha, A. Seza Doğruöz, Giovanni Da San Martino, Harish Tayyar Madabushi, Ritesh Kumar, and Elisa Sartori (Eds.). Association for Computational Linguistics, Toronto, Canada, 847–855. https://doi.org/10.18653/v1/2023.semeval-1.117 [236] Amalie Brogaard Pauli, Isabelle Augenstein, and Ira Assent. 2024. Measuring and Benchmarking Large Language Models’ Capabilities to Generate Persuasive Language. https://doi.org/10.48550/arXiv.2406.17753 arXiv:2406.17753. [237] John Pavlopoulos, Jeffrey Sorensen, Léo Laugier, and Ion Androutsopoulos. 2021. SemEval-2021 Task 5: Toxic Spans Detection. In Proceedings of the 15th International Workshop on Semantic Evaluation, SemEval@ACL/IJCNLP 2021, Virtual Event / Bangkok, Thailand, August 5-6, 2021, Alexis Palmer, Nathan Schneider, Natalie Schluter, Guy ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:73 Emerson, Aurélie Herbelot, and Xiaodan Zhu (Eds.). Association for Computational Linguistics, 59–69. https: //doi.org/10.18653/V1/2021.SEMEVAL-1.6 [238] Branislav Pecher, Jan Cegin, Robert Belanec, Jakub Simko, Ivan Srba, and Maria Bielikova. 2024. Fighting Randomness with Randomness: Mitigating Optimisation Instability of Fine-Tuning using Delayed Ensemble and Noisy Interpolation. In Findings of the Association for Computational Linguistics: EMNLP 2024, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 11005–11044. https: //doi.org/10.18653/v1/2024.findings-emnlp.644 [239] Branislav Pecher, Ivan Srba, and Maria Bielikova. 2024. Comparing Specialised Small and General Large Language Models on Text Classification: 100 Labelled Samples to Achieve Break-Even Performance. arXiv:2402.12819 [cs.CL] https://arxiv.org/abs/2402.12819 [240] Branislav Pecher, Ivan Srba, and Maria Bielikova. 2024. On Sensitivity of Learning with Limited Labelled Data to the Effects of Randomness: Impact of Interactions and Systematic Choices. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 522–556. https://doi.org/10.18653/v1/2024.emnlpmain.32 [241] Branislav Pecher, Ivan Srba, and Maria Bielikova. 2024. A Survey on Stability of Learning with Limited Labelled Data and its Sensitivity to the Effects of Randomness. ACM Comput. Surv. 57, 1, Article 19 (Oct. 2024), 40 pages. https://doi.org/10.1145/3691339 [242] Qiwei Peng, Robert Moro, Michal Gregor, Ivan Srba, Simon Ostermann, Marian Simko, Juraj Podrouzek, Matúš Mesarčík, Jaroslav Kopčan, and Anders Søgaard. 2025. SemEval-2025 Task 7: Multilingual and Crosslingual FactChecked Claim Retrieval. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Sara Rosenthal, Aiala Rosá, Debanjan Ghosh, and Marcos Zampieri (Eds.). Association for Computational Linguistics, Vienna, Austria, 2498–2511. https://aclanthology.org/2025.semeval-1.323/ [243] Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). 1532–1543. [244] Verónica Pérez-Rosas, Bennett Kleinberg, Alexandra Lefevre, and Rada Mihalcea. 2018. Automatic Detection of Fake News. In Proceedings of the 27th International Conference on Computational Linguistics, Emily M. Bender, Leon Derczynski, and Pierre Isabelle (Eds.). Association for Computational Linguistics, Santa Fe, New Mexico, USA, 3391–3401. https://aclanthology.org/C18-1287 [245] Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language Models as Knowledge Bases?. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (Eds.). Association for Computational Linguistics, Hong Kong, China, 2463–2473. https://doi.org/10.18653/v1/D19-1250 [246] Lara Piccolo, Azizah C Blackwood, Tracie Farrell, and Martino Mensio. 2021. Agents for fighting misinformation spread on Twitter: design challenges. In Proceedings of the 3rd Conference on Conversational User Interfaces. 1–7. [247] Matúš Pikuliak, Ivan Srba, Robert Moro, Timo Hromadka, Timotej Smoleň, Martin Melišek, Ivan Vykopal, Jakub Simko, Juraj Podroužek, and Maria Bielikova. 2023. Multilingual Previously Fact-Checked Claim Retrieval. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 16477–16500. https://doi.org/10.18653/v1/2023.emnlpmain.1027 [248] Dina Pisarevskaya and Arkaitz Zubiaga. 2025. Zero-shot and Few-shot Learning with Instruction-following LLMs for Claim Matching in Automated Fact-checking. In Proceedings of the 31st International Conference on Computational Linguistics, Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert (Eds.). Association for Computational Linguistics, Abu Dhabi, UAE, 9721–9736. https://aclanthology.org/ 2025.coling-main.650/ [249] Jakub Piskorski, Nicolas Stefanovitch, Giovanni Da San Martino, and Preslav Nakov. 2023. SemEval-2023 Task 3: Detecting the Category, the Framing, and the Persuasion Techniques in Online News in a Multi-lingual Setup. In Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), Atul Kr. Ojha, A. Seza Doğruöz, Giovanni Da San Martino, Harish Tayyar Madabushi, Ritesh Kumar, and Elisa Sartori (Eds.). Association for Computational Linguistics, Toronto, Canada, 2343–2361. https://doi.org/10.18653/v1/2023.semeval-1.317 [250] Emily Pitler and Ani Nenkova. 2008. Revisiting readability: a unified framework for predicting text quality. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (Honolulu, Hawaii) (EMNLP ’08). Association for Computational Linguistics, USA, 186–195. [251] León Podgurski, Karolina Zaczynska, and Georg Rehm. 2022. Evaluating Web Content Using the W3C Credibility Signals. https://doi.org/10.3233/SSW220005 ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
0:80 Srba et al. Akhtar, Rami Aly, Christos Christodoulopoulos, Oana Cocarascu, Zhijiang Guo, Arpit Mittal, Michael Schlichtkrull, James Thorne, and Andreas Vlachos (Eds.). Association for Computational Linguistics, Dubrovnik, Croatia, 29–37. https://doi.org/10.18653/v1/2023.fever-1.3 [355] Lingling Xu, Haoran Xie, Si-Zhao Joe Qin, Xiaohui Tao, and Fu Lee Wang. 2023. Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment. arXiv:2312.12148 [cs.CL] https: //arxiv.org/abs/2312.12148 [356] Rongwu Xu, Brian S. Lin, Shujian Yang, Tianqi Zhang, Weiyan Shi, Tianwei Zhang, Zhixuan Fang, Wei Xu, and Han Qiu. 2024. The Earth is Flat because...: Investigating LLMs’ Belief towards Misinformation via Persuasive Conversation. https://doi.org/10.48550/arXiv.2312.09085 arXiv:2312.09085. [357] Hao Yan, Sanmay Das, Allen Lavoie, Sirui Li, and Betsy Sinclair. 2019. The Congressional Classification Challenge: Domain Specificity and Partisan Intensity. In Proceedings of the 2019 ACM Conference on Economics and Computation (Phoenix, AZ, USA) (EC ’19). Association for Computing Machinery, New York, NY, USA, 71–89. https://doi.org/10. 1145/3328526.3329582 [358] Fumeng Yang, Zhuanyi Huang, Jean Scholtz, and Dustin L. Arendt. 2020. How do visual explanations foster end users’ appropriate trust in machine learning?. In Proceedings of the 25th International Conference on Intelligent User Interfaces (Cagliari, Italy) (IUI ’20). Association for Computing Machinery, New York, NY, USA, 189–201. https: //doi.org/10.1145/3377325.3377480 [359] Junting Ye and Steven Skiena. 2019. MediaRank: Computational Ranking of Online News Sources. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (Anchorage, AK, USA) (KDD ’19). Association for Computing Machinery, New York, NY, USA, 2469–2477. https://doi.org/10.1145/3292500.3330709 [360] Qingqing Yi, Zhong Qian, Peifeng Li, and Qiaoming Zhu. 2022. Chinese Sentence-level Event Factuality Identification with Recursive Neural Network. In 2022 International Joint Conference on Neural Networks (IJCNN). 1–8. https: //doi.org/10.1109/IJCNN55064.2022.9892209 [361] Shehel Yoosuf and Yin Yang. 2019. Fine-Grained Propaganda Detection with Fine-Tuned BERT. In Proceedings of the Second Workshop on Natural Language Processing for Internet Freedom: Censorship, Disinformation, and Propaganda, Anna Feldman, Giovanni Da San Martino, Alberto Barrón-Cedeño, Chris Brew, Chris Leberknight, and Preslav Nakov (Eds.). Association for Computational Linguistics, Hong Kong, China, 87–91. https://doi.org/10.18653/v1/D19-5011 [362] Himanshu Zade, Megan Woodruff, Erika Johnson, Mariah Stanley, Zhennan Zhou, Minh Tu Huynh, Alissa Elizabeth Acheson, Gary Hsieh, and Kate Starbird. 2023. Tweet Trajectory and AMPS-based Contextual Cues can Help Users Identify Misinformation. Proc. ACM Hum.-Comput. Interact. 7, CSCW1, Article 103 (apr 2023), 27 pages. https://doi.org/10.1145/3579536 [363] Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. 2019. SemEval2019 Task 6: Identifying and Categorizing Offensive Language in Social Media (OffensEval). In Proceedings of the 13th International Workshop on Semantic Evaluation. Association for Computational Linguistics, Minneapolis, Minnesota, USA, 75–86. https://doi.org/10.18653/v1/S19-2010 [364] Marcos Zampieri, Preslav Nakov, Sara Rosenthal, Pepa Atanasova, Georgi Karadzhov, Hamdy Mubarak, Leon Derczynski, Zeses Pitenis, and Çağrı Çöltekin. 2020. SemEval-2020 Task 12: Multilingual Offensive Language Identification in Social Media (OffensEval 2020). In Proceedings of SemEval. [365] Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019. Defending Against Neural Fake News. In Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Eds.), Vol. 32. Curran Associates, Inc. https: //proceedings.neurips.cc/paper_files/paper/2019/file/3e9f0fc9b2f89e043bc6233994dfcf76-Paper.pdf [366] Amy X. Zhang, Aditya Ranganathan, Sarah Emlen Metz, Scott Appling, Connie Moon Sehat, Norman Gilmore, Nick B. Adams, Emmanuel Vincent, Jennifer Lee, Martin Robbins, Ed Bice, Sandro Hawke, David Karger, and An Xiao Mina. 2018. A Structured Response to Misinformation: Defining and Annotating Credibility Indicators in News Articles. In Companion Proceedings of the The Web Conference 2018 (Lyon, France) (WWW ’18). International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, CHE, 603–612. https://doi.org/10.1145/3184558. 3188731 [367] Canyu Zhang and Paul D. Clough. 2020. Investigating clickbait in Chinese social media: A study of WeChat. Online Social Networks and Media 19 (2020), 100095. https://doi.org/10.1016/j.osnem.2020.100095 [368] Heng Zhang, Peifeng Li, Zhong Qian, and Xiaoxu Zhu. 2023. Incorporating factuality inference to identify documentlevel event factuality. In Findings of the Association for Computational Linguistics: ACL 2023. 13990–14002. [369] Yifan Zhang, Giovanni Da San Martino, Alberto Barrón-Cedeño, Salvatore Romeo, Jisun An, Haewoon Kwak, Todor Staykovski, Israa Jaradat, Georgi Karadzhov, Ramy Baly, Kareem Darwish, James Glass, and Preslav Nakov. 2019. Tanbih: Get To Know What You Are Reading. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): System Demonstrations, Sebastian Padó and Ruihong Huang (Eds.). Association for Computational Linguistics, Hong ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.
Survey on Automatic Credibility Assessment Using Textual Credibility Signals 0:81 Kong, China, 223–228. https://doi.org/10.18653/v1/D19-3038 [370] Liang Zhao, Ting Hua, Chang-Tien Lu, and Ing-Ray Chen. 2016. A topic-focused trust model for Twitter. Computer Communications 76 (2016), 1–11. https://doi.org/10.1016/j.comcom.2015.08.001 [371] Xinyi Zhou and Reza Zafarani. 2020. A Survey of Fake News: Fundamental Theories, Detection Methods, and Opportunities. Comput. Surveys 53, 5 (Oct. 2020), 1–40. https://doi.org/10.1145/3395046 arXiv:1812.00315 [372] Aneta Zugecova, Dominik Macko, Ivan Srba, Robert Moro, Jakub Kopál, Katarína Marcinčinová, and Matúš Mesarčík. 2025. Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 780–797. https://doi.org/10.18653/v1/2025.acl-long.38 [373] Hafzullah İŞ and Taner TUNCER. 2018. Confidence Index Analysis of Twitter Users Timeline. In 2018 International Conference on Artificial Intelligence and Data Processing (IDAP). 1–8. https://doi.org/10.1109/IDAP.2018.8620917 ACM Trans. Intell. Syst. Technol., Vol. 0, No. 0, Article 0. Publication date: 2025.