scieee AI-readable full text Open interactive document viewer

Critical thinking and artificial intelligence in academia: A qualitative matrix analysis procedure for evaluating AI systems

Codina, Lluís; Aguilera-Cora, Elisenda; Lopezosa, Carlos; Freixa, Pere

Abstract

This work introduces the Matrix AI Systems Analysis Procedure (MASIA), a qualitative, matrix- based method designed to evaluate the performance and quality of generative artificial intelligence systems within academic settings. MASIA centers on the analysis of three key components in AI-generated responses: narrative synthesis, source usage, and the formulation of new prompts. By doing so, it fosters critical thinking among users and offers valuable tools for both teaching and research.The procedure defines variables and analytical parameters that enable the comparison of different AI systems, thereby supporting informed decision- making in scholarly and research environments. Furthermore, MASIA integrates ethical considerations, including traceability, proper attribution, and plagiarism prevention, making it a flexible instrument adaptable to various academic needs and projects. The chapter concludes that MASIA is a straightforward yet powerful tool for enhancing critical thinking, improving teaching and learning processes, and providing a foundation for comparative research on artificial intelligence in academia.

Full text

161 Critical thinking and artificial intelligence in academia: A qualitative matrix analysis procedure for evaluating AI systems Lluís Codina Universitat Pompeu Fabra, Spain https://orcid.org/0000-0001-7020-1631 Elisenda Aguilera-Cora Universitat Pompeu Fabra, Spain https://orcid.org/0000-0003-0923-9192 Carlos Lopezosa Universitat de Barcelona, Spain https://orcid.org/0000-0001-8619-2194 Pere Freixa Universitat Pompeu Fabra, Spain https://orcid.org/0000-0002-9199-1270 Codina, L., Aguilera-Cora, E., Lopezosa, C., & Freixa, P. (2025). Critical thinking and artificial intelligence in academia: A qualitative matrix analysis procedure for evaluating AI systems. In J. Guallar, M. Vállez, & A. Ventura-Cisquella (Coords). Digital communication. Trends and good practices (pp. 161-173). Ediciones Profesionales de la Información. https://doi.org/10.3145/cuvicom.12.eng 162 Critical thinking and artificial intelligence in academia: A qualitative matrix analysis procedure for evaluating AI systems Lluís Codina; Elisenda Aguilera-Cora; Carlos Lopezosa; Pere Freixa Digital communication. Trends and good practices Abstract This work introduces the Matrix AI Systems Analysis Procedure (MASIA), a qualitative, matrix-based method designed to evaluate the performance and quality of generative artificial intelligence systems within academic settings. MASIA centers on the analysis of three key components in AI-generated responses: narrative synthesis, source usage, and the formulation of new prompts. By doing so, it fosters critical thinking among users and offers valuable tools for both teaching and research.The procedure defines variables and analytical parameters that enable the comparison of different AI systems, thereby supporting informed decision-making in scholarly and research environments. Furthermore, MASIA integrates ethical considerations, including traceability, proper attribution, and plagiarism prevention, making it a flexible instrument adaptable to various academic needs and projects. The chapter concludes that MASIA is a straightforward yet powerful tool for enhancing critical thinking, improving teaching and learning processes, and providing a foundation for comparative research on artificial intelligence in academia. Keywords Generative artificial intelligence; Qualitative evaluation; Critical thinking; Analysis matrices; Academic ethics; AI systems in academy; Evaluative methods. 1. Introduction This paper presents an analytical procedure to evaluate the performance and quality of generative artificial intelligence systems in academic environments. The procedure, which we call the Matrix AI Systems Analysis Procedure or MASIA, is designed to evaluate AI systems that, as part of their response, not only provide a narrative summary but also include citations and the bibliographic sources they used to generate the content. This method of analysis promotes critical thinking among AI system users, provides elements for teaching-learning processes, and can be the basis for developing data collection in research processes. The use of sources as part of the response is a necessity in the academic context because one can verify and expand the information provided by the AI, as well as maintain the chain of attributions (HighLevel Expert Group on Artificial Intelligence . 2019; Crompton and Burke, 2023¸ Kaebnick et al. 2023; Lund et al, 2023; Tilie et al., 2023; Gundersen et al., 2018; World Commission on the Ethics of Scientific Knowledge and Technology, 2019; Dwivedi et al., 2021; Bianchini et al. 2022). The latter is doubly convenient, because in addition to increasing the quality of the AI response, it prevents plagiarism or inadequate attribution, both of which are essential in academic work. The procedure presented here consists of analysis matrices based on a series of variables (Codina & Pedraza, 2016), which are determined according to the authors’ teaching experience and prior experience in analyzing and using AI systems for academic work, particularly due to the need to provide protocols for the critical use of AI systems to university students and predoctoral researchers. 163 Critical thinking and artificial intelligence in academia: A qualitative matrix analysis procedure for evaluating AI systems Lluís Codina; Elisenda Aguilera-Cora; Carlos Lopezosa; Pere Freixa Digital communication. Trends and good practices Since the launch of ChatGPT in late 2022, the authors have incorporated artificial intelligence into their teaching and research activities (Lopezosa & Codina, 2023; Lopezosa et al., 2023a; 2023b; Aguilera-Cora et al., 2024a; 2024b; Codina, 2025). This integration highlighted the need for an intellectual tool to train students and predoctoral researchers in the proper use of AI (Codina and Garde, 2023). There was also a need for an instrument that would allow comparative studies of the efficiency of artificial intelligence systems suitable for use in academia. The latter could be useful for research purposes, or for providing economic decision-makers at universities, for example, with information on which to evaluate the acquisition of AI systems (Bhatia, 2023; Whitfield & Hofmann, 2024; Elsevier 2024). Before presenting the components of the evaluation method, we must present some terminological clarifications, the consideration of which is part of the procedure itself, just as we must consider the composition of the results of AI when it responds to a user instruction. 2. Terminology The terminology presented in the table below is considered part of the procedure, so it is necessary to precisely establish the use of a set of terms for its proper application. Table 1 Terminology of the AI systems evaluation procedure Term Explanation Bibliographic sources In a RAG-type AI system (see definition below), this is the list of documents (journal articles, reports, web pages, etc.) that justify the answer. In the context of the evaluation procedure, this concept of source is the one taken into account unless otherwise indicated. Sources of information In a RAG-type AI system (see definition below), information sources are the resources used to locate the sources on which its response is based. Typical information sources for RAG-type AI systems can be academic databases or search engines like Google. Indicators In an evaluation procedure, indicators are characteristics that provide information about what is to be evaluated or compared. For example, in a comparative analysis of national economies, the unemployment rate, inflation, or GDP are indicators. By their very nature, they are also variables. Matrices Table-based information structures that allow data to be extracted, presented visually, and comparatively analysed. In the procedure presented here, the use of matrices is considered normative. The word “table” is equivalent. The use of the term “matrix” introduces the idea that a table meets certain conditions, the most important of which are that they are homogeneous tables and are organized so that rows are entities and columns are the properties of the entities. AI model Also called Large Language Model (LLM). This is the technical name for generative artificial intelligence, given its technological foundation. An AI model or LLM cannot be used by an end user, as its use requires programming and APIs. Therefore, end users generally work with AI through AI systems. Results page An AI system’s response is presented to the user on a results page that typically consists of at least three components: the narrative summary, a list of sources, and a set of suggested new prompts. Parameters In an evaluation procedure, parameters group indicators or variables. The idea is that a group of variables serves to characterize a significant aspect of a certain complexity of the subject being evaluated. For example, one parameter of nations in comparative analyses is their economy, another their demographics or political system, etc. But to characterize each parameter, it is necessary to use disaggregated indicators, such as the unemployment rate in the case of the economy, along with others such as GDP, etc. Prompt Natural language instruction used to obtain the response from an AI system. 164 Critical thinking and artificial intelligence in academia: A qualitative matrix analysis procedure for evaluating AI systems Lluís Codina; Elisenda Aguilera-Cora; Carlos Lopezosa; Pere Freixa Digital communication. Trends and good practices Term Explanation Suggested Prompts List of new prompts that some AI systems provide as part of their response to a prompt. Rapid review A type of review that omits some of the usual controls of systematic reviews to obtain immediate results with preliminary value. Narrative syntheses from AI can be compared to a form of rapid review. Retrieval Augmented Generation (RAG) Augmented Retrieval (ARG) involves improving the answers generated by AI systems by combining their training base knowledge with information retrieved in real time from external sources such as academic or specialized databases or general-purpose search engines like Google or Bing. Most academic AI systems are ARG. Some general-purpose AI systems, like Perplexity, are also ARG-type systems. Google, presumably, should become ARG-type once it effectively integrates its search engine with its AI. Literature review A literature review is a systematic process of searching, selecting, analyzing, and synthesizing existing information on a specific topic. It involves critically evaluating previous studies, identifying patterns, debates, and gaps in current knowledge, and presenting a comprehensive and organized overview of the state of the art in the field. , Producing literature reviews applied to academia is one of AI’s main functions. Scratchpad The scratchpad is the section preceding the narrative synthesis where the AI presents the chain of reasoning it followed to accomplish its tasks. In our case, the task consists of solving the four phases leading to a narrative synthesis. The ability to examine the scratchpad means, among other things, checking whether the AI understood the objectives of the task and correctly approached it. Narrative synthesis A narrative synthesis is the result of analyzing a set of sources using a well-defined framework and compiling the key insights derived from the analysis of these documents into a coherent textual summary. A narrative synthesis is both a byproduct of a literature review and part of the response of a generative AI system. AI system It is composed of one or more AI models (also called LLMs for Large Language Model), of one or more software layers, e.g. , for querying databases (e.g. , academic databases), for managing references, etc., as well as a user interface. Utilities These consist of complementary functions that software programs typically offer in addition to their core functions. For example, in an AI system focused on academia, a utility might consist of resources for managing references. Variables A variable is a property of an entity that can take on different values for each entity, or over time within the same entity. Recording these values allows entities to be analyzed, characterized, and compared. They can also be called indicators if their heuristic nature is desired. Source: own elaboration 3. Composition of an AI system’s results page In a results page of the type of AI system we are analyzing, we can determine the existence of three main components: 1. Narrative synthesis. 2. Sources consulted. 3. New prompts. We will present each of these components for evaluation purposes. But first, we must point out a fourth element that, although not part of the results page, can be found within the interface itself: 4. Additional utilities and functionalities specific to each system. 165 Critical thinking and artificial intelligence in academia: A qualitative matrix analysis procedure for evaluating AI systems Lluís Codina; Elisenda Aguilera-Cora; Carlos Lopezosa; Pere Freixa Digital communication. Trends and good practices 3.1. Narrative synthesis Narrative synthesis is the text generated by generative artificial intelligence It’s a synthesis because it’s the result of analyzing and synthesizing a series of previous pieces of information. It’s narrative because it’s presented as a more or less articulated narrative or discourse. Other types of syntheses are possible, such as those presented in the form of tables or graphs. However, when we talk about narrative syntheses, we consider any format to be included, by extension, unless otherwise stated. From here, we can establish a first block of criteria according to which AIs that present synthesis are preferable, which are: – Traceable through a visible scratchpad. The scratchpad is a section preceding the narrative summary in which the AI transparently and traceably presents the chain of reasoning it followed to solve the task. The ability to examine this chain of reasoning allows for the detection of potential bias or other errors, as well as verifying whether the AI correctly understood the task. In any case, it facilitates the traceability of the process followed. – Articulated. This means that the summary is presented in some sort of structure. For example, in separate sections, possibly organized by headings and following some sort of logical arrangement of sections. – They exhibit coherence and cohesion. Coherence consists of the interrelationship of the sentences that make up the text through their appropriate connection to the main theme. It is manifested by thematic unity, the (relative) absence of redundancy, and the logical progression of ideas. Cohesion, on the other hand, is manifested in the grammatical interrelationship of sentences. It is primarily determined by the connectives. – They exhibit connectivity. The end of each paragraph anticipates the next, and the beginnings of subsequent paragraphs connect to the previous ones. This property is made evident through the use of connectives. Connectivity is increased if there is a section that reunites the main ideas, or a section with an equivalent function. – They are (relatively) long. All other criteria being equal, long summaries are preferable. Since we’re talking about a range that can extend from 200 to 3,000 words, those that, if necessary, can be closer to this higher limit are preferable. – They are multimodal. In addition to text, they include some additional formatting, e.g., tables, cards, concept maps, mind maps, or diagrams. 3.2. Sources In academic AI, sources are typically documents, reports, and scientific journal articles. Ideally they allow the ideas and content comprising the synthesis to be attributed to their original creators. In the academic context, we have a categorical imperative to use AI systems that, along with the generated narrative synthesis, provide the sources on which they are based. This is the reason for preferring RAG-type AI systems. An additional reason for this preference is that an absence of sources in the answer would lead to a break in the attribution chain and, consequently, predispose to plagiarism. Note that we are separating the actual act (or lack thereof) of plagiarism from the fact that some AI systems facilitate or promote, de facto, plagiarism by offering unsourced answers. Logically, we should prefer AI systems that, at the very least, do not promote plagiarism. 166 Critical thinking and artificial intelligence in academia: A qualitative matrix analysis procedure for evaluating AI systems Lluís Codina; Elisenda Aguilera-Cora; Carlos Lopezosa; Pere Freixa Digital communication. Trends and good practices For the purposes of our proposed analysis, AIs are preferable [to what?] in relation to sources: – They exhibit capillarity. AIs that assign sources at the sentence level, or failing that, at the paragraph level, are preferable to a final list that affects the entire undifferentiated narrative synthesis. Capillarity also implies connectivity, since when a paragraph has (for example) three related ideas, each source is linked to each of the ideas, instead of placing the three sources at the end of the paragraph or at the end of the entire synthesis. – They provide well-formed citation formats, that is, with complete reference information and, where appropriate, viable links. Furthermore, the user of AI is obliged to verify and review sources, not only to evaluate arguments, but also to attribute third-party ideas and content to their true authors through the conventional citation system. 3.3. New prompts Some AIs offer, as a third prominent component of their responses, a list of new prompts or new questions. This approach may be of little interest or may be very incisive. In the latter case, they have obvious heuristic value. Preferable are those that generate additional, related prompts or questions as part of the results page, which we will evaluate: – Opportunity. That is, are the new suggested prompts appropriate for the search objectives? – Variety of approaches. Do they offer new facets or approaches not considered in the original prompt? 3.4. Utilities and Idiofunctions Some AI systems present one or more characteristics that are specific and unique to the system under consideration. In contrast to common functions, we can speak of idiofunctions; that is, functions unique to each particular system and therefore present only in the system under consideration. For example, an AI system may present a function that consists of extracting concepts, or another that consists of being able to design analysis matrices from references. These are called idiofunctions because they are unique to each AI. Although these functions may become standardized over time (and the concept may lose its meaning), these differences are significant at present and useful to consider. 3.5. Analysis matrices With the help of the previous concepts, we can now present the elements of analysis, which we articulate in parameters and variables as shown in Table 2. 167 Critical thinking and artificial intelligence in academia: A qualitative matrix analysis procedure for evaluating AI systems Lluís Codina; Elisenda Aguilera-Cora; Carlos Lopezosa; Pere Freixa Digital communication. Trends and good practices Table 2 Analysis variables. Parameter Code Variables / Check Question 1. Narrative synthesis 1.1 Scratchpad Does it present a scratchpad with the chain of reasoning followed by the AI system? Can this document be consulted after the task is completed? 1.2 Articulation Is the narrative synthesis presented organized or articulated in various sections or is it presented as a continuum without a defined structure? 1.3 Coherence and connection Is there a consistent thematic unity throughout the narrative summary and within each paragraph? Is there a connection between the sections, paragraphs, or sections of the narrative summary? 1.4 Extension Is the summary of the narrative adequate for its objectives? How many words does the narrative summary contain? Are there alternative versions of the extension? 1.5 Multimodality Does the results page include only text, or does it include other forms of information, such as diagrams? If not initially presented, are they offered as alternatives? 2. Sources 2.1 Number How many sources are cited? 2.2 Diversity Do the sources exhibit adequate diversity for the purposes of the prompt? Note: A single database is not an a priori limitation on diversity. Does the system allow you to differentiate whether the sources belong to academic texts, press, grey literature, or other unregulated sources? 23 Capillarity Are the sources connected at least at the paragraph or section level? 2.4 Well formed Are the sources presented in a format that is easy to export, manage, and cite sources? 3. Suggested prompts 3.1 Chance Do the suggested new prompts seem appropriate or timely given the information needs? 3.2 Variety Are the prompts varied and help broaden the focus of the topic? 4. Idiofunctions 4.1 Specific and exclusive functions of each system considered In addition to the common functions examined, does the system have any other specific functions? 168 Critical thinking and artificial intelligence in academia: A qualitative matrix analysis procedure for evaluating AI systems Lluís Codina; Elisenda Aguilera-Cora; Carlos Lopezosa; Pere Freixa Digital communication. Trends and good practices Table 3 Theoretical scores. Parameter Code Variables Theoretical score Narrative synthesis 1.1 Scratchpad 0-3 1.2 Joint 1.3 Connection 1.4 Extension 1.5 Multimodality Sources 2.1 Number 2.2 Diversity 23 Capillarity 2.4 Well formed Suggested Prompts 3.1 Chance 3.2 Variety Idiofunctions 4.1 Specific and exclusive functions of each system considered 0-3 The scoring scale is offered as an example. For each use, those responsible may (with justification) determine other scales. In this case, we have used a scale typical of heuristic evaluations in the field of information systems usability, and it corresponds to the following estimate: Punctuation Interpretation 0 Absence of function or variable considered 1 The function or variable appears in a minimal expression 2 The function or variable is correctly implemented but allows for improvements 3 The function or variable is fully implemented The initial scores in this scoring system are assigned intuitively and are adjusted as new cases are examined to allow comparisons. A final tally is made once all cases are examined. It is also common for two analysts to assign scores independently, then the scores are compared, and discrepancies are resolved by consensus. However, the scale and the specific procedure for assigning scores can be established for each specific project. Tabla 4 Data extraction table Cod. Variable Punctuation Narrative synthesis 1.1 Scratchpad 1.2 Joint 1.3 Connection 1.4 Extension 1.5 Multimodality Sources 2.1 Number 2.2 Diversity 23 Capillarity 2.4 Well formed 169 Critical thinking and artificial intelligence in academia: A qualitative matrix analysis procedure for evaluating AI systems Lluís Codina; Elisenda Aguilera-Cora; Carlos Lopezosa; Pere Freixa Digital communication. Trends and good practices Cod. Variable Punctuation Additional Prompts 3.1 Chance 3.2 Variety Idiofunctions 4.1 Specific and exclusive functions of each system considered TOTAL Comparative summary table System Synthesis Sources Prompts Idiofuncion TOTAL The tables above are common examples of matrix-based analysis systems. For each specific project, project managers can modify any aspects as appropriate. 3.6. Other evaluation modes It is clear that different evaluation methods can be developed. Task-based evaluation is one significant alternative, as in Font-Julián et al. (2024), in which an evaluation model is developed that combines qualitative and quantitative analysis procedures under specific tasks and with a group of two or more users as judges who agree on their evaluations. Another significant alternative are the benchmarking evaluation methods, such as those that can be seen on the Artificial Analysis portal (https://artificialanalysis.ai/) where dozens of language models are periodically compared based on a battery of tests. 3.7. Differential contribution of this evaluation mode Any evaluation method can be useful, depending on the context and objectives of each case. The qualitative evaluation method we propose here has a threefold function: – Strengthen users’ critical thinking regarding AI. – Provide a method for teaching/learning and acquiring skills in the use of AI systems. – Provide a procedure for evaluating and comparatively analyzing AI systems based on qualitative matrix analysis of parameters and indicators. 3.8. Variable geometry procedure The MASIA procedure provides analytical frameworks that can be applied as presented. However, it is not essential to use all variables, and new variables can be added or even other parameters can be considered. The essence of this evaluation method is as follows: – The designers of the analysis, with any of the three objectives stated above, may consider the convenience of adding new variables or removing some of the variables.