Full text
User Involvement in Recommender System Audits: A Systematic Review LORENZO PORCARO,AGNESE MACORI, and DANIEL RAFFINI,Department of Computer, Control and Management Engineering "A. Ruberti", Sapienza University of Rome, Italy TIZIANA CATARCI,ISTC-CNR and Department of Computer, Control and Management Engineering "A. Ruberti", Sapienza University of Rome, Italy Recommender systems play a pivotal role in shaping access to information online, yet audits of these systems rarely prioritize the perspectives of those most affected by them: the users. In the light of a growing body of work on recommender system audits, this systematic review examines the extent and forms of user involvement, classifying modes of participation and assessing how they influence audit practices. Our analysis finds that most audits treat users primarily as passive data sources, with only a small fraction engaging directly with their lived experiences. Where direct involvement occurs, it provides insights that purely technical approaches often miss, exposing critical blind spots in current methodologies. We argue that meaningful accountability requires a paradigm shift: sustained and active user participation, enhanced access to data for independent auditors, and stronger interdisciplinary collaboration across research, industry, and civil society. These findings highlight the urgent need for auditing frameworks capable of delivering more transparent and genuinely accountable recommender systems. 1 INTRODUCTION Recommender Systems (RSs) are increasingly recognized as complex sociotechnical systems, whose performance emerges from the interplay of technical mechanisms and social dynamics [ 76 ]. As RSs become ubiquitous in online spaces, across social media, streaming platforms, and e-commerce, a range of ethical and societal concerns has emerged [ 133 ], which have fueled growing demands for transparency and accountability, influencing academic research, industry practices, and regulatory frameworks. Within this context, algorithmic auditing has emerged as a key empirical approach for investigating system behavior and identifying potential harms. Historically, users have been pivotal to RS’s development, providing data, shaping design, and participating in evaluations. Yet, their role in the broader auditing ecosystem remains unclear. This paper presents a systematic review of user involvement in RS audits, arguing that while user participation alone is insufficient, it is essential for meaningful and accountable assessments. Despite being both data sources and primary stakeholders, users are often overlooked in audits, and practices that actively involve affected communities remain rare. In this work, we systematically investigate the extent and impact of user involvement in auditing recommender systems through the following overarching research question: RQ. To what extent have users been involved in auditing recommender systems, and how has their involvement impacted the audit process? To provide a more nuanced understanding, we explore the following sub-questions: RQa. To what degree do recommender system audits depend on user involvement? RQb. At which stages of the auditing process are users involved, if at all? RQc. What types of user data are collected and utilized for auditing purposes? This research is part of the project Algorithmic Auditing for Music Discoverability (AA4MD) which has received funding from the European Union’s Horizon Europe research and innovation programme under the Marie Skłodowska-Curie grant agreement No 101148443. Authors’ Contact Information: Lorenzo Porcaro, [email protected]; Agnese Macori, [email protected]; Daniel Raffini, raffini@diag. uniroma1.it, Department of Computer, Control and Management Engineering "A. Ruberti", Sapienza University of Rome, Rome, Italy; Tiziana Catarci, [email protected], ISTC-CNR and Department of Computer, Control and Management Engineering "A. Ruberti", Sapienza University of Rome, Rome, Italy. Preprint. Under review, 2025. 1
2 Porcaro et al. While the main research question enables a broad assessment of user involvement, RQa explores methodological approaches, RQb investigates stages of the audit lifecycle, and RQc focuses on responsible data practices. In this regard, our main contributions are: (i) a systematic evaluation of how users are involved in RS audits; (ii) an assessment of the extent to which audits rely on user input; (iii) identification of the stages where users actively contribute; and (iv) a critical discussion of future pathways for participatory auditing. The remainder of this paper is structured as follows. Section 2reviews stakeholder participation in algorithmic auditing and provides a historical perspective on the role of users in RS development. Section 3defines the three key concepts for this review: recommender systems, algorithmic auditing, and user involvement. Section 4details our review methodology, covering eligibility criteria, sources, search strategy, selection, quality appraisal, data extraction, and synthesis approach. Section 5presents the included audits and categorizes user involvement, while Section 6analyzes the audit processes in terms of their dependence on user input, stages of participation, and types of user data employed. Section 7outlines future directions, including improved data access, broader cultural and geographic representation, and enhanced interdisciplinary collaboration, before concluding in Section 8. 2 BACKGROUND Over the past decade, research has increasingly focused on investigating, understanding, reporting, and mitigating the problematic behaviors of algorithmic systems, which can significantly affect and harm people’s lives. This section does not aim to provide a comprehensive review of the interdisciplinary debate on what constitutes an algorithmic audit, nor to propose a universal definition. Scholars have already contributed extensively, from the seminal work by Sandvig et al. [ 118 ] to the literature review by Bandy [ 10 ] and the overview by Metaxa et al. [ 88 ]. Similarly, it would be redundant to revisit the historical origins of auditing in the social sciences, as explored by Gaddis [ 49 ], or its influence on current practices, as discussed by Vecchione et al. [ 135 ]. We also do not aim to situate algorithmic auditing within the broader ecosystem of legal, ethical, social, and technological considerations. For up-to-date perspectives, we refer to the field scan by Costanza-Chock et al. [ 31 ], Koshiyama et al.’s insights on the alignment of audit practices with emerging legal requirements [ 77 ], and Panigutti et al.’s exploration of algorithmic investigations in online platforms and search engines in relation to the EU Digital Services Act (DSA) [102]. Building on this scholarship, Section 2.1 focuses on stakeholder participation in algorithmic auditing, emphasizing users and their connection to algorithmic accountability, or its absence. Section 2.1.1 reviews the literature on end-user algorithmic auditing, a novel paradigm that places users at the center of the auditing process. Finally, Section 2.2 introduces the technological focus of this review: recommender systems. We trace their origins and highlight how users have historically played a central role in their design and evaluation. Overall, our goal is to demonstrate that user involvement, while not sufficient on its own, is essential for auditing recommender systems. 2.1 Stakeholders’ Participation in Algorithmic Audit Understanding the scope of algorithmic auditing requires considering the roles played by different stakeholders. Two key factors, (1) access to information about the audited algorithmic system, and (2) the auditor’s contractual relationship with the audit target, provide a useful basis for categorizing audits into first-, second-, and third-party audits. Here, we summarize the main characteristics of each category, and we point to the works led by Raji [ 110 – 112 ] for a more detailed discussion. First-party audits are conducted by members of the organization deploying the system. These auditors have direct access but also share vested interests in the company. Second-party audits are performed by external contractors hired Preprint. Under review, 2025.
User Involvement in Recommender System Audits 3 by the organization. While they have controlled access to the system, they retain some independence from the hiring company. Third-party audits are carried out by independent entities, such as academic researchers, investigative journalists, or civil society organizations. These auditors typically have neither direct access nor a contractual relationship with the audited organization. In summary, first-party auditors are internal stakeholders with internal access, second-party auditors are external stakeholders with internal access, and third-party auditors are external stakeholders with external access. Regardless of the audit type, the core objective is to evaluate whether an algorithmic system functions properly within its socio-cultural context, or whether deficiencies exist [88]. Therefore, including multiple stakeholders is essential to identify harmful behaviors and assess their impact. Defining relevant stakeholders and their interests is a fundamental step in any audit [ 20 ], and commonly includes policymakers, industry practitioners, academics, NGOs, civil society, and individual users. Among these stakeholders, the inclusion of affected communities should be a priority [ 31 ], yet despite broad recognition of the need to engage those “most at risk of harm,” such involvement is often neglected, relegating these communities to marginalized external roles [ 110 ]. However, their inclusion is critical for at least two reasons. First, individuals who experience harms firsthand provide situated knowledge that can help auditors identify root causes and design better solutions [ 74 ], a rationale central to the end-user audit paradigm discussed later. Second, audits aim to identify who should be held accountable for potential harms, and external stakeholders are part of the fora that can demand accountability. In this regard, researchers at the Data & Society institute have examined how Algorithmic Impact Assessments intersect with accountability, analyzing how impacts are co-constructed through accountability relations [ 89 , 90 ]. Two key considerations relevant to this work emerge from their analysis. First, audits alone are insufficient to hold an organization accountable for the disparate harms its systems may cause. Second, accountability is realized only when audit methods and findings are shared with stakeholders who can enforce changes and provide remedies. These points directly inform the focus of this review: to understand why, when, and how users, as stakeholders, have been involved in recommender system audits, contributing to the co-construction of accountability around the harmful behaviors such systems may exhibit. Overall, the importance of stakeholder involvement in both the audit process and the broader accountability network aligns with the goal of fostering trustworthy systems by extending stakeholder engagement across the entire system lifecycle [47]. Next, we discuss how audits may be designed to fulfill this function. 2.1.1 End-user Algorithmic Audit. The involvement of end-users in the design, development, and evaluation of systems is a well-established practice in computing [ 19 ]. As early as the 1980s, scholars emphasized that those with a direct need for a system can and should play an active role in its construction. The literature on user-centered design, end-user involvement, and participatory and community-led approaches is extensive, and a full review is beyond the scope of this work. Here, we focus on seminal contributions to end-user algorithmic audits published in recent years, highlighting the relationship between participation, users, and communities in auditing practices. One challenge in this area is the heterogeneity of nomenclatures used to describe practices that place users at the center of audits. While this diversity can initially be misleading, it also reflects important methodological nuances. Shen et al. [ 124 ] introduced the notion of everyday algorithmic audit, emphasizing how users’ day-to-day interactions with systems serve as a valuable source of knowledge for identifying problematic behaviors. This temporal dimension is especially relevant for systems like RSs, which are deeply embedded in daily life. They propose a four-stage audit process: (1) users initiate audits; (2) they raise awareness of experienced harms; (3) they generate and test hypotheses Preprint. Under review, 2025.
4 Porcaro et al. about the causes of problematic system behavior; and (4) they propose mitigation strategies. DeVos et al. [ 36 ] expanded this idea through user-driven algorithmic audit, in which users actively guide the audit process. Their three-phase methodology, illustrated through a case study on online image search biases, includes: (1) developing search strategies to autonomously identify problematic behaviors; (2) interpreting results and assessing outcomes based on users’ knowledge and beliefs; and (3) devising actions to mitigate harmful biases. Li et al. [ 83 ] further extend this work by examining different levels of user participation and the varying roles users may assume in such audits. Deng et al. [ 33 ] proposed the concept of user-engaged algorithmic audits, emphasizing collaboration between users and industry practitioners. While related to user-driven audits, this approach highlights how different forms of user engagement serve multiple purposes in investigating problematic behaviors and explores the benefits and challenges of such engagement from an industry perspective. These perspectives collectively highlight the diverse ways users can interact with other stakeholders and the audit target. Everyday audits occur when users independently initiate and lead the process, drawing on folk theories and situated knowledge of system functioning to provide insights into the system’s impact [ 35 ]. User-driven audits, in contrast, involve researchers actively supporting the process by facilitating interactions and providing epistemological tools to formalize users’ knowledge. User-engaged audits position users as collaborators, though they may still operate under practitioner-led initiatives, with users playing a supporting role. While these approaches foreground individuals or groups of users as key stakeholders, a more complex challenge arises when extending participation to affected and marginalized communities. First, designing tools that support large-scale audits remains a technical hurdle, as addressed by Lam et al. [ 78 ], who employ collaborative filtering to audit a text toxicity detection model. Second, supporting participatory, community-led processes requires more than extracting knowledge or lived experiences for external purposes [ 58 ]. This distinction is reflected in our analysis of the audits included in this review, underscoring the gap between user involvement and community-centered auditing. 2.2 Users and Recommender Systems At its core, Recommender Systems can be defined as “[...] software tools and techniques that provide suggestions for items that are most likely of interest to a particular user” [ 114 ]. While RS technology has advanced significantly over the past three decades, one fundamental aspect remains unchanged: the user is a central stakeholder. This is not to diminish the importance of other stakeholders, as highlighted by research on multi-stakeholder RSs [ 1 ], but users have historically played a pivotal role in their development. Since the first implementations of RSs in the early 1990s [ 84 ], managing the potentially unlimited flow of information has been essential to help users access and choose among diverse sources. As Konstan and Terveen note in their historical overview [ 76 ], RSs were designed to support users’ goals and preferences by analyzing their interactions and feedback. In the absence of pre-existing datasets, users themselves actively generated the data necessary for these systems to function effectively. At a later stage, between 1995 and 2005, as recommender systems were increasingly deployed in commercial settings, improving user experience remained a central challenge. A turning point came with the Netflix Prize in 2006 [ 15 , 95 ], when Netflix offered $1 million to anyone who could improve the accuracy of its recommendation algorithm by at least 10%. This focus on system-centered metrics partly overshadowed the broader goal of enhancing user experience, an issue that persists today. Nonetheless, even at that time, several researchers recognized that prioritizing accuracy metrics alone could hinder meaningful progress in the field [ 87 ]. Consequently, the RS community emphasized the importance of developing user-centric perspectives, noting that gains in accuracy did not necessarily translate into improved user experience. In the 2010s, Pu [ 109 ], Knijnenburg [ 75 ], and their colleagues presented comprehensive Preprint. Under review, 2025.
User Involvement in Recommender System Audits 5 frameworks for evaluating and explaining user experience when interacting with RSs, reinforcing the centrality of the user in design, development, and evaluation. Over the past two decades, two key trends have transformed RS research. First, RSs have become the backbone of diverse applications, from social media to streaming services and e-commerce platforms. This has generated vast amounts of user behavioral data, enabling more sophisticated solutions, though an exclusive focus on behavior has sometimes backfired [ 42 , 93 ]. Second, while traditional measures of user experience focused on choice satisfaction, decision difficulty, perceived accuracy, diversity, and recommendation effectiveness, new areas of interest have emerged. Human values such as fairness, non-discrimination, agency, oversight, alongside trustworthiness, privacy, and safety, have become central [ 133 ]. In parallel, the ubiquity and opacity of these systems have raised societal and ethical concerns, prompting scholars, policymakers, and civil society to critically examine their objectives and broader implications [ 63 , 94 ]. These issues influence both academic research and industry practice, and have shaped regulatory debates. For instance, recent EU policy initiatives highlight the oversight of algorithmic systems, including recommender systems, through frameworks such as the General Data Protection Regulation (GDPR) [ 63 ], the Digital Services Act (DSA) [ 94 ], and the Artificial Intelligence (AI) Act [37]. At the intersection of technological development and stakeholder involvement, the question of how users can actively contribute to auditing becomes particularly relevant. Rather than treating users solely as data generators or passive recipients of recommendations, we argue for recognizing their capacity to surface issues, provide feedback, and shape system accountability. This review builds on this foundation to examine how users have been involved in auditing these systems. 3 DEFINITIONS This section introduces the working definitions of the three core elements of this review: recommender systems, algorithmic auditing, and user involvement. As noted in the previous section, multiple definitions for each concept may coexist. Our goal is not to establish a universal definition, but to provide a clear and consistent reference point to contextualize the analysis and synthesis that follow. 3.1 Recommender Systems We define a Recommender System as a tool that helps users find items of interest in situations of information or choice overload, in line with Jannach and Bauer [ 69 ]. Given an individual or group of users, the system aims to predict the relevance of an item or set of items for those users. Beyond this Information Retrieval perspective, we emphasize that RSs are complex sociotechnical systems, where the interplay of social and technical factors shapes system performance [ 136 ]. Indeed, the RSs considered in this review are those embedded in commercial and public online platforms or applications, including feed-ranking algorithms in social media and recommendation engines in streaming services. In these contexts, complexity also arises from interactions between multiple interconnected systems. Theoretically, it is important to distinguish systems that, although they may interact with RSs, possess distinct characteristics and are therefore outside the scope of this review. First, we do not consider search engines. While related from a technical perspective, search relies on active user queries, whereas RSs provide personalized suggestions often without explicit input. Second, we exclude content moderation systems. Although critical to platform operation and often preceding recommenders by filtering inappropriate content, these systems typically operate according to broadly applied, predefined policies. Third, ad delivery and targeting systems are excluded. Despite similarities in profiling and personalization, their explicit promotional purposes differentiate Preprint. Under review, 2025.
6 Porcaro et al. them in goal and design. We note that, especially in the context of online platforms, disentangling the functioning of interconnected systems can be challenging. Nevertheless, clarifying these distinctions is essential to defining the scope and limitations of this review. 3.2 Algorithmic Audit We adopt Bandy’s definition of an algorithmic audit as “an empirical study investigating a public algorithmic system for potential problematic behavior” [10]. By empirical study, we mean any quantitative, qualitative, or mixed-method analysis that produces evidence-based claims. In this review, the algorithmic systems under consideration are RSs deployed in commercial or other public settings. Finally, we define problematic behavior as any action that causes, or has the potential to cause, harm, and we refer to Shelby et al. [ 123 ] for a taxonomy of potential sociotechnical harms caused by algorithmic systems. Several types of audits fall under this definition. Technical audits typically compare a system’s behavior against a benchmark to determine whether deviations remain within acceptable parameters [ 90 ]. Most audits adopt this approach, analyzing the system’s raw output to infer its potential implications for users [ 88 ]. This technical perspective deviates from the original focus of social science audits, which rely on controlled real-world experiments to observe how people actually interact with systems [ 135 ]. The concept of sociotechnical audits, proposed by Lam et al. [ 79 ], bridges these approaches by combining algorithmic analysis with user-centered evaluation. At the other end of the spectrum, in terms of user involvement, are end-user audits, as introduced in Section 2.1.1. Within the broader audit ecosystem, this review excludes compliance audits, i.e., formal evaluations of an organization’s adherence to regulatory frameworks or established standards. While such audits are increasingly important, especially given new regulatory efforts governing digital spaces, their design, purpose, and implementation differ significantly from those typically discussed in the scientific literature. To maintain a focused scope, compliance audits are therefore outside the boundaries of this review. 3.3 User Involvement The role of user participation in RS audits is the central focus of this systematic review. However, at the time of writing, no comprehensive taxonomy exists to describe the different forms and characteristics of such involvement. For the purposes of this review, we propose a practical distinction among three levels of user involvement: direct, indirect, and none (Figure 1). Direct involvement refers to the active inclusion of users in the auditing process, where they provide explicit input at one or more stages of the audit. Common techniques include interviews, questionnaires, and participatory design workshops. In these cases, users are aware of the type of input being requested and can independently choose whether or not to participate. Indirect involvement describes situations in which user data is used by auditors to conduct the audit, often without direct interaction. Techniques such as web crawling, data scraping, or certain forms of crowdsourcing fall into this category. Here, users function primarily as data subjects rather than active participants. Finally, audits that involve neither active nor passive user participation, such as those relying entirely on agent-based modeling or on API, i.e., collected data without user traceability, are categorized as having no involvement. We emphasize that the distinction between different kinds of involvement is not intended to imply a hierarchy or value judgment. A central aim of this review is to clarify these modes of user participation through an in-depth analysis of the collected RS audit corpus. Moreover, many audits employ multiple forms of involvement, particularly in mixed-method studies. To account for this, we introduce a fourth category, multiple involvement, which captures cases Preprint. Under review, 2025.
User Involvement in Recommender System Audits 7 in which several forms of user participation coexist and shape the audit’s design. We argue that the aforementioned categorization may provide a valuable lens for addressing the research questions outlined in Section 1. Fig. 1. Conceptual overview of user involvement types in recommender system audits. The schema depicts a spectrum of engagement, ordered by increasing levels of involvement: (a–b) no involvement, including simulations or agent-based/sock-puppet testing; (c–d) indirect involvement, where user data is leveraged through scraping, crowdsourcing, or data donation; and (e) direct involvement, where users actively participate in the audit process. 4 METHODS This review follows the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines [ 100 ]. Following, we begin by presenting our positionality with respect to the object of analysis in Section 4.1. Then, Section 4.2 outlines the eligibility criteria for including audits in the corpus, Section 4.3 describes the information sources considered and the search strategy employed, and Section 4.4 details the selection process and quality assessment. Aftewards, Section 4.5 introduces the data collection process and the data items extracted, and Section 4.6 summarizes the method used to synthesize the results. The protocol for this systematic review was preregistered on the Open Science Framework (OSF) on 13/12/2024 and is available at https://osf.io/8d4tz/?view_only=a571041cf406448c8a7e832f9e66a6b9. 4.1 Researchers Positionality Positionality statements provide context about the circumstances in which research is conducted and interpreted by its authors. By explicitly stating our positionality, we aim to remain reflexive about our role as researchers engaging with the sociotechnical dimensions of recommender systems. The lead author has a background in RS research and experience designing audits for a regulatory agency. Two co-authors come from a Digital Humanities background, focusing on AI Ethics. The fourth author has expertise in Human-Computer Interaction and experience in the field of Responsible AI. All four authors conducted this work while affiliated with academic institutions in Southern Europe. We acknowledge that our collective experiences may shape how we interpret the audits included in this review and that alternative interpretations may arise from different perspectives. Our European, computer science, HCI, and digital humanities perspective may therefore emphasize certain aspects of user involvement while undervaluing approaches from other disciplines or cultural contexts. 4.2 Eligibility Criteria Based on the working definitions outlined in Section 3, we applied the following criteria to determine the eligibility of publications for this review: (1) Empirical study: The publication must present an experiment or analysis (quantitative, qualitative, or mixedmethod) that generates evidence-based claims with clearly defined outcomes. Systematic reviews, position papers, theoretical frameworks, conceptual analyses, and other forms of non-empirical work were excluded. Preprint. Under review, 2025.
8 Porcaro et al. (2) Public and real-world algorithmic system: The study must focus on systems that are publicly deployed and accessible. We excluded studies based solely on simulated systems using real-world data or analyses of prototypes or systems not yet in production. (3) Recommender Systems: The algorithmic system under audit must be a Recommender System. Audits of related technologies, such as search engines or online advertising delivery systems, were excluded. (4) Problematic behavior: The study must analyze the system in relation to some form of problematic behavior, defined as causing, or having the potential to cause, harm to individuals interacting with the system. Studies focused solely on system performance metrics or user engagement were excluded unless these outcomes were explicitly linked to harmful effects. In addition to these core criteria, we applied three further filters: (i) Publication year: Publications prior to 2005 were excluded. Although RSs date back to the early 1990s, this cutoff reduces heterogeneity and focuses on the period when RSs became widely embedded in online platforms. (ii) Publication type: Only peer-reviewed conference and journal publications were included. While we recognize the value of grey literature, it was excluded from the core systematic comparison to ensure methodological consistency. Relevant grey literature is analyzed separately to complement the findings. (iii) Publication language: Publications not written in English were excluded. We acknowledge that this reinforces language bias in academic research, but English was the only non-native language fluently spoken by all authors, making this a practical necessity. 4.3 Information Sources and Search Strategy We consulted four main sources for the literature search: •ACM Full-Text Collection, accessed via the ACM Digital Library interface. •EBSCO, searching across 27 databases (as listed in the pre-registration). •ProQuest, searching across 9 databases (as listed in the pre-registration). •Scopus, using its online search interface. All searches were executed on December 9, 2024, and the queries are publicly available in the pre-registration. To ensure that the review remained up-to-date during its development, we set weekly alerts in each source to identify newly published, potentially relevant material until March 1, 2025. We designed the search queries following a three-tiered approach. First, we included keywords representing the type of studies under investigation, namely, audits (Table 1, K1). Second, we compiled keywords related to potential negative impacts of algorithmic systems (Table 1, K2). Third, we added keywords specific to recommender systems (Table 1, K3). The first two keyword sets (K1 and K2) were searched across all available fields, whereas K3 was limited to the title, abstract, and keyword fields. This restriction reduced noise from more general or ambiguous terms and ensured that the focus on recommender systems was prominent in each publication. Our initial query design was based on the working definitions outlined in Section 3, however, a preliminary search revealed that several relevant publications were missing. To refine and validate the search strategy, we consulted two related literature reviews: one focused on problematic content in YouTube recommendations [ 139 ], and another presenting a risk-utility meta-analysis of algorithmic recommendation [ 62 ]. Although differing in scope, both reviews contained records relevant to our investigation. Analysis of the missing works revealed two main limitations in our initial query design: (i) the K2 keyword set, focused on problematic behaviors, was too narrow and did not capture the full range of issues addressed in relevant studies; and (ii) many studies referring to online platforms did not Preprint. Under review, 2025.
User Involvement in Recommender System Audits 9 explicitly mention “recommender systems,” instead using terms such as “YouTube’s algorithm.” To address these gaps, we expanded the K2 keyword set and added terms for widely used online platforms. In particular, we included platforms classified as very large online platforms and search engines under the EU’s Digital Services Act, defined as those with more than 45 million users in Europe [ 46 ]. Table 1presents the three sets of keywords (K1, K2, and K3) used in both the initial and final versions of the search query. Table 1. Keywords used in the search strategy. Changes between the initial and final versions are highlighted in red. 1st version Final version K1 audit OR auditing OR "impact assessment" audit OR auditing OR "impact assessment" K2 problematic OR harm OR risk OR unsafe OR accountability OR transparency problematic OR harm OR risk OR unsafe OR accountability OR transparency OR disturbed OR inappropriate OR "negative impact" OR "negative effect" OR fraud OR bias OR mistrust OR inequality OR unfair OR imbalance OR discrimination OR prejudice OR offensive OR undesirable OR unwanted OR responsibility OR auditability OR trustworthy K3 "recommender system" OR "recommendation system" OR "algorithmic recommendation" OR "recommendation algorithm" OR recommender "recommender system" OR "recommendation system" OR "algorithmic recommendation" OR "recommendation algorithm" OR recommender OR AliExpress OR Amazon.com OR "Amazon Store" OR "Apple AppStore" OR Booking.com OR Facebook OR "Google Play" OR "Google Maps" OR "Google Shopping" OR Instagram OR LinkedIn OR Pinterest OR Pornhub OR Snapchat OR Shein OR Stripchat OR TikTok OR Temu OR Twitter OR Wikipedia OR YouTube OR Zalando OR XVideos OR XNXX 4.4 Selection Process and Quality Appraisal The initial screening was independently conducted by three authors, who assessed each publication based on its title and abstract against the eligibility criteria. When uncertainties arose, records were flagged for discussion. After the first screening round, all flagged records were jointly reviewed until consensus was reached. Publications that passed this step proceeded to full-text analysis, during which a quality appraisal (QA) was carried out to ensure that only high-quality studies were included in the review. The QA process drew on established appraisal frameworks while incorporating domain-specific considerations relevant to recommender system audits. Generic QA tools provide broad guidance for evaluating research but do not fully account for the unique challenges of auditing recommender systems, such as identifying problematic behaviors and assessing user involvement. To address this gap, we designed a QA tool around three main criteria: i) relevance: the study addresses auditing methodologies, user participation, and the identification of problematic behaviors; ii) clarity: transparency in research objectives, findings, and limitations; iii) rigor: the use of robust designs and tools to identify and measure problematic behaviors, with attention to reproducibility. The tool includes 10 questions, each scored on a three-point scale (2 = fully met, 1 = partially met, 0 = not met). To be included in the systematic review, records were required to score at least 70% of the maximum possible points. This Preprint. Under review, 2025.
16 Porcaro et al. focused on viewpoint exposure, in contrast, illustrate how recommender systems can shape access to political information and impact democratic processes, reinforcing the call for greater transparency and accountability in algorithmic decisionmaking. The only exception in this set is the audit of the room rental platform, which focuses on how the platform’s recommendation system contributes to disparities in exposure, revealing that certain user groups receive significantly more visibility than others. Table 4. Summary of audits with users’ indirect involvement (ordered by publication year). LEGEND: Problematic Behaviours: A&A = Agency and Awareness, AM = Amplification, RH = Representational Harm, VE = Viewpoint Exposure; Stages: PI = Problem Identification, HF = Hypothesis Formulation, A/T = Analysis/Testing, MI = Mitigation. Ref Design Platform Prob. Beh. Stages Data Population Size [9] QUAN Facebook VE A/T Internal US adult ∼10M [14] QUAN Facebook VE A/T API Danish adult 1,000 [113] QUAN YouTube VE A/T API Worldwide ∼6M [105] QUAN YouTube AM A/T API Worldwide ∼8M [129] QUAN Room rental RH A/T Internal Barcelona-based adult 61,997 [11] QUAN X/Twitter VE A/T API Worldwide (engaged in US politics) 10,000 [12] QUAN X/Twitter VE A/T API Worldwide (engaged in US politics) 10,000 [67] QUAN X/Twitter AM A/T Internal Worldwide 2M [18] QUAN X/Twitter AM A/T Crawler French adult 463 [27] QUAN YouTube AM A/T API US adult (quota) 1,181 [64] QUAN YouTube VE A/T API US adult (quota) 48,026 [39] QUAN X/Twitter VE A/T API Worldwide (engaged in US politics) 928 5.4 Multiple Involvement Seven audits rely on multiple forms of involvement to advance the understanding of recommender systems’ problematic behaviours (Table 5). These audits typically combine a step where users, via interviews, surveys, or, in one case, gamification, actively provide information to the auditors, akin to audits with direct involvement, with a subsequent step analyzing user interactions using behavioral data, as in audits with indirect involvement. This two-step process is clearly more resource-intensive, but the triangulation of data from different sources provides greater analytical depth and can yield richer insights. YouTube dominates this category, accounting for half of the audits, and regarding the types of problematic behaviours examined, amplification and viewpoint exposure remain the most common focus, similar to audits with indirect involvement. The complexity of implementing multi-part audits is reflected in sample sizes. Aside from one audit, the largest in our corpus, involving 30 million users, conducted in collaboration between industry and academic institutions, the remaining studies involve much smaller populations ( 𝜇 = 151 ±126). There is a predominance of US-based users (6/7), with two audits using quota sampling to approximate national demographic distributions. The only exception is a study involving WHO infodemic managers located globally. Survey methods are the most common approach for collecting information directly from users. Surveys are typically administered before users begin their monitored interactions with the recommendation system, enabling characterization of participants by attributes such as political affiliation or gender and racial resentment. In two audits, surveys are administered multiple times, before and after platform interaction, allowing the study of exposure effects longitudinally. Interviews, a more resource-intensive method, are also conducted either before or after platform interaction to provide Preprint. Under review, 2025.
User Involvement in Recommender System Audits 17 qualitative insights that complement quantitative analysis. Interviews can serve an exploratory function when conducted beforehand, or an explanatory function when conducted afterward. The only distinct approach is a gamification strategy in which participants intentionally locate problematic content on YouTube with minimal clicks, and their results are compared with recommendations generated via the YouTube API. Table 5. Summary of the audit with users’ multiple involvement (ordered by publication year). LEGEND: Problematic Behaviours: A&A = Agency and Awareness, AM = Amplification, RH = Representational Harm, VE = Viewpoint Exposure; Stages: PI = Problem Identification, HF = Hypothesis Formulation, A/T = Analysis/Testing, MI = Mitigation. Ref Design Platform Prob. Beh. Stages Data Population Size [44] MM Facebook A&A HF, A/T, MI Interview + API US adult (quota) 40 [16] QUAN YouTube AM A/T Survey + Crawler US adult 361 [70] QUAN YouTube AM A/T Survey + Crawler US adult 99 [96] QUAN YouTube AM A/T Gamification + Crawler Worldwide 113 [52] QUAN Facebook, Instagram VE A/T Survey + Internal US adult 30M [126] MM TikTok AM HF, A/T Interview + Manual US adult (quota) 50 [137] QUAN Twitter VE HF, A/T Survey + Crawler US adult 243 5.5 No involvement Most recommender system audits in our corpus (38/65) do not involve users in any form (Table 6). These audits rely on systematic and quantitative analyses, examining recommender system outcomes without access to internal workings. Such reverse-engineering black-box methods have been widely applied to platforms including YouTube, Amazon, Facebook, and X/Twitter, with YouTube, again, being the most studied. We remark that this strong focus on YouTube is also evident in the systematic reviews by Hilbert et al. [ 62 ], as well as in the one dedicated exclusively to this platform [ 139 ]. For the purpose of this review, audits without user involvement are broadly divided into two categories: pathways analysis and visibility analysis. Pathways analysis starts from a seed item, i.e., a predefined entry point, and traces the recommendation pathways generated by the system. The aim is to map the progression from the seed item to subsequent ones through intermediate recommendations. If the seed contains problematic content (e.g., books promoting vaccine misinformation), the analysis estimates how much additional problematic content is surfaced. Conversely, benign seeds (e.g., a children’s video) are used to assess the likelihood of encountering problematic content through subsequent recommendations. Seeds can be defined from prior audits, external sources, or keyword queries, or, in some scenario-specific studies, they may be cherry-picked. Importantly, seeds can be single items (e.g., a video) or groups of items (e.g., a channel), the latter being useful for thematic investigations such as extremist networks. Once seeds are set, recommendation pathways are extracted via platform APIs (if available) or by manually/automatically collecting recommendations. To minimize confounding factors, audits often employ controlled environments, such as VPNs, cleared histories, incognito mode, or virtual machines. After pathways are collected, the content must be classified as problematic or not. This classification can be manual (by researchers or annotators) or automated (using classifiers trained on labeled data). The analysis then examines distribution patterns of problematic content, either through descriptive statistics (e.g., share of pathways leading to problematic items, number of steps to reach them) Preprint. Under review, 2025.
18 Porcaro et al. or through network analysis, modeling pathways as graphs and applying tools like random walks to simulate user behavior. Visibility analysis, by contrast, does not follow trajectories but gathers the full set of content a hypothetical user may be exposed to. These users are simulated rather than real, typically through bots, personas, sock-puppet accounts, or agent-based testing. All such methods rely on assumptions about user attributes and behaviors, as defined by the audit designers. For example, personas may reflect political affiliations, bots may engage differently with conspiracy content, and agents may simulate geographic variation in access to information. Once simulated users and their interaction patterns are defined, recommendation data are collected, often over multiple days and sessions, to replicate realistic behavior, again within controlled environments. The resulting exposure is then classified and analyzed as in pathways analysis, but with a focus on how user characteristics and interactions shape the recommended content. In summary, while pathways analysis mimics what a user could encounter by following system recommendations, visibility analysis considers what a user could see and choose to interact with. Although we do not enter into the merits of these audits, which would require a much deeper and different analysis, we stress that these methods may serve only explanatory purposes, bounded by assumptions made at both system and user levels. As a result, problem identification and hypothesis generation are predetermined by the researchers conducting the audit. Table 6. Summary of audits with pathway and visibility analyses (ordered by publication year). LEGEND: Problematic Behaviors: A&A = Agency and Awareness, AM = Amplification, RH = Representational Harm, VE = Viewpoint Exposure Pathway Analysis Visibility Analysis Ref. Design Platform Prob. Beh. Ref. Design Platform Prob. Beh. [99] QUAN YouTube VE [54] QUAN Google News VE [130] QUAN YouTube AM [43] QUAN Spotify RH [122] QUAN YouTube AM [56] QUAN Facebook VE [127] QUAN Amazon AM [66] MM YouTube AM [72] QUAN YouTube VE [13] QUAN Twitter VE [103] QUAN YouTube AM [71] MM Amazon AM [2] QUAN YouTube AM [138] QUAN YouTube, Reddit, Gab AM [131] QUAN YouTube VE [5] QUAN YouTube AM [116] QUAN YouTube VE [81] QUAN YouTube VE [61] QUAN YouTube VE [121] QUAN YouTube AM [92] MM YouTube AM [53] QUAN TikTok VE [6] QUAN YouTube AM [132] QUAN YouTube AM [86] MM YouTube VE [119] QUAN YouTube AM [32] QUAN Amazon VE [82] QUAN TikTok VE [104] QUAN YouTube AM [125] QUAN Douyin VE [40] QUAN YouTube AM [140] MM YouTube AM [57] QUAN YouTube AM [51] QUAN Facebook, YouTube AM [80] QUAN YouTube VE [68] QUAN YouTube AM [65] QUAN YouTube VE [23] MM YouTube VE Preprint. Under review, 2025.
User Involvement in Recommender System Audits 19 6 DISCUSSION Users, if involved directly, may play a crucial, albeit often informal, role in the auditing of recommender systems, providing invaluable insights through their experiences and observations. This involvement extends beyond mere feedback, encompassing how users perceive, interact with, and even attempt to course-correct recommender system outputs. Through their development of algorithmic folk theories, personal understandings of how these systems function and impact their lives, users act as de facto auditors, constantly assessing the system’s behavior against their expectations. They may observe mismatches between their interests and the system’s design, identify biases and harms, and develop strategies to resist. These lived experiences directly inform the audit process and reveal the system’s responsiveness, or lack thereof, to user input. The impact of such direct involvement is profound: it provides rich, qualitative data that situates algorithms in real-world contexts and highlights their societal effects. Understanding these user-driven insights is crucial for informing future research and for guiding the development of more transparent and user-centered recommender systems. In contrast, indirect involvement positions users primarily as data subjects rather than active participants. Here, researchers analyze the digital traces users leave behind, what they view, click, or share, allowing large-scale, systematic examinations of recommender system behaviors that may be invisible to individuals. This approach enables audits to identify disparities in exposure and amplification, yet is constrained by the inability to capture the full spectrum of personalized interactions and by difficulties in establishing clear causal effects. While users are not active participants, their digital footprints provide the empirical basis for understanding recommender behaviors and holding platforms accountable. To address the limitations of both approaches, a few studies adopt multiple forms of user involvement. In these blended audits, users contribute both their lived experiences and their behavioral data, offering context that data alone cannot capture. Direct input helps confirm that digital choices are not always true reflections of preference and that users react sensitively to technical changes, challenging the notion of users as passive subjects. This combination strengthens the evidentiary basis of audits, moving beyond correlations to a more nuanced understanding of recommender systems’ influence. It highlights not only what content users are exposed to but also how they comprehend and react to it, providing insights for designing recommendations that are more transparent, responsive, and reflective of real human–recommender interactions. What has been discussed so far provides initial insights into the extent of user involvement in recommender system audits and how such involvement influences outcomes, summarized in Table 7. In the following sections, we analyze the findings in relation to the three sub-research questions. First, we examine to what degree recommender system audits depend on user involvement (RQa, Section 6.1). Second, we explore at which stages of the auditing process users are involved, if at all (RQb, Section 6.2). Third, we discuss the types of user data collected and used for auditing purposes (RQc, Section 6.3). 6.1 RQa: To what degree do recommender system audits depend on user involvement? The degree to which recommender system audits depend on user involvement varies considerably. Yet, a significant reliance on users as mere data subjects, or a complete lack of user involvement, remains common, despite widespread acknowledgment of users as key stakeholders. The current landscape shows a clear divide in both the extent and the purpose of user involvement, often determined by the audit’s scope and objectives. Preprint. Under review, 2025.
20 Porcaro et al. Table 7. Comparison of user involvement modes in RS audits. Direct Indirect Multiple How users are involved Users are active participants who provide insights based on their lived experiences and observations. They act as de facto auditors, constantly assessing the system’s behavior. Users are treated as passive data subjects. Their behavioral patterns and digital traces are often analyzed by researchers and auditors without direct interaction. Users contribute both their lived experiences and their behavioral data, integrating different types of information to create a more comprehensive audit. Data collected Self-reported data collected through interviews, surveys, and direct feedback. This includes folk theories, perceptions, and emotional responses to the system. Behavioral interaction data collected from online activities such as clickstreams, browsing histories, and other digital footprints, often without explicit consent. Both self-reported and behavioral data are combined to provide context that behavioral data alone cannot capture. Impact on the audit process Provides insights into real-world RS functions and societal effects. This involvement helps uncover problematic behaviors and how users adapt to unwanted outputs. Enables large-scale, systematic analysis of RS behaviors. This approach can reveal broad trends and systemic issues that may be invisible at the individual level. Bridges large-scale behavioral patterns with lived experiences. This strengthens the evidentiary basis for findings and helps move toward more contextualized explanations. Limitations Resource-intensive and often limited to small samples, making it difficult to generalize to systemic issues. User involvement is typically confined to early stages and rarely extends to validation or mitigation. Constrained by the difficulty of capturing personalized interactions and linking recommendations to user actions. User agency complicates causal claims, and ethical concerns remain significant. Complex to integrate different data types and to account for external influences on behavior. User participation is still uncommon in later audit stages, such as technical testing or developing fixes. A striking observation is that three-quarters of the audits in our corpus do not directly involve users. Instead, these audits predominantly simulate user behavior based on assumptions about preferences and attitudes, or rely on indirect tracking of user activity, effectively treating users as passive data sources. This approach enables the investigation of large-scale systemic issues caused by recommender systems, such as the spread of misinformation. By analyzing aggregated behavioral data, auditors can identify broad trends that may remain invisible at the individual level. However, such indirect or absent involvement typically excludes users from having a direct voice in the audit process and from participating in accountability mechanisms. This preference for indirect involvement is often a pragmatic choice. Participatory methods that entail direct user engagement are inherently more costly and resource-intensive. As a result, audits aiming for large-scale analysis tend to rely on indirect involvement, limited to data provision without active input, or, in more extreme cases, exclude users Preprint. Under review, 2025.
User Involvement in Recommender System Audits 21 altogether. This reveals a persistent tension between the pursuit of broad empirical insights and the practical constraints of conducting in-depth qualitative research with a large number of participants. Conversely, when the audit’s focus shifts to nuanced issues such as representational harms, direct user involvement becomes more prominent. These audits, which typically employ qualitative methods with smaller samples, foreground the lived experiences and perceptions of individuals. In this context, users are not merely data points: their perspectives provide essential insights into how algorithms affect, e.g., identity, fairness, and agency. Such active participation enables auditors to uncover subtle harms and unintended consequences that purely behavioral data would likely obscure, underscoring that for certain critical insights, direct user involvement is not only valuable but indispensable. The broader cognitive–behavioral focus that dominates much of recommender system technical research is also reflected in its audit methodologies. While understanding user behavior is undoubtedly important, an overreliance on behavioral proxies in audits risks obscuring the equally critical subjective and emotional dimensions of user experience. This partial mirroring of the field’s priorities in the audit literature highlights a missed opportunity to fully engage with and learn from users’ unique perspectives as key stakeholders. 6.2 RQb: At which stages of the auditing process are users involved, if at all? Users, when directly involved, are primarily engaged during the early stages of the auditing process for recommender systems, particularly in identifying problems and formulating hypotheses. This initial involvement draws on their lived experiences to uncover system limitations and highlight problematic behaviors. In cases of indirect involvement, such as through user data, this information is mainly used to test pre-identified hypotheses and analyze known issues, rather than to actively shape or validate potential solutions. A notable gap exists in user participation during the later stages of the audit process: (i) in technical testing, users are rarely involved in the validation of audit findings; and (ii) in mitigation, users typically play no role in the development or evaluation of proposed fixes or interventions. This limited scope of involvement means that, while audits can effectively harness user knowledge and lived experience to uncover previously unidentified problematic behaviors, the absence of user participation in later stages prevents validation of whether testing, analysis, and mitigation efforts are truly effective from the user’s perspective. The prevailing model reflects a one-way flow of information: user insights are extracted but seldom revisited or integrated into iterative refinement processes. In this sense, our review shows that user knowledge is consumed by the audit process without reciprocal feedback, and current practices rarely extend user involvement to more active, evaluative stages. 6.3 RQc: What types of user data are collected and utilized for auditing purposes? When auditing recommender systems, a variety of user data is collected and utilized, primarily falling into two broad categories: self-reported data and behavioral interaction data. The choice between these types often dictates the nature and depth of user involvement in the audit. Self-reported data, typically gathered through direct user engagement such as interviews or surveys, captures users’ subjective experiences and perceptions of recommender systems. This kind of data is essential for understanding how users make sense of algorithmic outputs, including their algorithmic folk theories and their awareness of potential biases or harms. It is particularly valuable for examining individual emotional responses to recommendations, dimensions often overlooked in audits based solely on behavioral traces. Behavioral interaction data, in contrast, is predominantly used in audits involving indirect user participation. This category includes clickstreams, browsing histories, and other forms of interaction data, often collected via scraping tools or crawlers. It enables large-scale analysis of consumption Preprint. Under review, 2025.
22 Porcaro et al. patterns, exposure dynamics, and content navigation. While this data is sometimes collected with user awareness, it is more frequently gathered without explicit input, treating individuals as passive data sources. A significant ethical concern arises around informed consent. Many of the audits analyzed do not report whether consent was obtained for data collection and use, raising critical questions of privacy, transparency, and data governance. Moreover, the type of audit, first- (platform-led), second- (collaborative), or third-party (external), often influences the nature and extent of data access. First-party audits generally benefit from unrestricted access to internal user data, while third-party audits must rely on publicly accessible data or user-contributed datasets, which can limit granularity and depth. Despite the range of data types used, certain dimensions of user experience remain underexplored. As noted, emotional and lived reactions to recommendations are rarely examined, particularly in audits that do not involve users directly. Given the profound influence recommender systems exert on preferences, attitudes, and decision-making, this oversight represents a significant gap in the current literature. Moreover, the fact that users are often unaware they are part of an audit further compounds the ethical and epistemic limitations, highlighting a disconnect between audit outcomes and the realities of those affected. 6.4 Limitations This systematic review provides insights into user involvement in RS audits, but has several limitations: Search Scope. Restricting to English-language publications likely excluded non-Western perspectives, reinforcing geographic concentration. Furthermore, excluding grey literature omitted significant contributions from NGOs and advocacy groups, potentially underestimating participatory auditing approaches. In fact, focusing on academic publications may underrepresent community-centered practices. Methodological and conceptual. Our taxonomy of user involvement simplifies nuanced methods and may obscure differences between tokenistic and genuinely collaborative approaches. Besides, quality appraisal and narrative synthesis, while systematic, limit causal inference and are influenced by interdisciplinary heterogeneity. Temporal and contextual. With most studies published 2020–2024, findings reflect recent practices rather than the field’s evolution. Moreover, treating audits as comparable across technological, regulatory, and social contexts may overlook factors affecting involvement strategies. Implications. Findings should be interpreted as a snapshot of academic auditing practices, not a comprehensive view of all user involvement. Future work should adopt multilingual searches, include grey literature, and foster cross-cultural and interdisciplinary collaboration to develop more inclusive and representative frameworks for understanding user participation in algorithmic auditing. 7 PATH FORWARD The advantages and disadvantages of different types of user involvement in recommender system audits have been highlighted by examining the kinds of data extracted, the stages of the audit in which users participate, and the degree to which audits depend on user involvement. Together, these aspects clarify the extent of user participation in audits and how such involvement may shape audit processes. Nevertheless, transversal issues identified in the reviewed literature merit further discussion to inform future audits and encourage better practices. These issues are not specific to any single form of user involvement but represent shared challenges across the field. For the sake of completeness, we highlight three critical aspects: i) persistent Preprint. Under review, 2025.
User Involvement in Recommender System Audits 23 challenges in accessing data; ii) the limited geographical and cultural scope of most audits; and iii) the need for stronger interdisciplinary collaboration, including recognition of the role of grey literature in shaping effective methodologies. 7.1 Facilitating Data Access for Research Purposes Whether self-reported or behavioral, qualitative or quantitative, user data is fundamental for auditing recommender systems. The characteristics of the data collected to analyze problematic behaviors directly affect the strength and robustness of the evidence produced, and, consequently, the validity of audits in informing stakeholders and supporting accountability. However, a clear asymmetry in data access between first-, second-, and third-party auditors presents significant challenges. Third-party auditors, in particular, often struggle to obtain the kind of robust evidence that internal audits, with privileged access to system data, can achieve. This observation does not undermine the value of adversarial audits conducted independently of platforms, rather, it underscores the importance of enabling external auditors to access internal data while preserving their scientific independence. There are examples of academia–industry collaboration that address this need, as well as recent regulatory developments, most notably the EU Digital Services Act (DSA) Article 40 [ 48 ]. These initiatives, along with ongoing debate among academics and regulators, represent a promising step forward. That said, our focus here is not on whether or how such access should be granted, but on the consequences that the lack of access has on current auditing practices. Indeed, the kind of data available profoundly influences the techniques employed, and without comprehensive internal data, auditors are often constrained to inductive, data-driven approaches. While not inherently problematic, this reliance risks tailoring the audit process to what is readily accessible rather than to what is most relevant or necessary for a complete assessment. This can bias research towards issues for which data is available, rather than those most critical from a societal perspective. This limitation is evident in the uneven distribution of audits across platforms. Some platforms are examined more extensively simply because their APIs and datasets are more accessible to researchers. Yet, even platform APIs pose significant challenges, as they are controlled by the platforms themselves and may not provide the scope or granularity required for robust academic work. For instance, Rieder et al. [ 115 ] highlight problems of query matching, completeness, temporal distance, and consistency in YouTube’s Data API, often used to infer behaviors of its recommender system. Similar concerns have been raised about the X/Twitter API [ 24 ] and the TikTok Research API [ 106 ]. Such limitations hinder comprehensive auditing, as APIs may filter, aggregate, or omit crucial data, potentially rendering audits less effective or even misleading. The challenge of data access also raises the critical question of how to balance the essential need for user data with the principles of privacy and protection against surveillance. Auditors must define proportionality, ensuring that data collected is strictly necessary and relevant for the audit’s purpose, without over-collecting personal information. This balance is vital for maintaining user trust and adhering to ethical standards. In response, innovative approaches like user data donation have emerged [ 17 , 60 , 98 ]. These initiatives allow users to consciously and voluntarily share their personal data with researchers, often through browser extensions. This method offers third-party auditors access to rich, ecologically valid data that better reflects real-world user experiences, while addressing privacy concerns more directly than passive tracking or broad scraping. Although the data collected may resemble that obtained through scraping, the key difference lies in active, informed consent and deliberate user involvement. This model represents a step towards more collaborative and ethically sound data acquisition, though it is important to stress that data donation does not necessarily equate to active participation in the audit process itself. Preprint. Under review, 2025.
24 Porcaro et al. 7.2 Broadening the Audit Geographical and Cultural Scope As introduced earlier, the core objective of any algorithmic audit should be to assess whether a system functions appropriately within its specific socio-cultural context. This is crucial because the behaviors of recommender systems affect people differently, not only due to the systems’ underlying design choices but also because of the social, cultural, and political environments in which they are deployed. A small number of audits in our review explicitly identify user groups as core stakeholders, for instance, South African teenage girls using Instagram, Black creators or LGBTQ+ users on TikTok, or Amharic-speaking women on YouTube. These cases demonstrate the unique insights that emerge when audits are grounded in specific communities. However, such examples remain rare. The vast majority of audits operate under the implicit assumption of an “average” user. This notion is not only questionable [ 21 , 117 ] but potentially harmful in this domain, as it risks reinforcing stereotypes and marginalizing underrepresented groups. The problem is compounded by the overwhelming concentration of audits on U.S.-based users, English-language data, and a handful of Western platforms. As highlighted by Urman et al. [ 134 ], this narrow scope severely limits the generalizability and relevance of audit findings for the global population affected by recommender systems. Methodologically, it also restricts researchers to data sources and scraping strategies optimized for Western contexts, sidelining smaller platforms, diverse linguistic settings, and culturally specific forms of content and interaction. Critical perspectives from Haraway’s contributions to Standpoint Theory [ 55 ], D’Ignazio and Klein’s Data Feminism [ 38 ], and Design Justice principles [ 30 , 34 ] provide valuable guidance here. Standpoint Theory underscores that marginalized communities often possess situated knowledge about sociotechnical systems’ limitations and harms. Data Feminism calls for pluralistic approaches that challenge dominant power structures in data practices, while Design Justice highlights the need for those most impacted to lead design and evaluation processes. Applied to auditing, these frameworks make clear that excluding diverse voices not only narrows the epistemic scope of audits but risks entrenching the very inequalities they aim to critique. Expanding the geographical and cultural reach of audits is thus both an ethical imperative and a methodological necessity: without situated perspectives from those most impacted, audits risk misdiagnosing harms, overlooking injustices, and ultimately legitimizing the asymmetries they intend to challenge. 7.3 Strengthening Interdisciplinary Collaboration and Grey Literature The sociotechnical nature of recommender systems demands audits that move beyond purely technical evaluations, and to fully understand the problematic behaviors that affect people’s lives, contributions from a range of disciplines are essential. Many studies in our corpus involve collaborations across social sciences, media studies, digital journalism, communication, and political science, often alongside computer scientists and engineers. Yet one notable absence is practitioners from the Recommender Systems and Information Retrieval communities. Even when audits adopt a black-box perspective, focusing on outputs and consequences rather than mechanisms, the insights of those who design, develop, and evaluate recommender systems remain invaluable. A recurring practice illustrates the risks of this gap: creating new user accounts to study problematic behaviors. This approach may be justified as a way to eliminate “the effect of real users’ data” and to “better isolate the algorithmic effect” [ 65 ]. However, from a RS perspective, this overlooks a fundamental issue: the cold-start problem, recognized in the literature since at least 1995 [ 85 ]. Cold-start effects mean that new users lack sufficient profile data, requiring a training period before recommendations align with preferences. Despite decades of technical literature on this issue Preprint. Under review, 2025.
User Involvement in Recommender System Audits 25 [ 101 ], none of the audits in our corpus explicitly recognize cold-start effects as a limitation. This raises important questions: when auditing under cold-start conditions, are we truly capturing the experience of real, long-term users? And can the “algorithmic effect” ever be meaningfully isolated from user data, given that recommender systems are inherently data-driven? These may appear to be minor technicalities, but they highlight a deeper point: robust auditing requires not only interdisciplinary breadth but also sufficient technical literacy in recommender system design. Another absence in this review, this time deliberate, concerns grey literature. We excluded it because its structures, purposes, and audiences differ significantly from academic publications, making systematic comparison challenging. Nevertheless, its relevance is indisputable. Grey literature has produced numerous audits that shed light on the societal impact of recommender systems, often through community-led and NGO-supported approaches. Examples include: Eticas Foundation’s audits of TikTok and YouTube from a migrant-rights perspective [ 45 ]; AI Forensics and Amnesty International’s work on TikTok and youth mental health [ 4 ]; investigations into Amazon bookstore recommendations with CheckFirst [ 3 ]; research from DCU’s Anti-Bullying Centre on male supremacist influencers [ 8 ]; AVAAZ’s documentation of YouTube’s climate misinformation problem [ 7 ]; Anti-Defamation League’s findings on extremist content exposure on YouTube [ 26 ]; the Center for Countering Digital Hate’s work on eating disorder and self-harm content on TikTok [ 25 ]; Data & Society’s introduction of data voids [ 50 ]; and the Mozilla Foundation’s “YouTube Regrets” project [91]. While we do not engage here with the scope or methodologies of these audits, they demonstrate the value of grounding investigations in the lived experiences of affected and marginalized communities. Collaborations with NGOs and CSOs, in particular, situate audits closer to those directly impacted by algorithmic harms, and address, albeit partially, the limitations highlighted in earlier sections on data access and cultural scope. Although grey literature faces its own methodological and data access challenges, its practices of involving users as key stakeholders offer critical lessons. For academic researchers, engaging with these approaches provides not only greater ecological validity but also a path toward producing results that can meaningfully challenge the accountability of recommender system owners and developers. 8 CONCLUSIONS This systematic review analyzed the involvement of users in 65 recommender system audits published between 2005 and 2024. Our findings reveal a significant gap between the recognized importance of user participation and its implementation in practice. Indeed, user involvement remains limited in scope and depth, where most of the audits rely on indirect or absent participation, treating users primarily as data sources. Direct involvement, when present, is mostly confined to early stages with minimal engagement in validation or mitigation. Quantitative, computational methods dominate, enabling large-scale analysis but often overlooking nuanced, lived experiences. Platform coverage is concentrated, particularly on YouTube, while those with restricted data access remain underexamined, reflecting structural inequalities in auditing. These findings carry important implications for multiple stakeholders. For researchers, our findings highlight the need to move beyond extractive approaches toward collaborative methodologies that treat users as co-auditors. Industry practitioners must recognize the limitations of purely technical audits, as excluding user perspectives risks blind spots. Policymakers should facilitate both data access and meaningful community involvement to enhance audit effectiveness and accountability. In our perspective, advancing user involvement requires: (i) developing scalable, participatory frameworks; (ii) expanding audits beyond Western, English-speaking contexts in ways responsive to local Preprint. Under review, 2025.
32 Porcaro et al. [126] Donghee Shin and Kulsawasd Jitkajornwanich. 2024. How Algorithms Promote Self-Radicalization: Audit of TikTok’s Algorithm Using a Reverse Engineering Method. Social Science Computer Review 42, 4 (2024), 1020–1040. doi:10.1177/08944393231225547 [127] Jieun Shin and Thomas Valente. 2020. Algorithms and Health Misinformation: A Case Study of Vaccine Books on Amazon. Journal of Health Communication 25, 5 (2020), 394–401. doi:10.1080/10810730.2020.1776423 [128] Ellen Simpson and Bryan Semaan. 2021. For You, or For"You"? Everyday LGBTQ+ Encounters with TikTok. Proc. ACM Hum.-Comput. Interact. 4, CSCW3, Article 252 (Jan. 2021), 34 pages. doi:10.1145/3432951 [129] David Solans, Francesco Fabbri, Caterina Calsamiglia, Carlos Castillo, and Francesco Bonchi. 2021. Comparing Equity and Effectiveness of Different Algorithms in an Application for the Room Rental Market. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (Virtual Event, USA) (AIES ’21). Association for Computing Machinery, New York, NY, USA, 978–988. doi:10.1145/3461702.3462600 [130] Melodie Yun-Ju Song and Anatoliy Gruzd. 2017. Examining Sentiments and Popularity of Proand Anti-Vaccination Videos on YouTube. In Proceedings of the 8th International Conference on Social Media & Society (Toronto, ON, Canada) (SMSociety17). Association for Computing Machinery, New York, NY, USA, Article 17, 8 pages. doi:10.1145/3097286.3097303 [131] Larissa Spinelli and Mark Crovella. 2020. How YouTube Leads Privacy-Seeking Users Away from Reliable Information. In Adjunct Publication of the 28th ACM Conference on User Modeling, Adaptation and Personalization (Genoa, Italy) (UMAP ’20 Adjunct). Association for Computing Machinery, New York, NY, USA, 244–251. doi:10.1145/3386392.3399566 [132] Ivan Srba, Robert Moro, Matus Tomlein, Branislav Pecher, Jakub Simko, Elena Stefancova, Michal Kompan, Andrea Hrckova, Juraj Podrouzek, Adrian Gavornik, and Maria Bielikova. 2023. Auditing YouTube’s Recommendation Algorithm for Misinformation Filter Bubbles. ACM Trans. Recomm. Syst. 1, 1, Article 6 (Jan. 2023), 33 pages. doi:10.1145/3568392 [133] Jonathan Stray, Alon Halevy, Parisa Assar, Dylan Hadfield-Menell, Craig Boutilier, Amar Ashar, Chloe Bakalar, Lex Beattie, Michael Ekstrand, Claire Leibowicz, Connie Moon Sehat, Sara Johansen, Lianne Kerlin, David Vickrey, Spandana Singh, Sanne Vrijenhoek, Amy Zhang, McKane Andrus, Natali Helberger, Polina Proutskova, Tanushree Mitra, and Nina Vasan. 2024. Building Human Values into Recommender Systems: An Interdisciplinary Synthesis. ACM Trans. Recomm. Syst. 2, 3, Article 20 (June 2024), 57 pages. doi:10.1145/3632297 [134] Aleksandra Urman, Mykola Makhortykh, and Aniko Hannak. 2025. WEIRD Audits? Research Trends, Linguistic and Geographical Disparities in the Algorithm Audits of Online Platforms - A Systematic Literature Review. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’25). Association for Computing Machinery, New York, NY, USA, 375–390. doi:10.1145/3715275.3732026 [135] Briana Vecchione, Karen Levy, and Solon Barocas. 2021. Algorithmic Auditing and Social Justice: Lessons from the History of Audit Studies. In Proceedings of the 1st ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (–, NY, USA) (EAAMO ’21). Association for Computing Machinery, New York, NY, USA, Article 19, 9 pages. doi:10.1145/3465416.3483294 [136] Guy H. Walker, Neville A. Stanton, Paul M. Salmon, and Daniel P. Jenkins. 2008. A review of sociotechnical systems theory: a classic concept for new command and control paradigms. Theoretical Issues in Ergonomics Science 9, 6 (2008), 479–499. doi:10.1080/14639220701635470 [137] Stephanie Wang, Shengchun Huang, Alvin Zhou, and Danaë Metaxa. 2024. Lower Quantity, Higher Quality: Auditing News Content and User Perceptions on Twitter/X Algorithmic versus Chronological Timelines. Proc. ACM Hum.-Comput. Interact. 8, CSCW2, Article 507 (Nov. 2024), 25 pages. doi:10.1145/3687046 [138] Joe Whittaker, Seán Looney, Alastair Reed, and Fabio Votta. 2020. Recommender systems and the amplification of extremist content. Internet Policy Review 10, 2 (2020). doi:10.14763/2021.2.1565 [139] Muhsin Yesilada and Stephan Lewandowsky. 2022. Systematic review: YouTube recommendations and problematic content. Internet Policy Review 11, 1 (March 2022). doi:10.14763/2022.1.1652 [140] Lisa Zieringer and Diana Rieger. 2023. Algorithmic Recommendations’ Role for the Interrelatedness of Counter-Messages and Polluted Content on YouTube – A Network Analysis. Computational Communication Research 5, 1 (2023), 109. doi:10.5117/CCR2023.1.005.ZIER Preprint. Under review, 2025.