Full text
148 AI and image banks: A research methodology Pere Freixa Universitat Pompeu Fabra, Spain https://orcid.org/0000-0002-9199-1270 Mar Redondo-Arolas Universitat de Barcelona, Spain https://orcid.org/0000-0002-0000-8593 Lluís Codina Universitat Pompeu Fabra, Spain https://orcid.org/0000-0001-7020-1631 Carlos Lopezosa Universitat de Barcelona, Spain https://orcid.org/0000-0001-8619-2194 Freixa, P., Redondo-Arolas, M., Codina, L., & Lopezosa, C. (2025). AI and image banks: A research methodology. In J. Guallar, M. Vállez, & A. Ventura-Cisquella (Coords). Digital communication. Trends and good practices (pp. 148-160). Ediciones Profesionales de la Información. https://doi.org/10.3145/cuvicom.11.eng
149 AI and image banks: A research methodology Pere Freixa; Mar Redondo-Arolas; Lluís Codina; Carlos Lopezosa Digital communication. Trends and good practices Abstract This chapter presents a methodological framework for analysing gender bias and the presence of sociocultural stereotypes in professional stock image banks, with a specific focus on the visual results returned by photographic and AI-generated platforms. The study is based on the hypothesis that neutral prompts — those lacking explicit references to gender, age, or ethnicity — should, in the absence of cultural or technical bias, yield a balanced visual representation across different social categories. Any significant deviation from such proportionality may indicate the existence of implicit biases or recurrent visual clichés. To explore this, the authors analysed images retrieved from four professional platforms — two based on conventional photography and two relying on AI image generation. A system of coded indicators was developed to classify the representations in terms of gender, age, ethnicity, functional diversity, beauty norms, and depicted actions. The methodology excluded group images and near-identical variants to ensure diversity and analytical rigour. The findings reveal that AI-based platforms more consistently align with user prompts (60.36%) compared to traditional photographic databases (44.84%). However, both types of platforms exhibit stereotypical patterns, suggesting a persistence of visual tropes and clichés. The proposed methodology proves effective in detecting these biases and offers a transferable analytical framework. The chapter aims to contribute to broader efforts towards more inclusive visual cultures, encouraging further interdisciplinary research on algorithmic image generation and representation in digital media. Keywords Gender bias; Stereotypes; Stock image platforms; Artificial intelligence; Visual representation; Image prompts; Algorithmic interpretation; Iconographic analysis; Media representation. 1. Introduction Image banks are among the main resources used by the media to visually complement the content they publish (Codina, 2011; Kamin, 2023). According to several authors (Gynnild, 2017; Mortensen & Gade, 2023), the downsizing of professional photography staff in traditional media outlets has led to a growing reliance on stock images to illustrate journalistic content, at the expense of commissioned photography (Mortensen et al., 2023; Ferry, 2023; Mortensen et al., 2024; Hugues, 2024). In his pioneering study, Tsang (1984) had already observed that by the late 1970s and early 1980s, only about one quarter of the photographs published were produced in-house. The remainder originated from agencies, image banks, and freelance photographers. It is therefore not surprising that image banks, alongside news agencies, have become the visual resource with the greatest capacity to influence and establish visual standards (Machin, 2004; Machin & Polzer, 2015; Frosh, 2015, 2020). Companies specialising in the distribution of stock photography have significantly expanded their capacity to generate all kinds of illustrations and visual resources — both photorealistic and fictional — through the incorporation of artificial intelligence (AI) technologies (PerdicesCastillo & Perianes-Rodríguez, 2011; Codina & Lopezosa, 2020; Vrabič-Dežman, 2024). AI
150 AI and image banks: A research methodology Pere Freixa; Mar Redondo-Arolas; Lluís Codina; Carlos Lopezosa Digital communication. Trends and good practices has also enabled the emergence of new players in the visual information market, who use image-generation tools to produce vast amounts of visual content. Through search interfaces, media outlets can access, locate, and select the images that best suit their needs within a highly competitive market (Allard, 2023; Bright, 2023; Sojit-Pejcha & Crabapple, 2023; Freixa & Redondo-Arolas, 2023; 2024). The growing prominence of stock images in the media has elicited mixed responses. Critics of the use of both stock images and AI-generated content in journalism argue, among other points, that such resources reproduce clichés and stereotypes, thus further perpetuating biases and undermining media credibility (Mortensen et al., 2023; Aiello et al., 2023). Moreover, they contend that these images lack editorial or documentary value, as they are not tied to specific events (Mortensen et al., 2024). As Frosh (2003) describes, “The most striking feature of stock images is their close relationship with classification. Inscribed within advertising ideology, these images participate in a representation of reality shaped by selective categorisation” (p. 91). Nonetheless, while providers of stock imagery are often criticised for reproducing standardised and stereotypical codes and formulas — regardless of the source of the photographs — Kwak and An (2016) demonstrated, using deep learning systems and large datasets, how media outlets use published press images to convey specific messages and often reinforce clichés and biases, irrespective of the images’ origins. Concerns about the reproduction of biases —be they gender, racial, cultural, or ideological— in press imagery have been widely acknowledged and explored both in general terms (Miller, 1975; Luebke, 1989; Rodgers & Thorson, 2000; Rodgers et al., 2007; Thurlow et al., 2020; Freixa et al., 2025) and in the context of political communication (Waldman & Devitt, 1998; Goodnow, 2010; Rönnback et al., 2025). These studies have employed combined qualitative and quantitative analytical methodologies, including statistical counting and both iconographic and iconological content analysis. Often, image analysis has been complemented by the examination of the accompanying text. Miller’s pioneering work, for example, involves counting the number of times men and women appear in the Los Angeles Times and The Washington Post over the course of a year, identifying the newspaper sections in which they are published and the roles depicted (Miller, 1975). By contrast, Goodnow (2010) focuses his study on a smaller sample of images, subjecting them to a detailed semiotic analysis that delves into the reading and interpretative codes inherent in journalistic imagery. As noted, Kwak and An (2016) adopted a quantitative approach, working with a much larger sample comprising two million images. Thurlow et al. (2020) and Freixa et al. (2025), meanwhile, applied semiotic analysis to samples of 600 images to assess the presence of stereotypes on stock image platforms. This chapter presents a methodological proposal for conducting visual research on images, clearly defining the observation parameters and the sample size. The proposed system has proven effective in studies examining gender bias and stereotypes in both photographic stock images and AI-generated visuals (Freixa et al., 2025). 2. Image banks and search platforms Image banks rely on digital platforms to make their collections available to clients. Through search interfaces, users can access, locate, and select those images that best suit their needs. This is a highly competitive market. Searches (conducted either through keywords or textual
151 AI and image banks: A research methodology Pere Freixa; Mar Redondo-Arolas; Lluís Codina; Carlos Lopezosa Digital communication. Trends and good practices prompts) generate visual galleries in response, presented as sets of thumbnail images which users can then select and download. Depending on the complexity of the search, these systems may return anything from a handful to an overwhelming number of images. To develop the proposed methodological model, we begin with the assumption that the pre-selections offered by these systems are designed to align with user queries. However, they may also reveal algorithmic biases and reflect the presence of stereotypes, such as those related to gender or race, by prioritising certain visual constructs over others. The working premise is that these platforms deploy trained algorithms to provide users with the most relevant possible results — those that best match their expectations. With some variations, nearly all search interfaces offer functionalities to contextualise and refine searches. Some of these features are technical, such as image size and format; others are commercial, such as copyright-related options. In addition, many platforms allow users to specify preferences related to attributes such as age, ethnicity, or gender of the individuals represented. All of them permit the use of keywords to help obtain more precise results. Let us consider the simplest —and most common— scenario: what happens when a search is conducted using an especially neutral prompt, without any indication of a specific group of people to be represented, and without using the tools available to filter or parameterise the search? How does the system operate to ensure that it meets the client’s expectations? The hypothesis is that the less information provided to the platform, the greater the degree of interpretative responsibility the system must assume to satisfy the query. When a prompt is vague or ambiguous, platforms tend to return numerous results with a wide range of nuances and variations, thereby increasing the likelihood that at least some of them will match the user’s needs. This methodology is specifically designed to provoke such a situation. We argue that this scenario provides the most suitable research environment in which to observe the presence of clichés, biases, and stereotypes within image banks. When a neutral prompt is submitted to a platform —without any additional parameters— the system is compelled to deliver a large and varied selection of responses to accommodate a broad range of potential users. The observation, classification, and characterisation of the variables produced by the system in response to an explicitly neutral request can reveal dominant visual constructs, highly stereotyped patterns of response, and ultimately expose the presence of biased codes and encoding systems. 3. Platform selection, prompt definition, and sample size The commercialisation of stock images takes place in an extremely competitive market, where a wide array of platforms offer their services to both private users and media organisations. These may be either paid or free of charge, and may specialise in either conventional photography or AI-generated imagery. Based on the criterion of professional reputation (Piironen, 2022; 2023; May, 2024), four professional platforms were selected for this research — two specialising in traditional photographic images and two offering AI-generated visuals. It was also ensured in the selection process that each platform is capable of returning at least 50 images in response to a given query, within the visual gallery provided, without requiring the user to reformulate the search or modify the prompt.
152 AI and image banks: A research methodology Pere Freixa; Mar Redondo-Arolas; Lluís Codina; Carlos Lopezosa Digital communication. Trends and good practices Table 1 Selection of professional stoxk image platforms Photographic image banks AI-Generated image banks Shutterstock https://www.shutterstock.com/photos Lexica (Stable diffusion) https://lexica.art/ Getty Images https://www.gettyimages.com/ Adobe Stock (opción IA) https://stock.adobe.com/es/ Source: Own elaboration. To formulate the search queries, we opted to use the natural language prompt functionality offered by these platforms, rather than relying on search equations based on Boolean operators (Battelle, 2006; Abadal & Codina, 2008; Codina, 2018). As previously mentioned, our hypothesis is that the use of neutral prompts activates the algorithmic interpretation systems embedded in these applications, to deliver what the system deems to be the most appropriate response to the request. This approach is considered particularly suitable for detecting potential biases and stereotypes: the results returned by the system are likely to align with what is considered common, normative, or standardised within its database or training corpus. The decision to use natural language prompts instead of search equations is motivated by three main reasons. First, because of their potential to trigger the algorithmic interpretive logic we aim to study. Second, because we are interested in exploring and testing this mode of interaction with image banks, given its emerging and innovative nature. And third, because we wish to analyse what occurs when users, despite sometimes having the option to carry out Boolean searches, choose instead to interact through prompts. To construct the experimental framework for the study, a preliminary test (pre-test) was conducted using two platforms: Lexica and Shutterstock. The objective was to assess the effectiveness of the prompts selected for the study. In this initial phase, six prompts were formulated combining one constant element with two variables. The constant element was the phrase “smiling person”, selected for the following connotations: – Its singular form favours the representation of a single individual. – The term does not imply a specific gender. – The term does not imply a specific age. – The term does not imply a specific ethnicity. – The term does not imply any particular activity. – It is expected that the person will be depicted frontally, despite no framing or compositional instructions being provided. As variables, two differentiating elements were introduced. First, a spatial indicator was added to provide geographical context to the queries. Three alternative settings were selected: city, landscape, and beach. The choice of the term landscape was based on a preliminary comparison between the use of landscape and countryside. As no significant differences were observed in the results returned by the platforms, landscape was selected as the more commonly used term. Second, we introduced variations regarding the formal rendering of the images, with the aim of analysing how AI systems interpret the concept of photographic representation. Three for-
153 AI and image banks: A research methodology Pere Freixa; Mar Redondo-Arolas; Lluís Codina; Carlos Lopezosa Digital communication. Trends and good practices mulations were tested: photography of, picture of, and image with photorealistic look. This comparison seeks to examine not only how the systems interpret and represent degrees of realism or hyperrealism, but also the compositional and framing elements that emerge in the images generated or selected in response to each prompt. The prompts used in the pre-test were as follows: – Picture of a smiling person in a landscape. – Photography of a smiling person in a landscape. – Image of a smiling person in a landscape. Photorealistic look. Similar tests were conducted for the two other settings, city and beach. In all cases, prompts containing the term photography yielded more hyperrealistic results than those using picture or image of (…) photorealistic look. The latter option occasionally produced caricature-like images with clearly distorted features, characteristic of graphic illustration. It was observed that incorporating the term photography in the prompt —added, as explained, to delimit the formal nature of the image in contrast to illustration— may introduce variations in the results, particularly when analysing the “activities” depicted in the scenes. To assess the impact of this variable, additional tests were carried out using two comparable prompts: Photography of a smiling person in the city and Smiling person in the city, applied across the four platforms under analysis. In the case of Shutterstock, including the term photography did indeed result in changes: several images featured people taking photographs, whereas such scenes were scarcely present when the term was omitted. By contrast, Getty Images returned virtually identical results in both cases, with no relevant variations attributable to the use of photography. On the Lexica platform, as previously noted, the term photography functioned as an effective filter to exclude unwanted illustrations from the results. However, its inclusion or omission did not significantly alter any of the analytical parameters under consideration. Similarly, on Adobe Stock, no substantial changes in visual outcomes were observed in relation to the use of the term. As with Lexica, the absence of photography often led to the appearance of illustrations among the retrieved images. In summary, it was found that the inclusion or exclusion of the term photography had no impact on the representation of biases linked to gender, race, age, or beauty-related stereotypes. Only in the case of Shutterstock was a slight deviation detected concerning the actions depicted, attributable to prompts explicitly including the term photography. Consequently, the following three prompts were ultimately submitted to all four platforms: – Photography of a smiling person in a landscape. – Photography of a smiling person in the city. – Photography of a smiling person on the beach. These prompt combinations yielded the analytical sample, composed of a total of 600 images. The sample was generated by collecting the first 50 images returned by each of the four platforms, provided they met the established criteria and constituted a coherent response
154 AI and image banks: A research methodology Pere Freixa; Mar Redondo-Arolas; Lluís Codina; Carlos Lopezosa Digital communication. Trends and good practices to the three proposed queries. The size of this sample —600 images— has previously been employed in similar studies, such as that of Thurlow et al. (2020), which were likewise aimed at identifying biases and stereotypes in digital image banks. 4. What can be observed and how This study is based on the hypothesis that prompts formulated without any explicit reference to the gender, age, or ethnicity of the individuals depicted should, in the absence of cultural or technical bias, yield a proportional distribution across the various observable categories. Any deviation from this theoretical balance is interpreted in the specialist literature as a potential indication of gender disparities and the presence of clichés or stereotypes (Cook & Cusack, 2010; Castillo-Mayén & Montes-Berges, 2014). The overrepresentation of particular groups —whether in terms of gender, phenotype, or age range— may reveal visual repetitions: images that are similar in nature, featuring recurring compositional structures or iconic elements that reflect culturally accepted forms of coding. These, therefore, can be interpreted as stereotypes (Ángeles-Galiano, 2023). In the field of media, such codified visual formulas, tropes, and stereotypes have proven to be highly effective communicative tools for illustrating events. They are part of professional routines in visual production and are especially prevalent in photographic imagery (Baeza, 2001; Freixa y Redondo-Arolas, 2022). Media outlets not only reproduce these visual constructs, but also continually update them through subtle variations, adapting them to the preferences and expectations of their audiences. In doing so, they contribute to the ongoing maintenance and symbolic renewal of these codes (Quin & McMahon, 1997; Novaes-Cirjanic, 2017). The observation of formal and thematic repetitions, as well as the identification of common elements across the images analysed, facilitates the detection and description of the most recurrent stereotypes within the visual corpus. For this purpose, iconological methods are employed, and more specifically, iconographic description is used to identify those attributes potentially associated with visual biases stemming from socio-cultural clichés or stereotypes (Drainville, 2018; Hariman & Lucaites, 2007, 2016; Panofsky, 1979). The application of a systematic analysis sheet is proposed as a methodological tool for assessing the presence of gender bias and stereotypes, structured around the parameters and indicators detailed in the following section. Any image featuring more than one person, clearly depicting a group scene, or lacking a human figure altogether, was excluded from the analysis in order to ensure consistency and comparability of the results. Additionally, the textual information associated with each image was used as support to aid in its classification. 4.1. Parameters and indicators: Biases and stereotypes Below is a description of the indicators proposed for observation, including the definition of each indicator, the observation procedure, and the values used.
155 AI and image banks: A research methodology Pere Freixa; Mar Redondo-Arolas; Lluís Codina; Carlos Lopezosa Digital communication. Trends and good practices 4.1.1. Indicator: Gender – Definition: The prompts are designed to retrieve images that may depict one or more individuals. There is no explicit instruction specifying the gender of the person(s) represented. – Procedure: If the image depicts a single individual, classify them as male, female, or undefined. Record the number of items representing each category. – Values: Female / Male / Not defined. 4.1.2. Indicator: Age – Definition: The prompts are designed to retrieve images that may depict one or more individuals. There is no explicit instruction regarding the age of the person(s) represented. – Procedure: If the image depicts a single individual, classify the age group represented. Record the number of items in each category. – Values: Child / Young person / Adult / Older adult. 4.1.3. Indicator: Ethnicity – Definition: The prompts are designed to retrieve images that may depict one or more individuals. There is no explicit instruction regarding the ethnicity or ethnicities of the person(s) represented. – Procedure: If the image depicts a single individual, classify the ethnicity represented. Record the number of items in each category. – Values: Caucasian / Asian / African American / Latino-Mediterranean / Indigenous American / Middle Eastern. 4.1.4. Indicator: Functional diversity – Definition: The prompts are designed to retrieve images that may depict one or more individuals. There is no explicit instruction as to whether the individuals shown may or may not present any form of functional diversity. – Procedure: If the image depicts a single individual, classify them according to different categories of functional diversity. Record the number of items in each category. – Values: Motor disability / Sensory disability / Intellectual disability / No disability. 4.1.5. Indicator: Beauty standards – Definition: The prompts are designed to retrieve images that may depict one or more individuals. There is no explicit instruction as to whether the individuals shown conform to particular beauty standards or codes of visual representation. – Procedure: If the image depicts a single individual, classify various visual elements that may be associated with beauty standards. The recurrence of such elements will allow for a shared labelling system. Record the number of items in each category. – A priori values: None. – Possible values: Make-up / Sunglasses / Glasses / Hats, caps, helmets / Jewellery / Watches / Summer clothing / Winter clothing / Backpacks / Handbags / Other.
156 AI and image banks: A research methodology Pere Freixa; Mar Redondo-Arolas; Lluís Codina; Carlos Lopezosa Digital communication. Trends and good practices 4.1.6. Indicator: Inventory of actions – Definition: The prompts are designed to retrieve images that may depict one or more individuals in undefined settings. There is no explicit instruction as to whether the individuals should be performing any particular action. Actions may serve as indicators of pre-established and stereotyped behavioural codes. – Procedure: If the image depicts a single individual, classify the actions being performed. The recurrence of elements will enable the establishment of a shared labelling system. Record the number of items in each category. – A priori values: None, although a higher incidence of representations involving photographic devices is expected due to the nature of the prompt. – Possible values: Taking a selfie / Looking at a map / Reading a book / Using a smartphone or tablet / Gesturing / Working / Eating or drinking / Using a camera or technology / Playing / Driving or being in a car / Dancing / Using public transport / Doing sport / Walking / Running / Other. Once the results are obtained, comparisons are proposed for each parameter and indicator across the different platforms to confirm or rule out the presence of biases and stereotypes. 5. Expected results and limitations The process of obtaining the sample is subject to particularities stemming from the algorithmic functioning of the platforms used. Despite the precision that prompts may convey, the systems do not always return images that strictly match the request. In this study, it was found that the result galleries frequently include group photographs, portraits of couples, or images in which no human figure appears at all. Furthermore, traditional photographic stock libraries tend to offer multiple versions of the same scene. These variations typically present minimal differences, usually related to framing or composition. For instance, in response to a prompt requesting a portrait of a woman in an urban street, the system may return five or six virtually identical images, differing only slightly in framing or angle. To construct the final sample, two types of images were removed: so-called false positives — that is, images not corresponding to the requested content— and redundant variations of the same shot, to avoid distortion in the analysis and to ensure greater visual diversity. It was observed that image banks based on artificial intelligence tend to return more accurate results in response to user prompts. By way of illustration, the results showed that AI-based platforms achieved a 60.36% rate of prompt alignment, compared to 44.84% for conventional photographic stock banks. This difference can largely be attributed to the different business models underpinning each type of platform. While traditional banks are based on the exploitation of pre-existing archives —which leads them to offer available images even when they do not fully match the query— AI-driven platforms are designed to generate new content that more closely aligns with the user’s request. As a result, they prioritise prompt accuracy over diversity of results. The proposed methodological system has proven effective for observing the defined indicators. The results obtained (Freixa et al., 2025) reveal small but significant variations in the