Are Games Always Fun and Fair? A Comparison of Reactions to Different Game‐Based Assessments
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Ohlms, Marie Luise; Melchers, Klaus G. Article — Published Version Are Games Always Fun and Fair? A Comparison of Reactions to Different Game‐Based Assessments International Journal of Selection and Assessment Provided in Cooperation with: John Wiley & Sons Suggested Citation: Ohlms, Marie Luise; Melchers, Klaus G. (2025) : Are Games Always Fun and Fair? A Comparison of Reactions to Different Game‐Based Assessments, International Journal of Selection and Assessment, ISSN 1468-2389, Wiley, Hoboken, NJ, Vol. 33, Iss. 1, https://doi.org/10.1111/ijsa.12520 This Version is available at: https://hdl.handle.net/10419/319267 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/
International Journal of Selection and Assessment RESEARCH ARTICLE Are Games Always Fun and Fair? A Comparison of Reactions to Different Game‐Based Assessments Marie Luise Ohlms 1,2 | Klaus G. Melchers 1 1 Department for Work and Organizational Psychology, Institut für Psychologie und Pädagogik, Universität Ulm, Ulm, Germany | 2 Department for Work and Organizational Psychology, Institut für Psychologie, Albert‐Ludwigs‐Universität Freiburg, Freiburg, Germany Correspondence: Marie Luise Ohlms ([email protected]) Received: 27 August 2024 | Revised: 19 December 2024 | Accepted: 6 January 2025 Funding: The research was supported by a doctoral scholarship of the Studienstiftung des deutschen Volkes of the first author. Keywords: applicant reactions | game | game‐based assessment | gamification | individual differences | personnel selection ABSTRACT Game‐based assessment (GBA) has garnered attention in the personnel selection and assessment context owing to its postulated potential to improve applicant reactions. However, GBAs can differ considerably depending on their specific design. Therefore, we sought to determine whether test taker reactions to GBAs vary owing to the different manifestations that GBAs may take on, and to test takers' individual preferences for such assessments. In an experimental study, each of N= 147 participants was shown six different GBAs and asked to rate several applicant reaction variables concerning these assessments. We found that reactions to GBAs were not inherently positive even though GBAs were generally perceived as enjoyable. However, perceptions of fairness and organizational attractiveness varied considerably between GBAs. Participants' age and experience with video games were related to reactions but had less impact than the different GBAs. Our results suggest that a technology‐as‐designed approach, which considers GBAs as a combination of multiple components (e.g., game elements), is crucial in GBA research to provide generalizable results for theory and practice. 1 | Introduction Commercial video games have gained tremendous popularity over the past decades, as they offer players enjoyment, mental stimulation, and stress relief, among other benefits (ESA 2022). The introduction of games into nongame contexts (Deterding et al. 2011) has been an attempt to exploit the advantages of video games for quite some time (e.g., Qian and Clark 2016). Accordingly, this trend has also found its way into the personnel selection and assessment context where it is referred to as game‐based assessment (Landers and Sanchez 2022). One reason for the use of game‐based assessment (GBA) is the difficulty in attracting highly qualified professionals to join an organization, which has become an increasingly challenging task for many organizations amid the prevailing skills shortage (Brunello and Wruuck 2021). In light of this difficulty, GBA is suggested as a way to support recruitment and selection by providing both a positive candidate experience and a valid measurement of applicants' relevant knowledge, skills, abilities, and other attributes (KSAOs, Bhatia and Ryan 2018). Given their engaging and gamified nature (Deterding et al. 2011), a possible advantage of GBAs claimed in the literature (e.g., Bhatia and Ryan 2018) is their potential to evoke more positive applicant reactions compared to traditional valid assessment methods. However, GBAs can differ tremendously, as their design offers flexibility owing to the variety of game elements, 1 such as storylines, levels, badges, avatars, and genres that can be chosen (Fetzer, McNamara, and Geimer 2017; Ohlms 2024). As a consequence of the considerable variability in GBA designs, two relevant questions arise: The first is whether This is an open access article under the terms of the Creative Commons Attribution License, which permits use, distribution and reproduction in any medium, provided the original work is properly cited. © 2025 The Author(s). International Journal of Selection and Assessment published by John Wiley & Sons Ltd. 1of16International Journal of Selection and Assessment, 2025; 33:e12520 https://doi.org/10.1111/ijsa.12520
applicant reactions to a specific GBA are comparable to reactions to another GBA? Second, one can ask whether certain individuals are more or less likely to react positively to GBAs, that is, whether there are individual characteristics associated with reactions to GBAs? Previous studies of GBAs have been unable to answer these questions because they have usually followed a technology‐as‐ causal (i.e., treating GBAs as a single entity) or a technology‐ as‐instrumental paradigm (i.e., treating GBAs as a single entity but considering potential interacting effects with other exogenous variables, Landers and Marin 2021). These studies compared a given GBA to its traditional counterpart (e.g., Landers et al. 2022; Ohlms, Melchers, and Kanning 2024a), sometimes considering potential moderators or mediators (e.g., Gkorezis et al. 2021), but did not compare different GBAs. Given that previous research has yielded mixed results regarding applicant reactions to GBAs (e.g., Landers et al. 2022; Ohlms, Melchers, and Kanning 2024a), the question arises as to whether a technology‐as‐designed perspective (Landers and Marin 2021),whichdecomposesGBAsintotheir specific design features and examines their influence on outcomes while considering other potentially influential exogenous variables, would be more appropriate. Thus, from a theoretical perspective, it is important to compare the same individuals' reactions to different GBAs to see whether reactions to GBAs are similar across different GBAs, suggesting a technology‐as‐causal perspective, or whether they vary across different GBAs, suggesting a technology‐as‐designed perspective. From an applied perspective, it would also be useful to know whether applicant reactions are always affected in the same way independent of the used GBA and whether applicants' personal characteristics are additionally associated with their reactions to GBAs. Such knowledge would enable organizations to make informed decisions about whether and what type of GBAs to employ to promote positive applicant reactions. To address these gaps, and following the call from a recent review by Ramos‐Villagrasa, Fernández‐Del‐Río and Castro (2022) to investigate factors that impact applicant reactions to GBAs, the present study represents a first step and employed a within‐subjects design to directly compare various test taker reactions to six rather different GBAs. Additionally, we aimed to examine how individual characteristics of potential applicants relate to their reactions to GBAs, building on Landers and Sanchez's (2022) proposed extension of Hausknecht, Day, and Thomas's (2004) applicant reaction model. Finally, given existing evidence of gender and age differences in preferences for commercial video games (ESA 2019,2022), we aimed to investigate whether these differences also exist in the context of GBAs. Hence, the present study provides insights for GBA research and practice by determining whether test takers react differently to different GBAs. This would allow a test whether a technology‐as‐causal or a technology‐as‐designed perspective is more appropriate when studying GBAs, thus providing important avenues for future GBA research. 2 | Game‐Based Assessment and Game Elements GBA refers to the use of (video‐)games as personnel selection instruments to measure applicants' relevant KSAOs for a certain position (Landers and Sanchez 2022). GBA serves as a stand‐ alone method, aiming to put applicants into a psychological state of gameful experience (Landers and Sanchez 2022). GBAs are designed through game development, which, owing to their often complex, sophisticated, and multimodal nature, typically involves higher costs than the development of more traditional assessment methods (Bhatia and Ryan 2018; Landers and Sanchez 2022). When crafting such playful assessments, one can generally make use of any game element or any combination of different elements that are employed in commercial games such as leaderboards (i.e., an indication about how well one performs compared to others), points (i.e., rewards awarded for the successful completion of specific tasks or the accomplishment of predefined targets), and/or a narrative (i.e., the storyline of the game) to name just a few. In addition to various game elements, the sphere of game genres also offers ample freedom in GBA development. A genre represents a way to categorize games based on their game elements, structure, challenge, and interactivity (Fetzer, McNamara, and Geimer 2017). Despite a plethora of different game genre classifications that has emerged to categorize games for entertainment purposes (e.g., King and Krzywinska 2002;Rollingsand Adams 2003), in the personnel selection context, six game genres are predominantly used: action, simulation, role‐ playing, adventure, mini‐game, and strategy (see Table 1in Fetzer, McNamara, and Geimer 2017, for a definition of the GBA genres). The combination of different game characteristics (e.g., game elements, game fiction) that define a game and contribute to a playful experience as well as to other desired outcomes (e.g., increased motivation) is defined differently across game taxonomies which are primarily determined by variations in their intended outcomes (Landers et al. 2018). Yet, game taxonomies specific to the personnel selection and assessment context are scarce, with one exception being the proposed GBA taxonomy by Hawkes, Cek and Handler (2017). This taxonomy aims to categorize GBAs based on seven categories. These categories are (1) fidelity (i.e., the degree of Summary •This study examines whether people vary in their reactions to different game‐based assessments (GBAs), exploring comparability among them. •Results suggest that potential applicants' reactions can differ considerably across different GBAs. •Participants' age correlated negatively with reactions to GBAs, while experience with video games was positively associated with reactions. •Depending on their age potential applicants differed in their preferences for certain GBAs. •Varying applicant reactions across different GBAs stress the need for a thoughtful choice of GBAs for personnel selection. 2of16 International Journal of Selection and Assessment, 2025
freedom when playing the GBA), (2) conflicting demands (i.e., the extent to which a player is confronted with multiple demands simultaneously), (3) variable path (i.e., the extent to which game progression is affected by actions taken during the GBA), (4) engagement (i.e., the extent to which the GBA has elements that users find enjoyable or that promote immersion), (5) suspensefulness (i.e., the extent to which a desired outcome is likely to occur and how it can be influenced by the user's actions), (6) gamefulness (i.e., the extent to which the assessment has elements and interactivity typical of a game), and (7) fidelity (i.e., the extent to which a game is reality‐based and job‐related). Thus, for example, one could apply Hawkes et al.'s taxonomy to a flight simulator GBA used for the selection of pilots. Following the taxonomy, one would rate such a GBA high on all dimensions except for gamefulness, as a flight simulator typically does not contain conventional game elements (see also Hawkes, Cek, and Handler 2017). 2.1 | Applicant Reactions to Game‐Based Assessments Given the higher development costs and programming complexity of GBAs compared to traditional selection methods (Landers and Sanchez 2022), one might question why organizations should incur these expenses to integrate GBAs into their selection processes. However, there are several benefits of GBAs that are postulated in the literature (e.g., Bhatia and Ryan 2018) and by GBA vendors. The main advantage of this playful method is claimed to lie in improvements of applicant reactions and candidate targeting through the entertaining and playful nature of these assessments (Bhatia and Ryan 2018). With many organizations that are currently struggling to find suitable talent to join their workforce (Brunello and Wruuck 2021), hopes are high that GBAs can improve the candidate experience, and thus provide a strategic advantage as applicants' perceptions of assessment procedures are linked to their attitudes to the organization and also to their job offer acceptance intentions. To examine whether GBAs do indeed have the hoped‐for positive effects on applicant reactions compared to their traditional counterparts and to offer generalizable implications for GBA research, researchers must be clear about the paradigmatic approach they are using to study GBAs, and which approach is appropriate to study this technology‐based selection method 2 (Landers and Marin 2021). Looking at the initial studies that have examined applicant reactions to GBAs, it becomes clear that these studies have tended to adopt a technology‐as‐causal (e.g., Ohlms et al. 2024)or technology as‐instrumental approach (e.g., Gkorezis et al. 2021; Ohlms, Melchers, and Kanning 2024a). According to Landers and Marin, in a technology‐as‐causal approach, researchers examine the impact of a technology (e.g., a GBA) on an outcome of interest (e.g., applicant reactions). Thus, in such a technology‐as‐causal paradigmatic approach, GBAs are treated as a single entity without regard to potential differences in their design. In a technology‐as‐instrumental approach, GBAs would still be considered as a uniform entity, but this approach considers potential interacting effects with other exogenous variables (e.g., applicant reactions to GBAs differ based on test takers' video game experience, Landers and Marin 2021). However, using a technology‐as‐causal or technology‐as‐ instrumental approach to examine the effects of GBAs on relevant personnel selection outcomes (e.g., applicant reactions, validity, reliability) may compromise the external validity of such research. This is because, as mentioned above, GBAs can vary considerably regarding their design (i.e., game elements or game genre) as well as the construct (s) they intend to measure (Fetzer, McNamara, and Geimer 2017). Given the diverse nature of GBAs, hardly any two GBAs are alike. Consequently, it seems unlikely that reactions to a specific GBA are comparable to those to another GBA. Moreover, the use of a GBA may not inherently result in applicants perceiving the selection process as more enjoyable, entertaining, and fair compared to traditional methods. In fact, the favorable outcomes for applicant reactions to a GBA are likely dependent on an effective game development (Landers and Sanchez 2022). Against this background, it may be argued that a technology‐ as‐designed approach may be more appropriate when TABLE 1 | Game‐based assessment genres. Genre Description GBA used in the present study Action Games demanding rapid and precise player reactions to tasks or stimuli US Air Force: Airman Challenge Simulation Games that emulate scenarios and tasks of a (semi‐)realistic virtual environment ReQiu: Workday simulation Role‐playing Games in which players command the actions of one or more characters and perform tasks on their behalf LTP: Building Docks Adventure Games in which players embark on a (virtual) adventure in which they face challenges and interact with avatars and/or objects Ohlms et al. (2024): Minecraft game Mini‐game Games of short duration, which can usually be solved within 10 min and are typically simple in their objectives Criteria Corp: Cognify; Landers et al. (2022) Strategy Games requiring strategic planning and problem‐solving skills to accomplish specific tasks in a virtual environment McKinsey Solve: Ecosystem Creation 3of16
examining the effects of GBAs and aiming to draw sound conclusions for GBA research and practice. So far, research has not yet clarified which paradigmatic approach is more appropriate for studying GBAs. If a technology‐as‐causal or technology‐as‐instrumental approach is sufficient for investigating GBAs, then applicant reactions to different GBAs should yield consistent results. This would allow GBAs to be treated as a single entity that always produces certain effects on relevant outcomes (i.e., applicant reactions) and to generalize from the results of one GBA to those of another. However, if applicant reactions differ from GBA to GBA, this would argue for a technology‐as‐designed approach. That is, looking at the effects of a GBA as a whole on applicant reactions would be of limited value and contribute little to theory and practice beyond a specific GBA. To determine whether applicant reactions to different GBAs are the same or different, it is necessary to have participants evaluate multiple GBAs. However, previous studies have only examined a single GBA at a time. Nevertheless, these initial studies of test takers' reactions to GBAs at least indirectly support the notion that reactions may differ between different GBAs. For instance, Landers et al. (2022) found more positive reactions to the GBA Cognify, which comprises several mini‐ games, compared to a paper‐pencil test measuring similar abilities. Conversely, Ohlms, Melchers and Kanning (2024a) reported more negative reactions for a GBA implemented in the Minecraft development environment and categorized in the adventure genre, compared to its traditional counterpart. Similarly, adding game elements to otherwise non‐gamified tests led to more positive reactions in some studies (e.g., Georgiou and Nikolaou 2020) but not in others (e.g., Ohlms et al. 2024). These findings suggest that a technology‐as designed approach for studying GBAs would be most appropriate. Furthermore, Gilliland (1993) theoretical model on applicant reactions, as well as its extensions by Hausknecht, Day and Thomas (2004) and Landers and Sanchez (2022), may also help to explain which paradigmatic approach may be appropriate for studying GBAs. Specifically, Hausknecht et al. posit various determinants (e.g., job or person characteristics) that may affect applicant reactions of the selection process (e.g., procedural justice perceptions or test anxiety), which, in turn, are assumed to result in individual and organizational outcomes such as the perceived organizational attractiveness or applicants' job offer acceptance intentions. Additionally, Hausknecht et al.'s model suggests several moderators that may influence these relationships. In the context of GBAs, Landers and Sanchez (2022)proposed an extension to Hausknecht, Day, and Thomas's (2004)model by adding game elements and game‐related moderators. Specifically, Landers and Sanchez introduced game elements such as leaderboards or storylines as antecedents that are intended to influence perceived procedural characteristics (e.g., justice rules, i.e., aspects that affect the perceived fairness of the selection process and decision, such as the perceived job‐relatedness or applicants' opportunity to perform), which, in turn, should influence applicant perceptions (e.g., perceived fairness) and related outcomes (e.g., attitudes toward the organization). Furthermore, in the context of GBAs, Landers and Sanchez extended Hausknecht et al.'s model by a series of game‐relatedmoderatorssuchasthe effectiveness of the GBA design. Thus, the extended model explains how individual game elements and their effective integration may influence applicants' perceptions of the justice rules. For instance, applicants might perceive a GBA in which they have to manage a fictional organization (e.g., Melchers and Basch 2022) as more job‐related than a traditional cognitive ability test that requires them to solve matrices. However, the opposite might be the case when applicants have to complete a GBA in which they have to save a cursed country (e.g., Ohlms, Melchers, and Kanning 2024a). Therefore, it is reasonable to assume that applicant reactions differ depending on the particular GBA. As argued above, to draw meaningful and generalizable conclusions for GBA research and practice, it is necessary to determine whether a technology‐as‐causal/technology‐as‐ instrumental or a technology‐as‐design approach should be taken to study GBAs, as this has fundamental implications for study designs and research directions. To answer this question, we will compare applicant reactions to different GBAs and to examine the following research question: Research Question 1 (RQ 1). Does the specific GBA influence applicants' reactions toward it? 3 2.2 | Individual Differences and Applicant Reactions to Game‐Based Assessments Not only might the GBA type influence applicant reactions, but also individual differences in preferences for games. Landers and Sanchez's (2022) extended applicant reactions model introduced the idea that applicants' skills and prior experience with game elements act as game‐related moderators to influence their reactions to GBAs. Thus, individuals who have more experience with video games and technical systems (e.g., computers) might have a higher computer self‐efficacy in dealing with computer‐based games and might react more positively to computer‐based GBAs, because they are already familiar with games from their leisure time. In line with this, previous research found that technology self‐efficacy and video game experience were positively associated with test takers' reactions to GBAs (Ellison et al. 2020; Gkorezis et al. 2021; Ohlms, Melchers, and Kanning 2024a). Likewise, studies examining applicant reactions to technology‐based selection instruments in general found that experience with computers were associated positively with applicant reactions to technology‐based selection procedures (Wiechmann and Ryan 2003). In addition, computer self‐efficacy appears to have a positive association with the perceived ease of use of digital systems (e.g., Oostrom et al. 2013). Thus, we hypothesize: Hypothesis 1. Video game experience is positively associated with applicant reactions to GBAs. Hypothesis 2. Computer self‐efficacy is positively associated with applicant reactions to GBAs. 4of16 International Journal of Selection and Assessment, 2025
Assuming that video game experience and computer self‐ efficacy influence reactions to GBAs, it seems relevant to consider how these explanatory variables are related to gender and age. For commercial video games, differences in preferences for video games in general, as well as for certain types of games in particular, were found based on gender and age. Specifically, women generally tend to spend less time playing games and demonstrate lower motivation toward gaming activities (e.g., González‐González et al. 2022; Hartmann and Klimmt 2006;Lucas and Sherry 2004). Moreover, there are gender differences in preferences for different game genres and elements (e.g., Lucas and Sherry 2004) so that men engage more with shooter, sports, and action games than women do, who instead tend to play more casualgameslikepuzzleandcardgamesthandomen(ESA2019). Research further indicates that women particularly tend to reject aggressive and violent content, sexualized portrayals of female characters in video games as well as a lack of social interaction (e.g., with avatars, Hartmann and Klimmt 2006). Additionally, women are less attracted to competitive game elements in video games than men are (Hartmann and Klimmt 2006). Initial studies in the context of GBAs also indicated that applicant reactions to GBAs vary based on applicants' gender. Hence, women tend to react less positively to GBAs compared to men (Ellison et al. 2020; Ohlms, Melchers, and Kanning 2024a). In terms of age, it may be assumed that test takers' age may be related to more negative reactions to GBAs owing to lower computer self‐efficacy and less video game experience among older individuals (Reed, Doty, and May 2005; Williams, Yee, and Caplan 2008). Regarding commercial video games, a representative survey by the Entertainment Software Association (ESA 2022) revealed that the majority of players (60%) are below the age of 35. Conversely, only 9% are between 55 and 64 years old, and a mere 4% are older than 65 years, suggesting that older people generally enjoy video games less than younger individuals. Furthermore, age is associated with game genre preferences, with individuals aged 18 to 34 tending to prefer action, shooter, and puzzle games, while those over 65 years tending to prefer puzzle as well as skill and chance games (ESA 2022). Based on the described findings, we aim to examine the following hypotheses: Hypothesis 3. Men show more positive reactions to GBAs than women. Hypothesis 4. Age is negatively associated with applicant reactions to GBAs. Hypothesis 5. GBA type moderates (a)the relationship between gender and applicant reactions and (b)the relationship between age and applicant reactions. 3 | Methods 3.1 | Stimulus Material The selection of the six different GBAs that participants had to evaluate within the study was based on a comprehensive search of GBAs developed for both commercial and research purposes. To identify these GBAs, we conducted an extensive search across multiple sources. Specifically, we conducted an online search to identify commercially available GBAs, such as those offered by assessment providers (e.g., Criteria Corp, Aon). In addition to commercial and publicly available GBAs, we reviewed the scientific literature on gamification and GBAs to further identify GBAs developed for research purposes. This search led to the identification of 78 different playful assessments (when a certain assessment consisted of several mini‐games/subgames they werecountedasmultipleseparateGBAs,e.g.,thesixmini‐games of the GBA Cognifiy [www.criteriacorp.com]wereconsideredas individual GBAs). To identify potential differences between the playful assessments found during our search, all assessments were coded by two independent coders (a Master level and a PhD student specializing in work and organizational psychology) based on two criteria. On the one hand, the coders assessed which game genre a particular procedure could be assigned to. The categorization of genre was based on Fetzer et al.'s (2017) classification, thus the playful procedures could be assigned to one of six different genres (action, adventure, role‐playing, simulation, strategy, mini‐game). On the other hand, to further classify the differences between identified proceduresbeyondgenre,weusedHawkes et al.'s (2017) taxonomy of GBAs. As mentioned above, this taxonomy comprises seven characteristics that can be used to categorize playful procedures. For example, according to Hawkes et al.'s taxonomy, a GBA in which test takers can move freely in a virtual game environment using the control keys would be rated high on the dimension freedom of action. In contrast, a multiple‐ choice questionnaire would represent a low level on this dimension. To reveal finer differences between GBAs, we used a 5‐point scale from 1 = low level (with a specific example for a low level for each dimension) to 5 = high level (with a specific example for a low level for each dimension). Based on the genre classifications and ratings, we selected six different GBAs for our study. When choosing the GBAs, we aimed to select procedures that differed in terms of both their genres and their classification according to Hawkes et al.'s (2017) taxonomy. Thus, from each of the six genres we picked one GBA, which further differed in terms of Hawkes et al.'s classification (see Table 2for the coding according to genre and Hawkes et al.'s taxonomy). The aim was to bring together a diverse range of GBAs for our study. These GBAs (see below for more information) were the (1) Airman Challenge, (2) Workday simulation, (3) Building Docks, (4) Minecraft game, (5) Grid Lock, and (6) McKinsey's Ecosystem Creation. For each of the GBAs, we wrote a text‐based description with the goal, storyline, individual tasks, game elements, and basic rules of the game. In addition, an excerpt of each game was shown in short (approx. 2‐min) videos to give participants a realistic impression of the game. These videos were edited from open‐access sources (e.g., YouTube or companies' website) and accompanied by a self‐recorded audio sequence that explained the GBA's unique features in detail. 3.1.1 | Airman Challenge The airman challenge falls under the action genre and is used by the US Air Force (https://www.airforce.com/airmanchallenge/) 5of16
as part of their recruitment process. In the airman challenge, applicants have to complete various missions around the world, taking on different roles within various air force job profiles. The goal of the game is to successfully complete as many missions as possible to gain ranks, badges, and achievements. Within each mission, several tasks have to be solved. Thereby, players have to slip into different air force job profiles and master their individual tasks. 3.1.2 | Workday Simulation The Workday simulation, developed by ReQiu (https://reqiu. eu) belongs to the simulation genre, simulating a trial day within the company ReQiu, during which test takers encounter various work‐related challenges. These tasks have to be completed and solved within a simulated office. For instance, one task involves scheduling workdays with the goal of scheduling as many workdays as possible. 3.1.3 | Building Docks Building Docks is categorized within the role‐playing genre. Applicants take on the role of a port owner and together with three other (computer‐controlled) characters have to turn a newly opened port profitable within a year. In doing so, they have to negotiate with various stakeholders (see Barends, De Vries, and Van Vugt 2022, for a comprehensive description). 3.1.4 | Minecraft Game The Minecraft game, developed by Ohlms, Melchers and Kanning (2024a), is from the adventure genre. In this GBA, applicants have to rescue a cursed country by completing various cognitive tasks, such as solving matrices (see Ohlms, Melchers, and Kanning 2024a, for details). 3.1.5 | Grid Lock Grid Lock is one of several mini‐games from the GBA Cognify, developed by Criteria Corp (https://www.criteriacorp.com). In Grid Lock, applicants have to insert various puzzle pieces into a given shape as quickly as possible (see Landers et al. 2022, for a comprehensive description). 3.1.6 | Ecosystem Creation McKinsey's Ecosystem Creation game, called Solve (https:// www.mckinsey.com), represents a strategy game. The goal of this GBA is to build a sustainable ecosystem in a coral reef. The main task is to establish a food chain with eight plant and animal species, in which each species finds enough food and suitable living conditions. Within the GBA, test takers receive a variety of information about different plant and animal species, from which they have to choose eight species and select the appropriate habitat for them. TABLE 2 | Coding of the six GBAs according to Hawkes (2017) GBA taxonomy. GBA used in the present study Genre GBA characteristics according to Hawke's GBA taxonomy Freedom of action Conflicting demands Variable paths Engagement Suspensefulness Gamefulness Fidelity Airman Challenge Action 5.0 3.0 4.5 5.0 4.5 5.0 4.5 Workday simulation Simulation 4.5 2.0 3.5 4.5 4.0 4.5 4.0 Building Docks Role‐playing 3.5 3.0 4.5 4.0 5.0 4.5 2.0 Minecraft game Adventure 4.5 2.5 3.0 4.5 3.0 5.0 1.0 Cognify: Grid Lock Mini‐game 2.5 1.0 1.5 3.0 2.5 3.0 1.0 Ecosystem Creation Strategy 4.5 5.0 5.0 4.5 5.0 5.0 1.0 Note: GBA characteristics were rated on a five‐point rating scale ranging from 1 = very low to 5 = very high. Values represent the mean rating of the two independent raters. 6of16 International Journal of Selection and Assessment, 2025
3.2 | Procedure and Sample The study used an experimental within‐subjects design. Thus, within the study, each participant rated all GBAs on six reaction variables (see below). We recruited German speaking participants via direct contacts (e.g., word of mouth, WhatsApp, e‐mail), social media (e.g., LinkedIn), and a departmental participant pool. Inclusion criterion for participation was correctly answering an attention check item (see below). After obtaining participants' informed consent, they were told to imagine that they were searching for a new job and had applied for an office position at six different organizations, which were all equally attractive to them in terms of payment and job tasks. Furthermore, participants were told that they had been invited to take an online assessment at each of the six organizations. Subsequently, participants were presented with the six different online assessments (i.e., GBAs) in randomized order. Participants did not play the GBAs themselves but were provided with a text‐based explanation as well as with a video of each GBA. Within each explanation text and video, the goal, starting situation, individual tasks, game elements, and basic rules of the specific GBA were described in detail. Additionally, each video aimed to provide participants with a comprehensive understanding regarding the visual representation of a particular GBA and its associated game elements. After each respective GBA was described using a video and explanation text, participants evaluated six different applicant reaction variables concerning this specific GBA. Participants' demographic data as well as their computer self‐efficacy, and video game experience were measured at the end of the survey. To determine the required sample size to test our research question with a power of 0.80, we conducted an a priori power analysis using G*Power (Faul et al. 2007). To be able to detect a small effect of f= 0.10 (corresponding to a dof 0.20) for a one‐ way multivariate analysis of variance with repeated measures (MANOVA; number of measurements = 6; assumed correlation among repeated measures = 0.50), this analysis revealed a required sample size of N= 114. Hypotheses 1to 5were analyzed using multilevel modeling, for which the concern about the sufficient sample size generally refers to the higher level (i.e., in our analyses to the number of participants). With a minimum sample size of 120 participants (i.e., at Level 2), we could expect sufficient power of our multilevel model (Maas and Hox 2005). Accordingly, we aimed to collect data from at least 120 participants. The initial sample even consisted of 160 participants. Yet, before the data analysis, we excluded 13 participants who failed the attention check, which led to a final sample of N=147. Inthis sample, 56.46% self‐reported as female, 42.86% as male, and 0.68% as diverse. Participants' mean age was 30.24 years (SD = 10.46) and 2.04% reported a secondary school certificate as their highest educational degree, 7.48% an intermediate school‐ leaving certificate, 24.49% a general qualification for university entrance (i.e., German Abitur), 43.54% a bachelor's degree, 21.77% a master's degree, and 0.68% a PhD. Of the participants, 44.90% were employed at the time of study (including 0.68% who were selfemployed). 44.90% were employed at the time of the study. 3.3 | Measures All items were measured using a five‐point Likert scales ranging from 1 = strongly disagree to 5 = strongly agree. Supporting Information Table S1 in the online supplement lists all items for the different scales and the respective item sources. 3.3.1 | Applicant Reaction Variables The different applicant reaction variables for a given GBAs were administered immediately after the GBA was presented. We measured perceived job‐relatedness (two items; mean α= 0.91), procedural fairness (three items; mean α= 0.92), and opportunity to perform (four items; mean α= 0.91) with three subscales of the Selection Procedural Justice Scale (Bauer et al. 2001) in the German translation by Manzey and Gurk (2005). Face validity (mean α=0.85)was assessed using a three‐item scale from Oostrom et al. (2013) using the German translation by Ohlms, Melchers and Kanning (2024a). To measure organizational attractiveness (mean α= 0.91), we used five items from Highhouse, Lievens and Sinar (2003) in the German translation by Basch, Melchers and Büttner (2022). Finally, enjoyment of the assessment (mean α= 0.92) was captured with three items from Wilde et al. (2009) using the German translation by Ohlms, Melchers and Kanning (2024a). To assess whether the six reaction variables indeed comprised separable constructs, we conducted a confirmatory factor analysis (CFA) with six correlated factors. This CFA showed an adequate fit across the different GBAs, mean χ 2 (155) = 295.06, p<0.001;meanCFI=0.95;mean SRMR = 0.06; mean RMSEA = 0.08. In contrast, a single‐ factor model had a poor fit, mean χ 2 (170) = 1060.14, p< 0.001; mean CFI = 0.68; mean SRMR = 0.10, mean RMSEA = 0.19. In addition, the six‐factor model fitted significantly better than the single‐factor model, all Δχ 2 s (15) > 598.00, all ps < 0.001. 3.3.2 | Individual Difference Variables In addition to the demographic variables, we used five items (α= 0.93) from Bourgonjon et al. (2010) to measure participants' video game experience. Furthermore, we used eight items from Wiechmann and Ryan (2003) to assess computer self‐ efficacy (α= 0.85). 3.3.3 | Attention Check To assess whether participants conscientiously completed the survey, we inserted an attention check item in between the applicant reaction items: “This is a test of your attention. Therefore, please tick ‘strongly agree’to show that you have read everything carefully.”As noted above, participants who failed to answer this item correctly were excluded from all analyses (i.e., 13 participants). 7of16
4 | Results Means, standard deviations, and intercorrelations between variables are presented in Table 3. Consistent with previous research (e.g., Gkorezis et al. 2021), age was negatively associated with video game experience (r=−0.18, p= 0.03). Furthermore, as in previous studies, men scored higher on video game experience (r= 0.43, p< 0.001) and computer self‐efficacy (r= 0.31, p< 0.001) than did women (e.g., Gkorezis et al. 2021). 4.1 | Reactions to the Different GBAs Our first objective was to assess whether the different GBAs evoked different reactions in test takers (RQ 1). To examine this research question, we conducted a series of one‐way analyses of variance (ANOVAs) with repeated measures. The independent variable was the GBA and the dependent variables were the six applicant reaction variables. As can be seen in Table 4, in all ANOVAs the GBAs had a significant effect on each of the six applicant reaction variables, all Fs > 5.22, all ps < 0.05. Furthermore, as can be seen in Figure 1, with the exception of enjoyment, the general pattern of results was similar for all the other applicant reaction variables. For these other variables, the corresponding η²s in Table 4also suggest that the effects of the GBAs were large, and subsequent post‐hoc tests using the Bonferroni procedure revealed various significant differences. Specifically, the same pattern of results was found for procedural fairness, perceived job‐relatedness, opportunity to perform, and face validity: The Workday simulation and Building Docks were rated most positively and significantly higher than the Ecosystem Creation, which was rated significantly higher than the Airman Challenge, the Minecraft game, and Grid Lock. A similar pattern of results emerged for organizational attractiveness, with the only difference being that ratings for the Ecosystem Creation and Minecraft game did not differ significantly. In particular, for all these applicant reaction variables there were medium to large differences between the Workday simulation on the one hand and the Airman Challenge (d= 0.77 to 1.12), Minecraft (d= 0.77 to 1.37), and Grid Lock (d= 0.62 to 1.41) on the other hand. Similarly, we found medium to large differences between Building Docks on the one hand and the Airman Challenge (d= 0.72 to 1.06), Minecraft (d= 0.66 to 1.31), and Grid Lock (d= 0.65 to 1.35) on the other hand. In contrast to all the previous applicant reaction variables, there were smaller differences between the different GBAs with regard to the perceived enjoyment of the GBAs. Specifically, all GBAs were rated positively (i.e., had a mean > 3) on this reaction variable and only very few GBAs differed significantly according to the post‐hoc tests (see Table 3and Figure 1). However, there were significant differences in enjoyment ratings between Building Docks and the Airman Challenge (d= 0.28), Building Docks and Grid Lock (d= 0.46), as well as the Minecraft game and Grid Lock (d= 0.31). Taken together, our results suggest that potential applicants perceived the different GBAs as relatively enjoyable. In contrast TABLE 3 | Descriptive information and correlations for age, gender, video game experience, and computer self‐efficacy. Variable MSDSD within 123456789 1. Procedural fairness 2.96 0.69 0.73 0.70*** 0.69*** 0.69*** 0.68*** 0.37*** 2. Job‐relatedness 2.57 0.61 0.86 0.76*** 0.82*** 0.80*** 0.72*** 0.40*** 3. Opportunity to perform 2.60 0.61 0.76 0.64*** 0.81*** 0.78*** 0.75*** 0.51*** 4. Face validity 2.72 0.55 0.93 0.59*** 0.76*** 0.70*** 0.73*** 0.40*** 5. Organizational attractiveness 2.95 0.64 0.71 0.73*** 0.74*** 0.67*** 0.62*** 0.59*** 6. Enjoyment 3.50 0.77 0.77 0.56*** 0.50*** 0.53*** 0.40*** 0.73*** 7. Age 30.24 10.46 −0.31*** −0.08 −0.11 −0.06 −0.19*−0.28*** 8. Gender a 0.43 0.50 0.02 −0.02 −0.02 −0.08 −0.04 0.02 0.10 9. Video game experience 2.13 1.08 0.13 0.12 0.20*0.17*0.18*0.26** −0.18*0.43*** 10. Computer self‐efficacy 4.40 0.53 −0.06 −0.14 −0.14 −0.14 −0.03 0.15 −0.11 0.31*** 0.12 Note: Correlations below the diagonal are person‐level correlations. To calculate these person‐level correlations, GBA‐level data below the diagonal were averaged across GBA (N= 147). Correlations above the diagonal are GBA‐level correlations (N= 882). Above the diagonal, GBA‐level data were centered around the respective person mean. Gender is coded 0 = female, 1 = male. a n= 146. *p< 0.05; **p< 0.01; ***p< 0.001. 8of16 International Journal of Selection and Assessment, 2025
Gkorezis, P., K. Georgiou, I. Nikolaou, and A. Kyriazati. 2021. “Gamified or Traditional Situational Judgement Test? A Moderated Mediation Model of Recommendation Intentions via Organizational Attractiveness.”European Journal of Work and Organizational Psychology 30, no. 2: 240–250. https://doi.org/10.1080/1359432X.2020.1746827. González‐González, C. S., P. A. Toledo‐Delgado, V. Muñoz‐Cruz, and J. Arnedo‐Moreno. 2022. “Gender and Age Differences in Preferences on Game Elements and Platforms.”Sensors 22, no. 9: 3567. https://doi. org/10.3390/s22093567. Hartmann, T., and C. Klimmt. 2006. “Gender and Computer Games: Exploring Females' Dislikes.”Journal of Computer‐Mediated Communication 11, no. 4: 910–931. https://doi.org/10.1111/j.10836101.2006.00301.x. Hausknecht, J. P., D. V. Day, and S. C. Thomas. 2004. “Applicant Reactions to Selection Procedures: An Updated Model and Meta‐ Analysis.”Personnel Psychology 57, no. 3: 639–683. https://doi.org/10. 1111/j.1744-6570.2004.00003.x. Hawkes, B., I. Cek, and C. Handler. 2017. “Next Generation Technology‐enhanced Assessment.”In The Gamification of Employee Selection Tools, edited by J. C. Scott, D. Bartram, and D. H. Reynolds, 288–314. Berlin: Cambridge University Press. https://doi.org/10.1017/ 9781316407547.013. Highhouse, S., F. Lievens, and E. F. Sinar. 2003. “Measuring Attraction to Organizations.”Educational and Psychological Measurement 63, no. 6: 986–1001. https://doi.org/10.1177/0013164403258403. Hox, J., and L. Wijngaards‐de Meij. 2015. “The Multilevel Regression Model.”In The Sage Handbook of Regression Analysis and Causal Inference, edited by H. Best and C. Wolf, 133–152. Sage. King, G., and T. Krzywinska. 2002. ScreenPlay: Cinema/Videogames/ Interfaces. New York, NY: Wallflower Press. Landers, R. N., M. B. Armstrong, A. B. Collmus, S. Mujcic, and J. Blaik. 2022. “Theory‐Driven Game‐Based Assessment of General Cognitive Ability: Design Theory, Measurement, Prediction of Performance, and Test Fairness.”Journal of Applied Psychology 107, no. 10: 1655–1677. https://doi.org/10.1037/apl0000954. Landers, R. N., E. M. Auer, A. B. Collmus, and M. B. Armstrong. 2018. “Gamification Science, its History and Future: Definitions and a Research Agenda.”Simulation & Gaming 49, no. 3: 315–337. https://doi. org/10.1177/1046878118774385. Landers, R. N., and S. Marin. 2021. “Theory and Technology in Organizational Psychology: A Review of Technology Integration Paradigms and Their Effects on the Validity of Theory.”Annual Review of Organizational Psychology and Organizational Behavior 8: 235–258. https://doi.org/10.1146/annurev-orgpsych-012420-060843. Landers, R. N., and D. R. Sanchez. 2022. “Game‐Based, Gamified, and Gamefully Designed Assessments for Employee Selection: Definitions, Distinctions, Design, and Validation.”International Journal of Selection and Assessment 30, no. 1: 1–13. https://doi.org/10.1111/ijsa.12376. Lievens, F., and P. R. Sackett. 2017. “The Effects of Predictor Method Factors on Selection Outcomes: A Modular Approach to Personnel Selection Procedures.”Journal of Applied Psychology 102, no. 1: 43–66. https://doi.org/10.1037/apl0000160. Lucas, K., and J. L. Sherry. 2004. “Sex Differences in Video Game Play: A Communication Based Explanation.”Communication Research 31, no. 5: 499–523. https://doi.org/10.1177/0093650204267. Maas, C. J. M., and J. J. Hox. 2005. “Sufficient Sample Sizes for Multilevel Modeling.”Methodology 1, no. 3: 86–92. https://doi.org/10.1027/ 1614-2241.1.3.86. Manzey, D., and S. Gurk. 2005. Prozedurale Gerechtigkeit von Personalauswahlmaßnahmen: Untersuchungen zu einer Deutschen Version der Selection Procedural Justice Scale (SPJS) von Bauer et al. (2001). [Procedural Justice of Personnel Selection Methods: Investigations of a German version of the Selection Procedural Justice Scale (SPJS) by Bauer et al. (2001)] [Paper presentation]. Bonn, Germany: 4th Annual Conference of the German Society for Work and Organizational Psychology. Melchers, K. G., and J. M. Basch. 2022. “Fair Play? Sex‐,Age‐,and Job‐Related Correlates of Performance in a Computer‐Based Simulation Game.”International Journal of Selection and Assessment 30, no. 1: 48–61. https://doi.org/10.1111/ijsa.12337. Ohlms, M. L. 2024. “Gamifizierung und Game-based Assessment [Gamification and Game-Based Assessment].”In Digitale Personalauswahl und Eignungsdiagnostik, edited by U. P. Kanning and M. L. Ohlms, 127–154. Springer. https://doi.org/10.1007/978-3-662-68211-1. Ohlms, M. L., K. G. Melchers, and U. P. Kanning. 2024a. “Can We Playfully Measure Cognitive Ability? Construct‐Related Validity and Applicant Reactions.”International Journal of Selection and Assessment 32, no. 1: 91–107. https://doi.org/10.1111/ijsa.12450. Ohlms, M. L., K. G. Melchers, and U. P. Kanning. 2024b. “Playful Personnel Selection: The Use of Traditional vs. Game‐Related Personnel Selection Methods and Their Perception From the Recruiters' and Applicants' Perspectives.”International Journal of Selection and Assessment 32, no. 3: 381–398. https://doi.org/10.1111/ijsa.12466. Ohlms, M. L., K. G. Melchers, and F. Lievens. 2025. “It's Just a Game! Effects of Fantasy in a Storified Test on Applicant Reactions.”Applied Psychology: An International Review 74, no. 1: e12569. https://doi.org/ 10.1111/apps.12569. Ohlms, M. L., E. Voigtländer, K. G. Melchers, and U. P. Kanning. 2024. “Is Gamification a Suitable Means to Improve Applicant Reactions and Convey Information During an Online Test?”Journal of Personnel Psychology 23, no. 4: 169–178. https://doi.org/10.1027/1866-5888/ a000343. Oostrom, J. K., D. Van Der Linden, M. P. Born, and H. T. Van Der Molen. 2013. “New Technology in Personnel Selection: How Recruiter Characteristics Affect the Adoption of New Selection Technology.”Computers in Human Behavior 29, no. 6: 2404–2415. https://doi.org/10.1016/j.chb.2013.05.025. Qian, M., and K. R. Clark. 2016. “Game‐Based Learning and 21st Century Skills: A Review of Recent Research.”Computers in Human Behavior 63: 50–58. https://doi.org/10.1016/j.chb.2016.05.023. Ramos‐Villagrasa, P. J., E. Fernández‐Del‐Río, and Á. Castro. 2022. “Game‐Related Assessments for Personnel Selection: A Systematic Review.”Frontiers in Psychology 13: 952002. https://doi.org/10.3389/ fpsyg.2022.952002. Reed, K., D. H. Doty, and D. R. May. 2005. “The Impact of Aging on Self‐Efficacy and Computer Skill Acquisition.”Journal of Managerial Issues 17, no. 2: 212–228. Rollings, A., and E. Adams. 2003. Andrew Rollings and Ernest Adams on Game Design. Indianapolis, IN: New Riders. Werbach, K., D. Hunter, and W. Dixon. 2012. For the Win: How Game Thinking Can Revolutionize Your Business. Philadelphia, Pennsylvania: Wharton Digital Press. Wiechmann, D., and A. M. Ryan. 2003. “Reactions to Computerized Testing in Selection Contexts.”International Journal of Selection and Assessment 11, no. 2–3: 215–229. https://doi.org/10.1111/1468-2389. 00245. Wilde, M., K. Bätz, A. Kovaleva, and D. Urhahne. 2009. “Überprüfung Einer Kurzskala Intrinsischer Motivation [Validation of a Short Scale of Intrinsic Motivation].”Zeitschrift für Didaktik der Naturwissenschaften 15: 31–45. https://pub.uni-bielefeld.de/record/2404161. Williams, D., N. Yee, and S. E. Caplan. 2008. “Who Plays, How Much, and Why? Debunking the Stereotypical Gamer Profile.”Journal of Computer‐Mediated Communication 13, no. 4: 993–1018. https://doi. org/10.1111/j.1083-6101.2008.00428.x. 15 of 16
Woods, S. A., S. Ahmed, I. Nikolaou, A. C. Costa, and N. R. Anderson. 2020. “Personnel Selection in the Digital Age: A Review of Validity and Applicant Reactions, and Future Research Challenges.”European Journal of Work and Organizational Psychology 29, no. 1: 64–77. https://doi.org/10.1080/1359432X.2019.1681401. Supporting Information Additional supporting information can be found online in the Supporting Information section. 16 of 16 International Journal of Selection and Assessment, 2025
