scieee AI-readable full text Open interactive document viewer

Game on, Faking off? Are Game-Based Assessments Less Susceptible to Faking Than Traditional Assessments?

Melchers, Klaus G.,Kanning, Uwe P.,Barends, Ard J.,Ohlms, Marie L.

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Melchers, Klaus G.; Kanning, Uwe P.; Barends, Ard J.; Ohlms, Marie L. Article — Published Version Game on, Faking off? Are Game-Based Assessments Less Susceptible to Faking Than Traditional Assessments? Journal of Business and Psychology Suggested Citation: Melchers, Klaus G.; Kanning, Uwe P.; Barends, Ard J.; Ohlms, Marie L. (2025) : Game on, Faking off? Are Game-Based Assessments Less Susceptible to Faking Than Traditional Assessments?, Journal of Business and Psychology, ISSN 1573-353X, Springer US, New York, Vol. 40, Iss. 6, pp. 1323-1335, https://doi.org/10.1007/s10869-025-10019-6 This Version is available at: https://hdl.handle.net/10419/333366 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/ Vol.:(0123456789) Journal of Business and Psychology (2025) 40:1323–1335 https://doi.org/10.1007/s10869-025-10019-6 ORIGINAL PAPER Game on, Faking off? Are Game‑Based Assessments Less Susceptible toFaking Than Traditional Assessments? MarieL.Ohlms1,2 · KlausG.Melchers1 · UweP.Kanning3 · ArdJ.Barends4,5 Accepted: 1 April 2025 / Published online: 9 April 2025 © The Author(s) 2025 Abstract Game-based assessment (GBA) in personnel selection and assessment has gained increasing attention among researchers and practitioners in recent years. A postulated advantage of this newselectionmethod is its suggested suitability to reduce applicant faking. However, currently, it is unclear whether GBAs are indeed less susceptible to faking than traditional assessments. To address this question, we conducted an experimental study to examine whether a GBA measuring honesty-humility does indeed reduce faking compared to a traditional personality test. N = 171 participants were randomly assigned to an honest or an applicant condition and then completed a GBA and a traditional test that both measured honesty-humility. Results showed that test takers were able to distort their responses in the GBA and in the traditional honesty-humility test. However, the faking effect in the honesty-humility GBA was significantly smaller than in the traditional test. Thus, our findings suggest that using GBAs to measure personality can reduce faking compared to traditional tests to some degree but that GBAs are not a panacea to completely prevent it. Keywords Personnel selection· Game-based assessment· Faking· Gamification· Serious game· Honesty-humility· Personality Applicants in a selection process usually try to create a positive impression to be perceived as competent and trustworthy, aiming to enhance their chances of getting a job offer (Marcus, 2009). To achieve this goal, they may emphasize strengths, omit negative information, or even fabricate information about their skills or personality. When applicants consciously stretch the truth, lie about qualifications, or distort test scores during the selection process to meet job requirements, they are engaging in faking (Donovan etal., 2014; Levashina & Campion, 2007; Tett & Simonet, 2011). Faking during selection processes is a persistent concern among researchers and practitioners because it can affect applicants’ assessment performance (e.g., Birkeland etal., 2006; Hu & Connelly, 2021; Viswesvaran & Ones, 1999) and thus who is getting hired (Rosse etal., 1998). Relatedly, there are widespread concerns that faking may impair the This study was pre-registered at: https:// aspre dicted. org/ 6y4zqtg7. pdf. An earlier version of this manuscript was used as part of the first author’s doctoral dissertation (Ohlms, M. L. (2024). Digitization in personnel selection: Adaptation and enhancement of traditional personnel selection instruments using gamification and game-based assessment [Doctoral dissertation]. kiz Universität Ulm, Germany). Additional supplementary materials may be found here by searching on article title https:// osf. io/ colle ctions/ jbp/ disco ver. * Marie L. Ohlms [email protected]g.de Klaus G. Melchers [email protected] Uwe P. Kanning [email protected] Ard J. Barends [email protected] 1 Department ofPsychology andEducation, Section Work andOrganizational Psychology, Universität Ulm, Ulm, Germany 2 Department ofPsychology, Section Work andOrganizational Psychology, University ofFreiburg, Engelbergerstraße 41, 79085Freiburg, Germany 3 University ofApplied Sciences Osnabrück, Osnabrück, Germany 4 Leiden University, Leiden, TheNetherlands 5 Vrije Universiteit Amsterdam, Amsterdam, TheNetherlands 1324 Journal of Business and Psychology (2025) 40:1323–1335 psychometric properties of an assessment and thus the quality of selection decisions (Robie etal., 2021; Salgado, 2016; Tett etal., 2006), particularly for assessments that are rather easy to fake, such as personality tests (Viswesvaran & Ones, 1999). A relatively new selection method that is claimed to reduce faking is game-based assessment (GBA, Bhatia & Ryan, 2018). GBA involves the use of games (often webor internet-based) specifically designed for personnel selection to assess important job-related constructs (Armstrong etal., 2016). GBAs—together with gamified assessments—fall under the broader category of game-related assessments (Ohlms, 2024;Landers & Sanchez, 2022). However, while GBAs use various game elements and represent actual games, gamified assessment is a redesign strategy that only adds one or more game elements to gamify an otherwise traditional non-gamified assessment (e.g., see the gamified assessments by Landers & Collmus, 2022; or by Ohlms etal., 2025). Thus, a gamified assessment leaves the actual test content unchanged and only adds a limited number of game elements to the assessment. Accordingly, gamified assessments typically feel and function much less like a game than a GBA. For fakeable constructs such as personality traits, GBAs may be less susceptible to faking than traditional selection tests, as test takers may be less aware that they are being tested when they are immersed in a game, and may thus be less motivated to fake their responses (Bhatia & Ryan, 2018). Furthermore, within a GBA, it might be less transparent to applicants which constructs measured. Thus, faking might be more difficult as the most desirable response is less apparent to them which reduces the opportunity to fake (Landers & Sanchez, 2022). However, hardly any research has examined whether the use of GBAs can indeed reduce faking (Ramos-Villagrasa etal., 2022). To address this research gap, we examined whether a personality GBA is less prone to faking than a traditional test. Specifically, we conducted an experiment to compare participants’ honesty-humility scores from a GBA and a traditional test in an honest vs. an applicant condition. By doing so, we advanced the understanding of whether GBAs, compared to traditional tests, reduce faking susceptibility. Based on theoretical models of faking (e.g., Levashina & Campion, 2006; Tett & Simonet, 2011), we posit that motivational and situational factors affect the motivation and the opportunity to fake, and that these might contribute to reduced faking susceptibility of GBAs. For organizations, the present study highlights whether GBAs can be used to restrict applicants’ intentional response inflation. Theoretical Background andPrevious Research Research on faking in selection instruments has a long history, especially for personality inventories (e.g., Birkeland etal., 2006; Hu & Connelly, 2021; Viswesvaran & Ones, 1999) but also for other selection methods such as interviews (e.g., Melchers etal., 2020) or situational judgment tests (SJTs) (e.g., Peeters & Lievens, 2005). In the context of personnel selection, faking is defined as “an intentional response behavior aimed at exerting a positive influence on the hiring decision” (Griffith etal., 2011, p. 345). Faking research has shown that the intentional distortion of responses may be problematic from both a theoretical and an applied perspective, as it can inflate performance scores (e.g., Birkeland etal., 2006; Hu & Connelly, 2021; Viswesvaran & Ones, 1999), impact hiring decisions (Rosse etal., 1998), and negatively affect the construct- (e.g., Boss etal., 2015; Schmit & Ryan, 1993) and criterion-related validity of assessments (e.g., Donovan etal., 2014; Jeong etal., 2017; Peterson etal., 2011). Consequently, response distortion remains a concern for both researchers and practitioners who continue to seek effective mitigation strategies (e.g., Bill & Melchers, 2022; Dwight & Donovan, 2003; Martinez & Salgado, 2021). In the literature, it is suggested that GBA reduces applicant faking, thereby potentially facilitating a more accurate measurement of the targeted constructs (e.g., Landers & Sanchez, 2022; Woods etal., 2020). However, empirical research examining this suggestion is scarce. Therefore, we set up the present study to evaluate the extent to which this claim holds true. But to better understand the potential of personality GBAs to reduce faking, it is first necessary to consider the theoretical background of faking. Theoretical Models ofFaking Different theoretical faking models (e.g., Ellingson & McFarland, 2011; Levashina & Campion, 2006; Tett & Simonet, 2011) assume that three key factors affect faking behavior and its success: (1) applicants’ ability to fake, (2) their motivation to fake, and (3) the opportunity to fake in a specific situation. First, faking poses a cognitively challenging task, as one has to identify which response or behavior is considered the most socially desirable in a certain situation (Tett & Simonet, 2011). Accordingly, the ability to fake, which is usually included as an individual difference variable in the different faking models, is needed for effective faking. In line with this, individuals scoring higher on cognitive ability (Geiger etal., 2018; Levashina etal., 2009) and on the ability to identify criteria (Kleinmann etal., 2011) tend to be more successful in inflating their responses because they are better at discerning what an assessment intends to measure and then to adjust their responses in line with these targeted assessment dimensions. In line with these suggestions, Buehl etal. (2019) found that these abilities were indeed related to “better” faking (i.e., larger score changes) in simulated selection interviews. Second, the motivation to fake is influenced by applicants’ personality (e.g., low 1325Journal of Business and Psychology (2025) 40:1323–1335 conscientiousness, McFarland & Ryan, 2000, or high scores on competitive worldviews or dark personality traits, (Roulin & Krings, 2016) but can also be affected by situational aspects such as the organizational attractiveness (Buehl & Melchers, 2018). Finally, the opportunity to fake refers to the extent to which a certain situation enables faking or the extent to which the content of an assessment is fakeable. In the context of GBAs, both the motivation and the opportunity to successfully fake may be lowered. The rationale for this is, first, that while playing a GBA and being immersed in the game environment, applicants might become absorbed in gameplay and may be less aware that they are being tested (Bhatia & Ryan, 2018). Consequently, the motivation to portray themselves in a favorable light and to engage in faking behavior may be weaker. Additionally, in a GBA, the individual items can be seamlessly embedded in the game environment and applicants may have to react simultaneously to dynamically shifting scenarios within the game. Accordingly, it may be less evident which behaviors are being assessed in a GBA and thus what might be a socially desirable behavior in this situation. Consequently, the opportunity to successfully present oneself favorably in GBAs may be reduced if the constructs measured are less transparent to applicants (Woods etal., 2020). Furthermore, GBAs may impose a higher cognitive load compared to traditional assessments due to their interactive and dynamic nature. The increased cognitive load may further reduce the opportunity to fake, as it requires applicants to process multiple pieces of information simultaneously, making it more difficult to consciously adjust responses in a socially desirable manner. Review ofPrevious Research A large body of research on personality tests has shown that test takers can (Viswesvaran & Ones, 1999) and do (Birkeland etal., 2006; Hu & Connelly, 2021) positively distort their scores in personality inventories to enhance their chances of receiving a job offer. Specifically, Viswesvaran and Ones found that test takers instructed to fake good in lab settings can inflate their Big Five personality test scores compared to those who responded honestly and the corresponding effect sizes were moderate to large. Similarly, studies on HEXACO personality inventories, which supplements the Big Five by honesty-humility (Ashton & Lee, 2007), also found faking effects for the honesty-humility factor (MacCann, 2013). Furthermore, studies that compared scores from real applicants (high-stakes setting) with scores from non-applicants (low-stakes setting) or from another low-stakes setting also found substantial response inflation even though the mean differences were smaller than lab studies (Birkeland etal., 2006; Hu & Connelly, 2021). Notably, research shows that honesty-humility is particularly susceptible to response distortion compared to the other HEXACO personality factors (Anglim etal., 2017; Holtrop etal., 2021). Given that honesty-humility is a highly valid predictor of counterproductive work behavior (Lee etal., 2019; Pletzer etal., 2019), this might lead to concerns about negative effects of faking on the validity of applicants’ honesty-humility scores. In addition to personality inventories, there is evidence that people can also inflate their scores in other, non-selfreport selection instruments, but the faking effects are often smaller than in personality tests—even when these other instruments are also designed to measure the same personality constructs. Thus, test takers are able to inflate their scores in other methods such as SJTs (Kasten etal., 2020) and structured interviews (Van Iddekinge etal., 2005), but the faking-related mean differences were smaller than in the corresponding personality tests. In addition to this, recent research found that faking effects in interviews were stronger when these were less structured (Bill etal., 2024) which provides further evidence that the choice and the design of an assessment instrument can affect the opportunity to fake. Despite the postulated advantage of GBAs to reduce faking (e.g., Bhatia & Ryan, 2018; Woods etal., 2020), research in this domain remains scarce and has found heterogeneous results. In particular, Landers and Collmus (2022) gamified a traditional conscientiousness and openness test by embedding it into a storyline and compared its faking susceptibility with a traditional test. The gamified conscientiousness test showed decreased faking susceptibility compared to the non-gamified version, whereas the gamified openness test did not. In contrast, Barends etal. (2019) found that gamified personality cues (e.g., choosing one from seven cars varying in their level of luxury) intended to measure honesty-humility were as fakeable as a traditional selfreport honesty-humility inventory. However, both Landers and Collmus and Barends etal. used gamified assessments instead of fully-fledged GBAs. Thus, their assessments had both only used a single game element and likely came across much more like a test or a traditional assessment rather than a game. Although Landers and Collmus (2022) and Barends etal. (2019) provide initial insights into faking on gamified assessments, to the best of our knowledge, to date, only Barends and De Vries (2023) investigated faking in an actual GBA. Specifically, they examined the fakeability of a GBA measuring honesty-humility, using a fake good (Study 1) and a fake specific (Study 2) instruction, and compared scores in the faking conditions with those in a control condition. In both studies, Barends and De Vries found no differences in honesty-humility game scores between the fake good and control conditions, suggesting that test takers were unable to fake in the GBA. However, three limitations may potentially restrict the 1326 Journal of Business and Psychology (2025) 40:1323–1335 generalizability of these results. First, as noted in their article, the faking treatment had only a limited effect, as test takers in the faking and control condition equally indicated answering realistically (in contrast to making a good impression in the faking condition). Secondly, as participants in the control condition received no specific instruction (i.e., no honesty instruction), they may not have answered honestly, which might have led to inflated scores in the control condition. Thirdly, to answer the question of whether GBAs are less prone to faking than their traditional counterparts, a direct comparison of GBA scores in the faking and honest condition with those obtained in a traditional test would be necessary. This is particularly relevant as comparing effect sizes from different studies would assume that manipulations across different studies are equally strong; however, this may not be the case. Therefore, the current study aims to provide a clearer comparison of faking effects in different methods measuring the same constructs. Current Research As mentioned above, GBA is suggested as a countermeasure for faking. Hence, the primary goal of our study was to determine whether a GBA assessing honesty-humility does indeed reduce faking compared to a traditional honestyhumility test. To answer this question, we used a 2 (betweensubjects instruction: applicant vs. honest instruction) × 2 (within-subjects selection instrument: honesty-humility GBA vs. traditional honesty-humility test) mixed-subjects design. In our study, we focused on the personality trait honestyhumility because it is a valid predictor of all three major dimensions of job performance—task performance, organizational citizenship behavior, and counterproductive work behavior (Lee etal., 2019). Furthermore, for counterproductive work behavior, honesty-humility was not only the most valid personality predictor in several meta-analyses but even had incremental validity above and beyond the Big Five, cognitive ability, and integrity (Lee etal., 2019; Pletzer etal., 2019, 2020). However, as mentioned above, honesty-humility is particularly prone to faking compared to the other HEXACO personality traits (Anglim etal., 2017; Holtrop etal., 2021). Additionally, we also have to acknowledge that we only had access to a personality GBA that assesses honesty-humility but not the other HEXACO personality traits. Nonetheless, the results of the current study may identify an upper bound of faking effects of other personality GBAs. Building on this reasoning and on research on faking in personality tests reviewed above, we hypothesize that irrespective of the selection instrument (i.e., traditional test vs. GBA), assessments measuring honesty-humility are prone to response inflation. Hypothesis 1: Participants in the applicant condition will receive more favorable honesty-humility test scores both (a) in a traditional honesty-humility test and (b) in an honesty-humility GBA than participants in the honest condition.1 As explained above, the use of personality GBAs, in contrast to traditional personality inventories, might potentially limit the overall degree of faking. Specifically, when immersed in a GBA, test takers might feel less motivated to inflate their responses, as they may be less aware of being tested (Bhatia & Ryan, 2018). Furthermore, in the complex, dynamic game environment of a GBA, it may be less transparent to test takers what constitutes desirable behavior. Consequently, personality GBAs may be less susceptible to faking compared to traditional personality tests in which they can conveniently adapt their responses on the rating scale to portray themselves in a positive light. Accordingly, we aim to examine whether GBAs might indeed reduce faking susceptibility. Hypothesis 2: The difference in honesty-humility test scores between the application condition and the honest condition will be larger for a traditional test than for a GBA of honesty-humility. Method Sample We conducted an a-priori power analysis to determine the sample size for a power of 0.80 to test our hypotheses. The assumed effect sizes were based on previous faking research (Viswesvaran & Ones, 1999). This analysis revealed an N of 126 for a 2 × 2 (between-subjects: instruction [i.e., honest vs. applicant condition] × within-subjects: selection instrument [i.e., GBA vs. traditional test]) mixed analysis of variance (ANOVA) to detect an intermediate-sized within-subjects main effect and a small to intermediate-sized within-between interaction (i.e., f = 0.15 or d = 0.30), assuming a correlation among repeated measures of r = 0.30 based on previous 1 This study was pre-registrated at: https:// aspre dicted. org/ 6y4zqtg7. pdf.. For the sake of transparency, we would like to point out that in response to feedback received during the review process, we have made some changes to our pre-registered analysis plan. These changes have been made to improve clarity and to better align with the reviewers’ suggestions, while maintaining the integrity of our original hypotheses. 1327Journal of Business and Psychology (2025) 40:1323–1335 research (Viswesvaran & Ones, 1999). To ensure that our sample was large enough to test our hypotheses after the potential exclusion of participants who did not meet our inclusion criterion of being able to correctly answer a manipulation check, we aimed to collect data from 200 participants. Participants were recruited via direct contact. Two hundred university students from different disciplines and universities in Germany voluntarily participated in the study, with 100 participants each in the honest and the applicant condition. Due to failing the manipulation check (see below), we excluded 26 participants before data analyses. Furthermore, due to technical problems with the honesty-humility GBA server, the results of three participants were lost. Therefore, we excluded these three participants for all analyses. Thus, the total N was 171 (113 women, 57 men, and 1 diverse participant), with 89 participants in the honest condition and 82 in the applicant condition. Age ranged from 18 to 41 years (M = 22.89, SD = 3.18). The majority of the participants reported a German Abitur (the final secondary-school examination in Germany qualifying for university admission) as their highest educational degree (89.48%), whereas 7.60% already had a bachelor’s degree, and 2.92% had a master’s degree. Procedure andExperimental Conditions The study used a mixed 2 × 2 design with instruction (honest vs. applicant condition) as a between-subjects factor and selection instrument (a traditional test vs. GBA assessing honesty-humility) as a within-subjects factor. Data collection took place in university facilities. First, we obtained participants’ informed consent. Then, we randomly assigned them to either the honest or the applicant condition. Participants in the honest condition were told that the goal of the study was to validate two traditional tests and a computer game. They were asked to answer all items honestly and were assured that their answers would be kept confidential and be used for research purposes only. In contrast, participants in the applicant condition were instructed to imagine that they had applied for a trainee program at a fictitious organization and had now been invited to a recruitment event at the organization as part of the selection process. Specifically, they were instructed to act as highly interested and suitable applicants. Furthermore, to simulate a high-stakes setting and to enhance their motivation to put their best foot forward, participants in the applicant condition were told that the top three participants would each receive 50 Euros. After receiving either the honest or the applicant instruction, all participants completed a GBA and a paper–pencil test assessing honesty-humility. To control for possible order effects, the order in which participants completed these assessments was randomized. In the last part of the study, participants completed an online questionnaire assessing their demographics as well as the manipulation check. Measures Honesty‑Humility Game‑Based Assessment As the personality GBA, we used the Building Docks webbased assessment game designed to assess honesty‐humility (see Barends etal., 2022, for a detailed description of the GBA). Prior studies on this GBA (Barends & De Vries, 2023; Barends etal., 2022) found convergent correlations of about 0.33/0.34 between honesty-humility GBA scores and self-reported honest-humility scores as well as support for the discriminant validity of the GBA in comparison to self-reports of the other five HEXACO traits and cognitive ability. Furthermore, the GBA was also a valid predictor of cheating for financial gain (Barends etal., 2022). In the Building Docks GBA, test takers take on the role of a port owner with the objective of collaboratively building a port together with three computer-controlled characters and turning it profitable. The behaviors of the computer-controlled characters were fixed to make the game as standardized as possible. The GBA is played as a turn-based game where events (i.e., items referring to a specific subtask) are presented in a fixed order to the test takers and the game continues until all items are completed even though test takers generally turn the port profitable about halfway throughout the game (the exact timing is dependent on the behavior of the player). This GBA contains three subtasks: economic games, virtual cues, and SJT items. Most items of the subtasks were created to measure honesty-humility, but the subtasks also included several filler items (targeting other HEXACO traits, however, too few to reliably measure these traits; see Barends etal., 2022). The main game element is the economic games, where money is transferred between characters and cooperative behavior is associated with a higher honesty-humility score. The items are scenario-based adaptations of wellknown economic games such as the dictator game and the public goods game (Van Lange etal., 2014). Furthermore, virtual cues (Barends etal., 2019) are used as rewards during the game, allowing players to personalize their character and the port. Here, the choice of less luxurious items (e.g., a basic automobile) is coded as higher honesty-humility than the choice of more luxurious items (e.g., a sports car). Moreover, the SJT items aim to create a separate storyline within the GBA, and the answer options represent different degrees of honesty-humility (e.g., whether or not the player makes a false promise to obtain a permit for an event in the port).2 Internal 2 We used Version 1.1 of the GBA that decreased the length and complexity of several instructions and items compared to the version used in the original publications. However, the meaning of item content was not changed except for one SJT item that was completely rewritten as it was considered too subtle in the original version of the GBA. 1328 Journal of Business and Psychology (2025) 40:1323–1335 consistency for the overall honesty-humility score based on 39 items was α = 0.85. The internal consistencies for the three subtasks were 0.59 for the SJT items, 0.85 for the virtual cues, and 0.67 for the economic games. Coefficient alpha for the SJT items was comparable to other SJTs assessing honestyhumility (e.g., α = 0.50 by Oostrom etal., 2019). Traditional Personality Test As the traditional test measuring honesty-humility, we used the 16 honesty-humility items from the German version of the HEXACO100 (Lee & Ashton, 2018). Items (e.g., “I would never accept a bribe, even if it were very large”) were answered on 5-point Likert scales from 1 = strongly disagree to 5 = strongly agree (α = 0.88). Manipulation Checks We included two manipulation checks. First, at the end of the study, participants had to answer the question “Which of the following answer options best describes the instructions you received at the beginning of the study?” by choosing one of the four options: “answer honestly,” “answer as if I was the most suitable candidate,” “there was no specific instruction,” or “I don’t remember the instruction.” As noted above, participants who failed to answer this manipulation check correctly were excluded from further analyses (8 participants in the honest condition and 18 in the applicant condition). And as a second manipulation check to assess the effectiveness of our manipulation in inducing higher faking in the applicant condition compared to the honest condition, we used five items from the faking attempt scale developed by Barends etal. (2019). These items used a 5-point Likert scale ranging from 1 = strongly disagree to 5 = strongly agree (α = 0.86, see Table3 from the Appendix for these and the following items). The items were translated into German and checked with back-translation. Video Game Experience Participants’ video game experience was measured using five items (α = 0.94) from Bourgonjon etal. (2010) in the German translation from Ohlms etal. (2024a). Table 1 Descriptive information and correlations among study variables Note. N = 89 (honest condition) and 82 (applicant condition). Intercorrelations for the honest condition are presented below the diagonal, and intercorrelations for the applicant condition are presented above the diagonal. Mean scores for performance in the honesty-humility GBA and honesty-humility test are presented as z-scores. Gender is coded as 0 = male, 1 = female GBA =game-based assessment, trad. =traditional, SJT =situational judgment test, VC =virtual cues, EG= economic games a n = 81 (the diverse participant was dropped for all correlations involving gender) *p < 0.05; **p < 0.01; ***p < 0.001 Variable M SD 1 2 3 4 5 6 7 8 M22.78 0.74 1.94 0.66 0.45 0.19 0.48 0.22 SD 2.73 0.44 1.04 0.67 0.93 1.04 0.91 0.95 1. Age 22.99 3.56 − 0.05 0.33** 0.03 0.18 0.05 0.14 0.20 2. Gendera0.60 0.49 0.14 − 0.54*** 0.11 0.06 0.05 0.12 − 0.11 3. Video game experience 2.32 1.26 − 0.09 − 0.61*** 0.06 0.08 0.04 − 0.05 0.28* 4. Honesty-humility trad. test overall − 0.60 0.86 0.15 0.10 − 0.02 0.33** 0.30** 0.24* 0.25* 5. Honesty-humility GBA overall − 0.41 0.88 − 0.05 − 0.06 − 0.00 0.43*** 0.64*** 0.88*** 0.67*** 6. Honesty-humility GBA: SJT − 0.17 0.93 − 0.10 0.01 0.01 0.48*** 0.74*** 0.36*** 0.32** 7. Honesty-humility GBA: VC − 0.45 0.87 − 0.08 − 0.00 − 0.07 0.36*** 0.85*** 0.49*** 0.34** 8. Honesty-humility GBA: EG − 0.21 1.01 0.08 − 0.15 0.09 0.15 0.60*** 0.37*** 0.15 Table 2 Means, standard deviations, and effect sizes for the comparison of scores in the selection instruments in the honest vs. the applicant condition Note. nhonest = 89; napplicant = 82. Values are presented as z-scores per instrument to allow comparison of scores across instruments a Values for Cohen’s d refer to the difference in independent sample t-tests between the honest and applicant conditions; positive values for Cohen’s d indicate higher means in the applicant condition **p < 0.01; ***p < 0.001 Selection instrument Honest Applicant Cohen’s da M (SD)M (SD) Traditional testoverall score − 0.60 (0.86) 0.66 (0.67) 1.62*** GBA overall score − 0.41 (0.88) 0.45 (0.93) 0.95** GBA:SJT items − 0.17 (0.93) 0.19 (1.04) 0.37* GBA:virtual cues − 0.45 (0.87) 0.48 (0.91) 1.05*** GBA:economic games − 0.21 (1.01) 0.22 (0.95) 0.44** 1329Journal of Business and Psychology (2025) 40:1323–1335 Results Means, standard deviations, intercorrelations, and reliabilities for the study variables are shown in Table1. Information for the honest condition is presented below the diagonal and for the applicant condition above the diagonal. Mean test scores for the two assessment instruments are shown as z-scores per instrument to allow a comparison of scores across instruments. Compared to previous studies by Barends and De Vries (2023) and Barends etal. (2022), we found similar correlations between the honesty-humility GBA and the traditional honesty-humility test. The correlation between the two assessment instruments was also comparable to Oostrom etal. (2019), who developed an honesty-humility SJT and validated it against selfreported honesty-humility. In addition, our results are also in line with meta-analytic results that found convergent validities ranging from 0.31 to 0.56 between different traditional personality self-reports measuring the same personality trait (Pace & Brannick, 2010). When compared to the meta-analytic results of the conceptually closest Big Five trait, agreeableness, this average correlation between self-reports was r = 0.31. Moreover, the correlation between self-reports and GBA measures of the same trait is also expected to be somewhat lower as there should be less common method bias between a GBA and a self-report in comparison to different selfreport instruments as there are no shared response styles, common scale formats, or identical scale anchors (Podsakoff etal., 2024) between these two methods. Furthermore, the correlation between the GBA score and the traditional test score of honesty-humility was descriptively slightly weaker in the applicant condition (r = 0.33, p < 0.01) compared to the honest condition (r = 0.43, p < 0.001). However, a one-tailed Steiger’s (1980) test for independent groups showed that the correlations were not significantly different, z = 0.75, p = 0.45. In a preliminary analysis, we assessed potential differences between participants in the honest and applicant conditions regarding age and video game experience using t-tests for independent groups and regarding gender using a χ2-test. Participants in both conditions were comparable in terms of age, t(169) = 0.43, p = 0.67, d = 0.07, and gender, χ2(2) = 5.10, p = 0.08. However, groups differed regarding their video game experience, t(166.81) = 2.15, p = 0.03, d = 0.33 (equality of variances could not be assumed). Therefore, we included video game experience as a covariate for subsequent analyses. To test the effectiveness of our applicant instruction in promoting faking attempts in the applicant condition relative to the honest condition, we conducted an independent samples Welch’s t-test (equality of variances could not be assumed) to compare scores on the faking attempt scale. This test confirmed that our manipulation was successful so that the faking attempts were significantly higher in the applicant condition (M = 4.58, SD = 0.49) than in the honest condition (M = 2.92, SD = 0.88), t(124.20) = 15.07, p < 0.001, d = 2.36. Tests ofHypotheses To assess whether participants in the applicant condition received more favorable honesty-humility scores for the traditional test and the GBA than in the honest condition, we conducted a 2 (instruction: honest vs. applicant condition) × 2 (selection instrument: traditional honestyhumility test vs. honesty-humility GBA) mixed ANCOVA with test scores on the different selection instruments as dependent variables and video game experience as a covariate. In line with Hypothesis 1, there was a large significant main effect with participants in the applicant condition scoring higher than participants in the honest condition, F(1, 168) = 96.01, p < 0.001, η2p = 0.36 (see Table2 and Fig.1). In contrast, there was no main effect of the selection instrument, F(1, 168) = 0.09, p = 0.76, η2p = 0.001. Additionally, we examined Hypothesis 2, which predicted a larger difference between participants’ honestyhumility scores in the applicant vs. honest condition for the traditional test than for the GBA. To do so, we considered the instruction × selection instrument interaction of the 2 × 2 ANCOVA, which also turned out to be significant, F(1, 168) = 7.07, p = 0.009, η2p = 0.04. This result reflected a smaller faking effect in the GBA (d = 0.95, p < 0.001) than in the traditional test (d = 1.62, p < 0.001).3 Furthermore, we examined whether participants in the honest and applicant condition differed in their scores on the three types of tasks included in the honesty-humility GBA: SJTs, virtual cues, and economic games (these analyses were not pre-registered). To do so, we conducted a one-way MANCOVA with the instruction (honest vs. applicant) as the independent variable, the task types as dependent variables, and video game experience as covariate. Results revealed a significant multivariate effect, Wilks’ λ = 0.78, F(3, 166) 3 We also repeated the analyses without including video game experience as a covariate. This did not change the results meaningfully; main effect instruction: F(1, 169) = 98.05, p <.001, η2p =.37; main effect selection instrument: F(1, 169) = 0.01, p =.91, η2p <.001; interaction effect: F(1, 169) = 7.56, p =.007, η2p =.04. 1330 Journal of Business and Psychology (2025) 40:1323–1335 = 15.50, p < 0.001, η2p = 0.22. Furthermore, separate repeated measures ANCOVAs also revealed significant effects for the SJT items, F(1, 168) = 5.83, p = 0.02, d = 0.37, virtual cues, F(1, 168) = 43.78, p < 0.001, d = 1.05, and economic games, F(1, 168) = 10.37, p = 0.002, d = 0.44.4 Thus, the virtual cues yielded a larger faking effect than the other two task types.5 Discussion We aimed to examine whether GBAs, a relatively new method in personnel selection and assessment, reduce faking as postulated in the literature (Bhatia & Ryan, 2018; Woods etal., 2020). Our results revealed two important things. First, we found that participants were able to distort their honesty-humility scores both in a self-report paper–pencil honesty-humility test as well as in a GBA assessing honestyhumility. Second, however, and in line with the claims in the literature, we found that the honesty-humility GBA was less susceptible to faking than a traditional honesty-humility test. Theoretical andPractical Implications Our study advances the understanding of faking in GBAs in several aspects. First, in line with previous meta-analytic results (Birkeland etal., 2006; Hu & Connelly, 2021; Fig. 1 Results of two-way mixed ANCOVA for honestyhumility scores as dependent variables, selection instrument as a within-subjects factor, and experimental condition as a between-subjects factor Note. nhonest = 89; napplicant= 82. Values are presented as z-scores per instrument to allow comparison of scores across instruments. ** p < .01. *** p < .001. 4 Again, an analysis without including video game experience as a covariate did not change the results qualitatively for the multivariate test or the subsequent ANOVAs. 5 Additionally, we checked whether the overall GBA score covered all honesty-humility facets equally as these may differ in their fakeability. We therefore reanalyzed the Barends etal. (2022) Study 2 data that used the full-length honesty-humility scale with the GBA. The comparisons were made using the procedure developed by De Vries etal. (2020) to check for masking and cancelation effects. The reports showed there were no significant differences between the correlations of the facets and overall honesty-humility to the GBA after correcting for multiple comparisons (see Barends etal., 2024, and also see TableS1 in the online supplement).