scieee AI-readable full text Open interactive document viewer

A checklist to assess the quality of survey studies in psychology

Protogerou, Cleo,Hagger, Martin S.

Full text

This is a self-archived version of an original article. This version may differ from the original in pagination and typographic details. Author(s): Title: Year: Version: Copyright: Rights: Rights url: Please cite the original version: CC BY 4.0 https://creativecommons.org/licenses/by/4.0/ A checklist to assess the quality of survey studies in psychology © 2020 the Author(s) Published version Protogerou, Cleo; Hagger, Martin S. Protogerou, C., & Hagger, M. S. (2020). A checklist to assess the quality of survey studies in psychology. Methods in Psychology, 3, 100031. https://doi.org/10.1016/j.metip.2020.100031 2020 A checklist to assess the quality of survey studies in psychology Cleo Protogerou a , b , * , Martin S. Hagger a , c a University of California, Merced, USA b University of Cape Town, South Africa c University of Jyv€ askyl€ a, Finland ARTICLE INFO Keywords: Study quality assessment Psychology Correlational studies Survey studies Evidence syntheses Transparency ABSTRACT Study quality is emerging as an essential component of evidence syntheses. It allows practitioners and policymakers to make informed decisions based on the quality of the evidence reviewed. Study quality is typically assessed by checklists of pre-determined quality criteria. Few study quality checklists have been systematically evaluated, and none have been developed specifically for survey studies in psychology. The present study addresses this evidence gap by developing the quality of survey studies in psychology (Q-SSP) checklistusing an expert-consensus method. An international panel of experts in psychology research and quality assessment (N¼53) evaluated the inclusion and importance of candidate quality items and offered commentary. The resulting checklist was used to evaluate a set of survey studies and inter-rater reliability of checklist scores was computed. A preliminary test of criterion validity of checklist scores was conducted using on a sample of survey studies with ‘known differences’in study quality verified by experts. Experts exhibited high agreement on inclusion and importance ratings of the candidate items. Minor adjustments were made to the candidate items based on experts' feedback. Inter-rater reliability of study quality scores using the checklist was high. Some evidence for criterion validity of scores using the checklist was obtained. Overall, we provide preliminary data to support the Q-SSP checklist as a potential means to evaluate the quality of survey studies in psychology. We recommend a future large-scale study using the Q-SSP checklist to assess study quality in studies with known differences in quality verified by experts. As research evidence for psychological phenomena accumulates, scientific communities, stakeholders, and policymakers are becoming increasingly reliant on research syntheses, such systematic reviews and meta-analyses, to provide pithy summaries of effects of interest, and to inform evidence-based policy and practice. While innovation in methods and analytic techniques for evidence syntheses provides increasingly sophisticated means to summarize research and test effects of interest, these methods are highly dependent on the quality of the evidence included in the analyses. Ways to evaluate the quality of research evidence for evidence syntheses are therefore increasingly recognized as essential components of evidence syntheses (Greenhalgh and Brown, 2017;Higgins and Altman, 2008;Lipsey and Wilson, 2001). Coupled with the imperative of assessing study quality for syntheses of research, there is also an increased need to evaluate the quality of individual studies. Study quality assessment can facilitate comparisons across individual studies and optimize the quality of future studies and their replication. Study quality is typically assessed using checklists, in which trained reviewers assess studies on a set of pre-determined quality components. Numerous study quality checklists or ‘tools’exist (e.g., Higgins et al., 2011;Jadad et al., 1996;Oxman and Guyatt, 1988). However, to date, no tool has been developed for the expressed purpose of assessing the quality of psychological research, and researchers in psychology have fulfilled the need for study quality assessment by adapting existing quality measures originally developed in other disciplines (e.g., Hagger et al., 2017;Husebøet al., 2012;Protogerou et al., 2018). As these tools have not been specifically developed to evaluate psychological studies, they may lack validity and provide insufficient coverage of the appropriate study quality components. The purpose of the present study is to fill this evidence gap by developing a study quality tool for psychological studies using survey designs. We focus on survey research as it is one of the predominant methods of research in psychology (Singleton and Straits, 2009). Furthermore, studies adopting survey methods are frequently the subject * Corresponding author. Psychological Sciences and Health Sciences Research Institute (HSRI), University of California, Merced, USA. E-mail address: cprotogero[email protected] (C. Protogerou). Contents lists available at ScienceDirect Methods in Psychology journal homepage: www.journals.elsevier.com/methods-in-psychology https://doi.org/10.1016/j.metip.2020.100031 Received 25 December 2019; Received in revised form 6 June 2020; Accepted 14 July 2020 Available online 17 July 2020 2590-2601/©2020 The Author(s). Published by Elsevier Ltd. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/). Methods in Psychology 3 (2020) 100031 of research syntheses in psychology (Hunter and Schmidt, 2004). Specifically, the present study aimed to develop tool for researchers to assess psychology survey studies using an expert consensus approach. The primary focus of the study was to establish the Q-SSP checklist as a means to provide assessments of study quality with evidence for the face and content validity of its items, as well as inter-rater reliability, and a secondary focus was to provide preliminary evaluation of the criterion validity of Q-SSP checklist scores with a goal of establishing whether the tool was effective in differentiating between studies of acceptable (or higher) versus questionable (or lower) quality. We expect the tool to improve the precision of research syntheses by enabling researchers to incorporate assessment of study quality as a key component of the sample of studies under review, and test effects of study quality on findings of the syntheses. We also anticipate the tool will inform the development of higher quality survey studies and research syntheses, as well as replications, by highlighting deficiencies in currently-available studies. Study quality: definitions and assessment Study quality reflects the extent to which a study has taken appropriate measures to minimize bias and error from inception to reporting of findings (Khan et al., 2011). It has been estimated that only approximately 20% of published studies across fields of behaviour health research are of sufficient quality (Ciliska and Buffet, 2008). Assessment of study quality 1 –also known as critical appraisal –is the systematic evaluation of the degree to which a study has been conducted to the highest possible quality standards (Higgins and Green, 2008). A study of acceptable quality provides assurances that the research was conducted in line with a set of pre-specified discipline-appropriate standards, and that findings may be legitimately generalized to populations of interest and implemented in practice. Therefore, research that has been assessed as having ‘good’quality based on a formal appraisal against specified quality standards, may allow researchers, clinicians, policy makers, and other interested stakeholders to make informed decisions based on the available evidence (Oxman and Guyatt, 1991). Other than providing a means to identify the strengths and weaknesses of a body of evidence, assessment of study quality also entails a number of other outcomes such as: selection of studies for inclusion in evidence syntheses; identifying potential sources of bias in the results of evidence syntheses; and gauging the impact of study quality on the results of a meta-analysis by incorporating study quality in subgroup and sensitivity analyses. Finally, quality assessment can also assist in improving research and publication standards by highlighting common deficiencies in the available evidence and possible means to improve the quality of subsequent studies (Greenhalgh and Brown, 2017;Greenhalgh, 2014;Johnson et al., 2014). Numerous checklists or ‘tools’designed to assess the quality of research study have been developed. Although there is idiosyncratic variability in content across the available tools, there is a degree of commonality in the general categories of quality criteria adopted. Typical categories of quality components relate to the population under investigation (e.g., sampling and recruiting strategies, sample size); study design (e.g., ‘appropriateness’of methodology, ethical review procedures); data collection (e.g., validation of instrument/measures used, detailed descriptions of data collection process); data analyses (e.g., ‘appropriateness’of statistical tests employed, dealing with attrition); and reporting and interpretation of results (e.g., completeness of results reported, suggestions for further research and practice) (see Crowe and Sheppard, 2011;Durant, 1994;Glynn, 2006 for examples of quality criteria used). Reviews of the literature have identified nearly 200 tools used to assess study quality across health and social sciences research (Deeks et al., 2003;Katrak et al., 2004). Extant tools have been developed to appraise experimental studies (e.g., Jadad et al., 1996), systematic reviews and meta-analyses (e.g., Oxman and Guyatt, 1988,1991), and qualitative studies (e.g., Zhang et al., 2019). Generic quality appraisal tools also exist (e.g., Glynn, 2006;National Institutes of Health, 2014). It has been argued, however, that most quality assessment tools have not been developed with sufficient scientific rigor (Crowe and Sheppard, 2011;Johnson et al., 2014;Katrak et al., 2004;Khan et al., 2011;Moyer and Finney, 2005). A long-standing argument against the standing of extant critical appraisal tools that has still yet to be resolved is that the tools omit key quality domains (Crowe and Sheppard, 2011;Deeks et al., 2003), and that no tool can be recommended without reservation (Alderson et al., 2003). The most prominent criticisms of extant tools relate to the absence of validity and reliability checks in their development. For example, in their review of 44 published quality appraisal tools, Crowe and Sheppard (2011) found that only six tools had been tested for concurrent validity, only two for construct validity, and only 12 for reliability. A further 11 tools had not been tested for any type of validity or reliability, and 17 tools provided no explanation on how they were developed and did not include details on how they should be administered and scored. Recommendations for developing credible quality assessment tools have highlighted the need to systematically identify relevant domains of study quality, include appropriate validity and reliability checks, account for discipline-specific research principles, and provide a guide with precise explanations of the terms used and scoring strategies (Crowe and Sheppard, 2011;Moyer and Finney, 2005). In addition, we note the absence of quality assessment tools designed specifically for survey research in psychology. 2 As survey research is one of the most frequently-used methods in psychology (Ponto, 2015; Singleton and Straits, 2009), a dedicated, fit-for-purpose quality tool is needed (Protogerou and Hagger, 2019). The lack of a tool to evaluate survey research has been noted by prominent methodologists in the field (Faragher et al., 2005;Hoffmann et al., 2017). Given the absence of relevant tools, researchers have adapted tools from disciplines outside psychology, or developed bespoke tools, in order to evaluate study quality (e.g., Faragher et al., 2005;Hagger et al., 2017;Hoffmann et al., 2017;Young et al., 2014). One problem with these adapted tools is that they reflect a ‘retrofitting’of tool content designed to assess the quality of other types of research and in other fields (e.g., medicine, health sciences), and, as a consequence, they are often lacking in some way. Such adapted tools may omit essential criteria relevant to psychology or survey methods, or the criteria are not specified in such a way that they relevant to, or sufficiently tailored to, the particular discipline. For example, many research quality checklists include items relating to participant recruitment or sampling methods, but this is often in the context of randomized-controlled or cross-sectional designs in settings like medical research, few make explicit reference to the specific information necessary to judge the quality of the recruitment/sampling methods in psychology survey studies (e.g., rates of participants refusing an initial invitation to participate, rates of attrition from survey completion). This makes the development of a psychology discipline-specific tool that satisfies the specific criteria for studies adopting survey methods essential for comprehensive coverage of study quality assessment in this domain. Development of disciplineand method-specific quality assessment tools is also important to ensure consistency in the criteria content and ratings of studies. The absence of a disciplineand method-specific tool that provides valid and reliable scores on study quality means researchers 1 Study quality should be differentiated from risk of bias, a related concept which reflects the extent to which systematic error in research will lead researchers to draw incorrect conclusions from the findings (Higgins and Altman, 2008). 2 Our definition of survey research is based on that provided by Check and Schutt (2012):“the collection of information from a sample of individuals through their responses to questions”(p. 160). C. Protogerou, M.S. Hagger Methods in Psychology 3 (2020) 100031 2 have fallen back on the use of multiple, diverse tools with idiosyncratic content to assess study quality. This presents a considerable challenge to researchers attempting to assess study quality across multiple studies, such as in the context of systematic reviews and meta-analyses, overviews of psychology survey research. The application of diverse tools to the same body of evidence may result in researchers arriving at different conclusions on the quality of the evidence, which can have ramifications for subsequent interpretations of the evidence. For example, variation in quality assessment scores may influence effect sizes across moderator groups defined by methodological quality scores in meta-analyses and affect conclusions drawn (Protogerou and Hagger, 2019). The imperative of precisely and reliably distinguishing between studies of acceptable and questionable quality in survey studies in psychology and the deficiencies of retrofitted tools highlights the need for a purpose-developed study quality tool. Study overview Recognizing the challenges presented by the lack of a dedicated tool to assess the quality of survey studies in psychology, we aimed to develop a tool to assess the quality of survey studies in psychology. The tool, the quality of survey studies in psychology (Q-SSP) checklist, was based on a comprehensive review of previous methodological quality assessment tools and checklists followed by an expert consensus method to evaluate and refine its content. The purpose of expert consensus methods is to define levels of agreement in a wide variety of settings, especially when insufficient or conflicting evidence exists (Fink et al., 1991;Jones and Hunter, 1995). Expert agreement pools the collective expertise of those with in-depth knowledge and training applied to the subject of interest (Hasson et al., 2000;Michie et al., 2017). Although it is acknowledged that variation and disagreements will occur in expert ratings, the consensus approach provides a summary of the convergence of expert knowledge. Furthermore, consensus approaches capitalize on the accumulated knowledge and practical experience of experts to obtain information that is culturally apt and rapidly implemented (Minas and Jorm, 2010;Stephens et al., 2017). For example, expert consensus studies have been used broadly across many disciplines to develop content of instruments and measures based on the pooled expertise in research (Herdman et al., 2002;Michie et al., 2005,2013;Stephens et al., 2017; Velligan et al., 2010), including the development of quality appraisal tools (Burnett et al., 2005;Jadad et al., 1996;Pace et al., 2012). Expert selection is a pertinent issue when it comes to using expert consensus to judge the validity of the content of measures and tools. While there is no established definition of an ‘expert’in a particular field or discipline, or rule as to who should be included as an expert in an expert consensus panel, some published guidelines exist. Broadly, it is recommended that experts are pooled from “relevant, backgrounds, and experiences”,“pertinent specialties”, and, when appropriate, members of relevant advocate groups and general public (Fink et al., 1984;Hsu and Sandford, 2007). Other recommendations suggest identifying experts through their involvement in relevant research and authorship of relevant publications (Addington et al., 2013;Yap et al., 2014). In addition to the criterion of ‘relevance’of experts' experience to the phenomenon under investigation, the ‘diversity’of experts has also been proposed as important; extant consensus studies have aimed to include experts from diverse geographical locations, professional ranks, genders, and age groups (Jorm, 2015). A common theme in the literature is that expertise and selected experts should be guided by the questions, aims, and needs of the consensus study in question (Jorm, 2015). Our definition of ‘expertise’is consistent with these extant practices in the literature using expert consensus methods. The Q-SSP checklist was developed in four stages. First, we developed an initial set of candidate quality items for the checklist based on a review of existing study quality appraisal tools and recommendations of consensus statements on quality requirements in psychology (Appelbaum et al., 2018;Asendorpf et al., 2013;Finkel et al., 2017). Based on this review, we developed clear language descriptions and assessment criteria for each item. Second, we used an expert consensus method to provide external, independent evaluations of the candidate set of quality items. A panel of experienced researchers with expertise in survey research, evidence synthesis, and quality appraisal evaluated the initial item set, descriptions, and assessment criteria in terms of their necessity, appropriateness, and importance for inclusion. Third, the tool was refined based on the results of the expert consensus ratings and open-ended comments to produce a final prototype of the tool. This version was used by the two authors to evaluate a sample of survey studies from three meta-analyses (Hagger et al., 2017;Hoffmann et al., 2017;Young et al., 2014). In the fourth and final stage, we aimed to provide preliminary support for the criterion validity of scores produced by the Q-SSP checklist using a ‘known differences’approach. We evaluated whether the tool could be used to effectively distinguish between groups of studies known to be of “acceptable”and “questionable”quality. First, a set of 20 candidate studies was identified from the aforementioned meta-analyses based on scores on the bespoke quality assessment tools used in the meta-analyses from which they were drawn. Next, a second expert panel provided independent appraisal of the studies and rated them as “acceptable”and “questionable”in quality based on their expert judgment. Subsequently, a final expert panel used the Q-SSP checklist to assess the quality of the same set of 20 studies. Ratings of the studies using the Q-SSP checklist the final panel were compared to the expert judgment ratings and scores from the previously used tools for the set of studies. Method Participants Participants comprised research-active faculty members and research staff from psychology, behavioral science, and health science faculties, with expertise in the application of psychological methods, survey research, study quality appraisal, and evidence synthesis. Participants were primarily identified through their publications. Specifically, literature searches were conducted with Google Scholar search engine, using the terms “questionnaire”,“survey”,“correlation”,“psychology”,“social” “behavior*“, methodological quality”,“study quality”,“meta-analysis”, “review”, and “evidence synthesis”. Relevant authors of retrieved studies were entered as candidates on our initial list of experts. Relevant expertise of candidates was based on their overall scholarly profile and experience including, but not limited to, number and breadth of research articles in high impact peer-reviewed discipline-relevant journals, citation ratings, rigor of publication design, previous experience with study quality assessment using methodological, study quality, or risk-of-bias tools or checklists. As we aimed to include a diverse panel of experts, and we did not restrict our search to country or language. In order to ensure international coverage, we conducted a further search of psychology/social/health science departments in the countries that did not feature in our search to identify relevant experts, through lists of publications appearing on the departmental websites. Recommendations on the optimal sample size in expert consensus vary. It has been suggested, for example, that 10 to 15 experts are sufficient if there is sufficient homogeneity in the group (e.g., members have similar professional background, education, training, and expertise; Delbecq et al., 1986). Overviews of consensus studies indicate that most studies employ between 15 and 20 experts (Hsu and Sandford, 2007). Based on these guidelines, we aimed for a final sample of at least 20 experts. Given that response rates to online surveys are approximately 33% (Nulty, 2008), we aimed to contact at least 100 eligible University faculty members and researchers with the requisite expertise and diversity in country and discipline coverage. Procedure The Q-SSP checklist was developed in four stages. An overview of the C. Protogerou, M.S. Hagger Methods in Psychology 3 (2020) 100031 3 aims and outcomes of each stage are provided in Table 1. Stage 1 –Development of Candidate Items. In the first stage of development of the Q-SSP checklist, we identified a set of candidate items for the initial version of the checklist based on previous research and recommendations (Asendorpf et al., 2013;Crowe and Sheppard, 2011;Durant, 1994;Moyer and Finney, 2005;Zeng et al., 2015). First, we searched the literature for existing study quality appraisal tools, and for overviews, systematic reviews, and meta-analyses, evaluating these tools. Second, we studied the content of the items in each tool and identified domains of study quality. Third, we studied reviews that appraised the rigor of existing quality assessment tools and took into account their recommendations for quality appraisal tool development. We also considered general published psychological research and publication standards (Appelbaum et al., 2018), and other recommendations for enhancing quality in psychological research (Asendorpf et al., 2013; Finkel et al., 2017). The procedure was conducted by the two authors who have extensive experience in evidence synthesis and quality assessment in the fields of social psychology, health psychology, and behavioral medicine. Stage 2 –Refining Item Pool Using Expert Consensus. The expert consensus study adopted a cross-sectional design comprising an online questionnaire administered using the Qualtrics TM online survey platform. The online expert consensus approach has been adopted in previous consensus studies (e.g., Connell et al., 2018), and has several advantages, including efficient, cost-effective means to recruit an appropriate panel of experts; assurance of participant anonymity; promotion of efficient dialogue with participants and means to prompt responses; and streamlined data collection, collation, and analysis (Wright, 2005). We were guided by Waggoner et al.’s (2016) best practice guidelines for online consensus research, which specify that studies should state inclusion criteria, recruit a minimum of 11 experts, provide a predetermined definition of consensus, and conduct comprehensive analysis of consensus data. It should be acknowledged that unanimous agreement in consensus surveys is rare and not expected. Previous research have adopted different criterion values for acceptable agreement in consensus studies (range 51%– 80%) (Hasson et al., 2000;Keeney et al., 2006). In the present study we adopted a conservative 80% criterion for acceptable agreement. In January 2018, eligible participants (N¼167) were sent an email invitation to participate in the consensus survey. The invitation provided full information about the study, expectations for participation, and a URL directing them to a welcome page that contained information about the study followed by a consent statement. Participants agreeing with the consent statement were automatically directed to the first page of the survey. Two more email reminders were sent in February 2018. Of the 167 experts contacted, 40 agreed to participate (24% response rate) and 33 completed the whole survey (17% attrition rate). Consequently, we exceeded the recommended number of experts (Waggoner et al., 2016). None of the seven ‘non-completers’proceeded beyond the demographic questions, so they were excluded from the analysis. Participant characteristics are presented in Table 2. Data collection was completed by the Table 1 Purpose and outcome/findings of each stage of development of the Q-SSP checklist. Stage Purpose Outcome/findings 1 Development of Q-SSP checklist items and scoring scheme by study authors. Shortlist of 20 candidate items and scoring scheme by study authors. 2 Expert consensus (agreement or disagreement) on (1) inclusion of shortlisted candidate items; (2) importance of candidate items; and (3) appropriateness of scoring system. Expert provision of feedback on any aspect of Q-SSP checklist. High inter-rater agreement on inclusion of 18/20 of items and importance of 16/20 items. 82% of experts agreed on scoring system. 3 Q-SSP checklist refinement based on (1) goodness-of-fit analysis of experts' responses to the agreement with and importance of items; (2) content analysis of experts' feedback; and (3) quality assessment of survey studies with Q-SSP checklist and inter-rater agreement analysis. Q-SSP checklist item refinement based on goodness-of-fit, content, and interrater agreement analyses. 4 Establishing the capacity of the Q-SSP checklist to distinguish between studies that vary in quality (criterion validity) based on (1) experts' quality assessments; (2) inter-rater agreement analyses; and (3) goodness-of-fit analyses. Some evidence for the criterion validity of the Q-SSP checklist was obtained. Note. Q-SSP ¼Quality of survey studies in psychology. Table 2 Expert panel participant characteristics for each stage of the Q-SSP checklist development. Characteristic Expert Panel Stage 2 Stage 4a Stage 4 b n%n%n% Gender Male 19 57.58 9 90.00 1 10.00 Female 13 39.40 1 10.00 9 90.00 Unspecified 1 3.03 0 0.00 0 0.00 Occupation Assistant professor 3 9.10 0 0.00 0 0.00 Associate professor/Reader 2 6.10 2 20.00 0 0.00 Full professor 4 12.10 0 0.00 0 0.00 Research professor 1 3.03 0 0.00 0 0.00 Professor 2 6.10 0 0.00 0 0.00 Lecturer 1 3.03 1 10.00 2 20.00 Senior lecturer 1 3.03 2 20.00 0 0.00 Research fellow/associate 1 3.03 4 40.00 2 20.00 Research assistant/PhD student 0 0.00 1 10.00 6 60.00 Unspecified 18 54.50 0 0.00 0 0.00 Region and country of residence Europe 19 57.60 6 60.00 5 50.00 Finland 2 6.10 1 10.00 0 0.00 France 0 0.00 1 10.00 0 0.00 Greece 1 3.03 0 0.00 0 0.00 Ireland 2 6.10 0 0.00 0 0.00 Italy 3 9.10 0 0.00 0 0.00 The Netherlands 3 9.10 0 0.00 1 10.00 Spain 0 0.00 0 0.00 1 10.00 Switzerland 1 3.03 0 0.00 0 0.00 UK 7 21.20 4 40.00 3 30.00 North America 3 9.10 0 0.00 3 30.00 Canada 1 3.03 0 0.00 0 0.00 USA 2 6.10 0 0.00 3 30.00 Asia-Pacific 7 21.20 2 20.00 1 10.00 Australia 4 12.10 1 10.00 1 10.00 China and Hong Kong 2 6.10 1 10.00 0 0.00 Singapore 1 3.03 0 0.00 0 0.00 Africa and Middle East 1 3.03 2 20.00 1 10.00 South Africa 1 3.03 0 0.00 0 0.00 Oman 0 0.00 0 0.00 1 10.00 Unspecified 0 0.00 2 20.00 0 0.00 Area of expertise a Psychology 31 93.9 5 50.00 10 100.00 Health psychology 3 9.10 1 10.00 0 0.00 Sport/exercise psychology 3 9.10 3 30.00 0 0.00 Health science 7 21.20 1 10.00 0 0.00 Social science 2 6.10 0 0.00 0 0.00 Evidence synthesis 6 18.20 1 10.00 0 0.00 Quality appraisal 2 6.10 0 0.00 0 0.00 Unspecified 0 0.00 3 30.00 0 0.00 Note. Q-SSP ¼Quality of survey studies in psychology. a Participants could check any of the available options for this characteristic, so categories are not mutually exclusive. The pool of experts was unique at each stage. C. Protogerou, M.S. Hagger Methods in Psychology 3 (2020) 100031 4 end of March 2018. Upon finishing the survey, participants received a closing message, thanking them for their participation and encouraging them to contact the authors if they had any further questions or comments. Participants' responses were recorded by the Qualtrics TM software and downloaded into data spreadsheets for analysis. The study was approved by the Research Ethics Committee of the Psychology Department, [Institution name redacted for masked review] (reference number PSY2017-057). The online consensus survey was divided into five sections. The first section included questions that described participants in terms of age, gender, geographical location, area of expertise, job title, and place of employment. The second section included the candidate set of items each accompanied by scales for participants to rate their agreement for inclusion of the item in the survey and its importance to evaluating study quality. Participants were also prompted to provide further comments and suggestions for modifications, via an accompanying open-ended freeresponse text box. Specifically, for each item participants were prompted to rate (1) their agreement on whether the item should be included in the checklist on a binary agree-disagree scale; and (2) their evaluation of the importance of the item to study quality on a continuous four-point scale (1 ¼not important and 4 ¼very important). The third section included questions prompting participants to rate their agreement with the proposed scoring system for the tool on a binary agree-disagree scale, and comment on the scoring system via an open-ended free-response box. A guide accompanied the tool items, which included definitions of each item, terms used, and details of the scoring system. The guide was downloadable, available throughout the online questionnaire, and participants received periodic reminders to consult it when responding to the items. In the fourth section, participants were asked to state whether or not they had used the guide during the survey and whether they had found it useful. A final question prompted participants to rate their agreement with the proposed title and acronym of the tool. The initial candidate tool items and the guide are available in the online supplement (see Appendices A and B). Stage 3 –Refining the Q-SSP Checklist. The initial pool of checklist items and descriptions were refined based on results from the Stage 2 expert consensus study. We considered 80% agreement our minimum criterion for consensus on participants' ratings for inclusion and importance of each item, based on previous recommendations (Hasson et al., 2000;Keeney et al., 2006). Items falling short of the 80% cut-point for consensus were considered candidates for revision or elimination. We computed goodness-of-fit of participants' responses to the agreement and importance scales with our a priori 80% criterion using chi-square analyses. For the purposes of the goodness-of-fit test, participants’importance ratings were dichotomized. Specifically, “important”and “very important”responses were classified as “high importance”category, and “not important”and “somewhat important”responses were classified as “low importance”. We also content analyzed participants' responses to the open-ended questions on the survey for each item. Content analysis provides new knowledge, insights, conceptual models and practical guides to action (Krippendorff, 1980). We followed Elo and Kyng€ as' (2008) approach, who describe content analysis as a research method for making replicable and valid inferences from data, through a systematic classification process of coding and identifying patterns or themes. Our approach was inductive, i.e., moving from the specific(participants' written comments on quality items) to the general (creating categories and themes describing participants' expectations and requirements about the tool). The first step of the analysis was an open coding procedure, in which entailed multiple readings of participants' responses with extensive notes taken. During open coding, initial categories were generated based on participants' responses that were semantically similar, and by checking the prominence of responses through its co-occurrence. After open coding, the initial categories were grouped under higher-level, broader, abstract categories or themes. The final stage involved applying labels to the extracted themes. Labels were descriptors capturing the essence of each theme. The first author carried out the content analysis and the second author reviewed the results and offered suggestions for revision and refinement. The content analysis is available online (https://osf .io/xgy69). A further step in the refinement stage involved establishing whether independent raters could produce reliable study quality ratings using the tool. The tool was used to assess the quality of 30 survey studies, extracted from three meta-analyses: Hagger et al. (2017),Hoffmann et al. (2017), and Young et al. (2014). These meta-analyses were chosen because (1) they included psychological survey studies (i.e., the study design that the Q-SSP checklist aims to assess); (2) study quality was assessed in the included studies; and (3) the research included studies representing different psychology fields (e.g., health psychology, social psychology, environmental psychology, traffic/transport psychology, sport psychology and social cognition). Furthermore, these meta-analyses provided their complete methods and procedures in online supplements. The authors of the present study, both with expertise in conducting research syntheses (systematic reviews, meta-analyses) in psychology, assessed all 30 studies independently. Inter-rater reliability was computed to evaluate agreement on each of the tool items using Gwet's (2008) AC 1 coefficient. The AC 1 is an alternative to the kappa statistic, used in situations when the extent of agreement between two raters is high but kappa does not appropriately reflect the extent of the agreement (Cicchetti and Feinstein, 1990;Gwet, 2008). Values equal to or greater than 0.50, 0.70, and 0.80 on the AC 1 coefficient denote moderate, good, and very good levels of agreement, respectively (Gwet, 2008). Stage 4 –Criterion Validity of Q-SSP Checklist Scores. An important criterion any study quality assessment checklist is that it can be used to effectively and reliably distinguish between studies that vary in quality. We therefore aimed to evaluate the effectiveness of the Q-SSP checklist prototype in distinguishing between “acceptable”and “questionable”studies from a ‘criterion set’of studies with established quality scores. However, establishing the quality of a criterion set of studies against which the tool is to be assessed, presented considerable challenges. In order to do this, a first step in this process (Stage 4a) was to randomly select a set of studies with known differences in quality based on two criteria: quality assessments using bespoke quality assessment tools from previous studies and expert consensus. The random selection of studies was done with the use of the online random number generator https://www.random.org/. We selected 20 studies from three previous meta-analyses (Hagger et al., 2017;Hoffmann et al., 2017;Young et al., 2014), 10 that were rated as having good (“acceptable”) quality and 10 that were rated as having poor (“questionable”) quality based on the assessment tools used in the individual studies. We then asked a panel of judges (N¼10), independent from the panels from the previous stages, with expertise in evidence synthesis and/or study quality appraisal to provide an assessment of the quality of each of the 20 studies based on their expert opinion. The subsequent step (Stage 4 b) aimed to examine whether Q-SSP scores were able to distinguish between studies identified as acceptable and questionable in quality in Stage 4a. Eligible experts (N¼33) were identified through their previous publication track record in the field (also refer to our Participants section) and invited to participate in the study by email. Ten agreed to participate (response rate ¼33.33%), and all who agreed subsequently completed their assessments (completion rate ¼100%). As a guide, judges were provided with a brief narrative identifying the typical expected criteria used to evaluate study quality; a summary of criteria derived from previous quality assessment tools. However, judges were asked to use their own judgment and bring to bear their experience and expertise in making their evaluations. The judges' appraisals were pooled and consistency evaluated using ICC. It was expected that this would provide a set of studies on which there was general consensus from the judges, along with the evaluations from the bespoke study quality assessment tools used in the original meta-analyses from which the set of studies was drawn, on their quality. Then, a final panel of experts (N¼10), also independent of the previous panels in previous C. Protogerou, M.S. Hagger Methods in Psychology 3 (2020) 100031 5 stages, with similar experience in evidence synthesis and/or quality assessment was then asked to use the Q-SSP checklist provide quality assessment scores for each of the studies. This last panel was also identified through their publications and research track record in the field. Twenty judges were invited, by email, to participate; 10 agreed to participate (response rate ¼50%) and all carried out assessments to completion (completion rate ¼100%). Table 2 presents participants' characteristics in Stages 4a and 4 b. We then compared overall study quality scores (“acceptable”vs. “questionable”) derived from the Q-SSP checklist with the consensus quality judgments of the experts and the assessments from the tools used in the meta-analyses from which the set of studies was drawn using percentage agreement and Gwet's (2008) AC 1 coefficient. We also evaluated the goodness-of-fit of the quality score for each study using the Q-SSP checklist with scores from the expert judges. High agreement and close fit for the Q-SSP checklist and expert judgement scores would provide preliminary support for the criterion validity Table 3 Inter-rater agreement and consensus survey agreement statistics for the Q-SSP checklist. Item# Domain and item description a Inter-rater reliability Consensus Agreement AC 1 PIncluded Importance Agreement χ 2 p M SD Mdn. Agreement χ 2 P 1. Introduction 1. Were hypotheses or aims explicitly stated? 96.67% .963 <.001 100.00% ––3.606 0.659 4 93.33% 2.455 .117 2. Introduction 2. Were operational definitions of predictor (independent) and outcome (dependent) variables provided? 93.33% .880 <.001 87.87% 1.280 .258 3.273 0.801 3 66.67% 0.030 .862 3. Introduction 3. Were participant eligibility criteria (inclusion and exclusion) explicitly stated? 86.67% .781 <.001 100.00% ––3.394 0.747 4 83.33% 0.485 .486 4. Introduction 4. Were participants recruited using an acceptable recruitment strategy? 93.33% .876 <.001 87.87% 1.280 .258 3.091 0.843 3 63.3% 0.371 .542 5. Participants 1. Were participants selected by a random/probability sampling strategy? 90.00% .817 <.001 75.76% 0.371 .542 2.818 0.983 3 70.00% 3.667 .056 6. Participants 2. Was the sample size appropriate? 90.00% .849 <.001 96.97% 5.939 .015 3.576 0.614 4 63.3% 4.008 .045 7. Participants 3. Were participants randomly assigned into groups/ conditions? 96.67% .958 <.001 81.82% 0.068 .794 3.121 1.023 3 83.33% 1.091 .296 8. Data 1. Was the response/ participation/recruitment rate provided? 83.33% .719 <.001 87.87% 1.280 .258 3.121 0.857 3 83.33% 1.280 .258 9. Data 2. Was the attrition rate acceptable? 73.33% .856 <.001 84.84% 0.485 .486 3.091 0.765 3 80.0% 0.068 .794 10. Data 3. Was the attrition rate treated appropriately in data analyses? 86.67% .815 .004 87.87% 1.280 .258 3.242 0.708 3 66.67% 0.485 .486 11. Data 4. Were the chosen statistical tests appropriate to address hypotheses or research questions? 100.00% 1.000 <.001 100.00% ––3.727 0.517 4 93.33% 5.939 .015 12. Data 5. Did the study include a formative research or pilot phase? 83.33% .719 <.001 69.70% 2.189 .139 2.303 0.810 2 73.3% 34.008 <.001 13. Data 6. Were the measures provided in the report (or in a supplement) in full? 80.00% .723 <.001 84.84% 0.485 .486 2.909 0.980 3 80.0% 1.091 .296 14. Data 7. Were all measures of established validity, or was a validation procedure undertaken by the authors? 96.67% .944 <.001 90.91% 2.455 .117 3.212 0.857 3 46.6% 0.030 .862 15. Data 8. Was the study sample described in terms of key demographic characteristics? 90.00% .817 <.001 96.97% 5.939 .015 3.303 0.684 3 86.6% 1.280 .258 16. Data 9. Was the data collection process described with sufficient detail for it to be replicated? 80.00% .706 <.001 96.97% 5.939 .015 3.424 0.830 4 76.6% 0.485 .486 17. Data 10. Were generalizations of findings restricted to the population from which the sample was drawn? 90.00% .805 <.001 78.87% 0.030 .862 2.909 0.914 3 56.6% 3.667 .056 18. Ethics 1. Was the study approved by a relevant institutional review board or research ethics committee? 100.00% 1.000 <.001 90.91% 2.455 .117 3.364 0.822 4 96.67% 0.030 .862 19. Ethics 2. Did participants provide informed consent (or assent, where relevant)? 96.67% .933 <.001 84.84% 0.485 .486 3.061 0.966 3 96.67% 2.189 .139 20. Ethics 3. Were funding sources or conflicts of interest disclosed? 93.33% .880 <.001 90.91% 2.455 .117 3.242 0.902 4 83.33% 0.371 .542 Note. Q-SSP ¼Quality of survey studies in psychology; AC1 ¼Gwet (2008) AC 1 agreement coefficient; χ 2 ¼Goodness of fit chi-square; t¼Independent samples t-test of difference from scale mid-point. a Items listed in the table are those presented to participants in the expert consensus study prior to revision. C. Protogerou, M.S. Hagger Methods in Psychology 3 (2020) 100031 6 of scores produced by the Q-SSP checklist. Results Stage 1 –initial item pool The review of the literature and existing study quality tools produced the initial list of candidate items for the subsequent development stage of the study using expert consensus. The initial list is available online (https://osf.io/xgy69). Stage 2 –expert consensus Participants in the expert panels for the second stage of the Q-SSP checklist development were university faculty and researchers (N¼33; age M¼45.30, SD ¼8.31) from fourteen countries. Participant characteristics including region and country of origin, academic rank, areas of expertise, and gender are presented in Table 2. Agreement among raters on the inclusion and importance ratings for each quality assessment item are presented in Table 3, and the data files and analysis scripts are available online (https://osf.io/xgy69). Participants demonstrated very high agreement on inclusion and importance ratings for the majority of the items. Specifically, agreement ratings on whether the item should be included in the checklist was above our 80% criterion for 18 out of the 20 items. Exceptions were items 5 (“Were participants selected by a random/probability sampling strategy?”) and 12 (“Did the study include a formative research or pilot phase?”), which had 75.8% and 69.7% agreement, respectively. These agreement proportions fell short of our 80% criterion, although a 70% agreement criterion is often considered acceptable (Hasson et al., 2000;Keeney et al., 2006). Goodness-of-fit chi-square analysis revealed statistically non-significant values for all but three of the items. Specifically, agreement was significantly lower than the 80% criterion for items 5 (“Were participants selected by a random/probability sampling strategy?”; χ 2 (1) ¼5.939, p<.015), 15 (“Was the study sample described in terms of key demographic characteristics”; χ 2 (1) ¼5.939, p<.015), and 16 (“Was the data collection process described with sufficient detail for it to be replicated”; χ 2 (1) ¼5.939, p<.015). Regarding the agreement ratings for the importance of each study quality item, results indicated that 16 of the 20 items were rated 3 or above on the 4-point scale. Mean scores for items 5 (“Were participants selected by a random/probability sampling strategy?“,12“Did the study include a formative research or pilot phase?“,13“Were the measures provided in the report (or in a supplement) in full?“, and 7 “Were generalizations of findings restricted to the population from which the sample was drawn?”ranged between 2.82 and 2.91. Chi-square tests indicated that agreement was high, with non-significant chi-square values indicating no difference from the 80% criterion for all but two items: item 6 (“Was the sample size appropriate?“; χ 2 (1) ¼4.008, p<.045), and 11 (“Were the chosen statistical tests appropriate to address hypotheses or research questions”; χ 2 (1) ¼5.939, p<.015). The lower rates of agreement for these items suggested that further scrutiny of participants’responses to the open-ended comments for these items was warranted. Finally, twenty-seven participants (82%) agreed with the proposed scoring system and the majority of participants (n¼25, 76%) consulted the guide when making their assessments. Stage 3 –- Q-SSP checklist refinement based on content-analysis and interrater agreement analysis Content analysis. Four themes emerged from the content analysis of participants' written responses to the open-ended questions: clarity, generalizability, transparency, and scoring flexibility. The themes summarize participants' comments, suggestions, and expectations concerning the content of tool and guide. Details of the content analysis including participants' comments on each item, the co-occurrence of comments, suggestions for improvement, emerging themes, and the steps undertaken to meet participants’expectations are presented online (https://osf .io/xgy69). The most prominent theme emerging from the content analysis was the need for clarity in the terminology and wording of the quality items and the guide. Participants identified ambiguity and lack of clarity in some of the terms used. In particular, the terms “appropriate”,“sufficient”,and“acceptable”were flagged as problematic, due to their potential to confuse users of the tool, the provision of relevant definitions in the accompanying guide notwithstanding. The expectation for the tool to be generalizable across survey designs (e.g., pencil-andpaper, online, quantitative, qualitative), research questions, and regulations of academic institutions, was a consistent theme. For example, some participants indicated that not all universities have ethics committees or IRBs, and that in some countries survey studies are exempt from committee or IRB approval. It was also prominently indicated that including random assignment and probability sampling as quality criteria would not be relevant to surveys employing other sampling and assignment methods. The imperative of transparency in reporting emerged as a theme, with suggestions to rephrase items to prioritize transparency as a study quality criterion. Some participants suggested that published studies in psychology tend not to report crucial information (e.g., attrition rates, a priori sample estimation), and such nonreporting diminishes study quality. It was therefore suggested that identifying whether particular quality criteria are reported “at all”may be a more appropriate for some of the items. The expectation that the tool allows for a degree of flexibility in scoring was suggested. For example, it was recommended that the scoring of the quality domains could be non-numerical, and even optional. Refinement of Checklist Items. Based on the inter-rater agreement Table 4 Final Q-SSP checklist items. Item # Domain Item 1 Introduction Was the problem or phenomenon under investigation defined, described, and justified? 2 Introduction Was the population under investigation defined, described, and justified? 3 Introduction Were specific research questions or hypotheses stated? 4 Introduction Were operational definitions of all study variables provided? 5 Participants Were participant inclusion criteria stated? 6 Participants Was the participant recruitment strategy described? 7 Participants Was a justification/rationale for the sample size provided? 8 Data Was the attrition rate provided? (applies to cross-sectional and prospective studies) 9 Data Was a method of treating attrition provided? (applies to crosssectional and prospective studies) 10 Data Were the data analysis techniques justified (i.e., was the link between hypotheses/aims/research questions and data analyses explained)? 11 Data Were the measures provided in the report (or in a supplement) in full? 12 Data Was evidence provided for the validity of all the measures (or instrument) used? 13 Data Was information provided about the person(s) who collected the data (e.g., training, expertise, other demographic characteristics)? 14 Data Was information provided about the context (e.g., place) of data collection? 15 Data Was information provided about the duration (or start and end date) of data collection? 16 Data Was the study sample described in terms of key demographic characteristics? 17 Data Was discussion of findings confined to the population from which the sample was drawn? 18 Ethics Were participants asked to provide (informed) consent or assent? 19 Ethics Were participants debriefed at the end of data collection? 20 Ethics Were funding sources or conflicts of interest disclosed? Note. Q-SSP ¼Quality of survey studies in psychology. C. Protogerou, M.S. Hagger Methods in Psychology 3 (2020) 100031 7 analysis and the written responses of the expert panel, items from the initial version of the Q-SSP checklist were revised. The revised items are presented in Table 4. 3 Revisions primarily involved some re-phrasing of checklist items and the guide, with the goal of improving clarity, generalizability, transparency in reporting, and flexibility in scoring. Most revisions were made on the basis of participants’responses to the open-ended comments for each item checked against responses to the inclusion and importance ratings. In addition, items 5 (“Were participants selected by a random/probability sampling strategy?”), 7 (“Were participants randomly assigned into groups/conditions?”), and 12 (“Did the study include a formative research or pilot phase?”) were removed in response to specific feedback provided by participants. Our experts pointed out that random assignment and random sampling, as well as the inclusion of formative research elements, do not often apply to studies adopting survey designs, and the absence of these elements may not necessarily impact study quality. While items 16 (“Was the data collection process described with sufficient detail for it to be replicated?”) and 18 (“Was the study approved by a relevant institutional review board or research ethics committee?”) were considered important items, they were substituted for other items. Specifically, item 16 was considered to encompass more than one quality dimension and was therefore replaced with items assessing separate criteria deemed essential to study replication. The criteria were based on recommendations and guidelines from reviews and commentaries on replication (Asendorpf et al., 2013;Norris et al., 2016;Schroter et al., 2012). The item was divided into four separate items: “Was information provided about the person(s) who collected the data (e.g., training, expertise, other demographic characteristics)?”; “Was information provided about the context (e.g., place) of data collection?”;“Was information provided about the duration (or start and end date) of data collection?”; and “Was the participant recruitment strategy described?” Similarly, item 18 was replaced by two items reflecting ethical conduct: “Were participants asked to provide (informed) consent or assent?”and “Were participants debriefed at the end of data collection?”Participants had alerted us to the fact that ethics committees do not exist in all countries or academic departments, and that survey studies are sometimes exempt from ethical or IRB approval. The substitute items gauge ethical procedures considered essential in human survey research, even in the absence of formal ethical approval, based on published guidelines (Appelbaum et al., 2018). According to these guidelines, informed consent and debriefing procedures are sufficient to cover issues surrounding distress, deception, lack of confidentiality and participant rights. Finally, we made minor changes based on inter-rater agreement results from the previous stage (see Table 3). Gwet (2008) AC 1 coefficients indicated acceptable agreement (AC 1 >0.70) across the rated studies for all items, with overall agreement levels >80% for all items. Disagreements were resolved through discussion. Without exception, disagreements stemmed from minor variations in the interpretation of quality criteria. Resolution of disagreements resulted in minor revisions to the Q-SSP checklist guide to clarify issues that led to the disagreements. For example, the term ‘context’used in the item “Was information provided about the context (e.g., place) of data collection?”was sometimes misinterpreted in studies that collected data via phone and internet. This led to adding text in the guide, further explaining the meaning of data collection “context”and “place”. The final version of the QSSP and its accompanying guide are presented in Appendix A (supplemental materials). Scoring System Development. Quality items in the QSSP are scored with the options: “yes”,or“no”,“not stated clearly”,or“not applicable”, based on the information provided in the research report (e.g., article, poster, protocol, thesis) and supplemental material, if available. Quality appraisal and scoring is expected to be based solely on information provided in the published study and any accompanying supplemental materials, instead of raters’interpretation of study elements that may be absent or missing from the report. The quality criteria are grouped into four domains: introduction (study rationale and variables), participants (sampling and recruitment), data (data collection, analyses, results and discussion), and ethics (consent, debrief, and funding/conflicts of interest). The domains represent groups of items designed to assess conceptually-similar aspects of study quality. For example, items gauging procedures of consent, assent, and debriefing, as well as the disclosure of funding sources and conflicts of interest, are grouped into the ethics domain, and items gauging participant inclusion criteria, recruitment strategies and sample size rationale, are grouped into the participants domain. These domains are colour coded on the scoring sheet to facilitate scoring. An overall quality numerical score is a percentage calculated by dividing the “yes”answers to the quality items by the total number of applicable items. Based on this scoring system, studies are categorized as having “questionable”quality if they do not receive “yes”responses for five or more checklist items, otherwise studies are classified as having “acceptable”quality. Depending on the number of applicable items, a study should receive a “yes”response to between 70% and 75% of items to receive an overall “acceptable”quality score. This criterion corresponds well with recommended cut-offs offered by other general study quality assessment tools (e.g., Glynn, 2006;Husebøet al., 2012). However, it must be stressed that such cut-off values are arbitrary, and other less stringent cut-off values have been proposed. Cut-off values should also be viewed in light of the concerns regarding the use of overall quality scores rather than domain or individual item scores. Domain-specific scores are simple ratios calculated by dividing “yes” scores by the number of applicable items in each domain. Non-applicable choices are shaded on the scoring sheet to ensure that only appropriate options are selected during scoring. As some items are considered essential to all studies, the “not stated clearly”or “not applicable”options are not considered appropriate (e.g., “Was the problem or phenomenon under investigation defined, described, and justified?”and “Was the population under investigation defined, described, and justified?”). Numerical scoring is at the discretion of users of the Q-SSP checklist. The QSSP checklist comes with a guide providing definitions and examples of the terms used in the checklist, and guidance on scoring (see checklist in Appendix A). Stage 4 –criterion validity of Q-SSP checklist scores The final stage examined the effectiveness of the Q-SSP checklist in distinguishing between studies of known difference in quality. Ten experts (age M¼33.70, SD ¼4.19) decided whether a set of 10 studies with known differences in quality were of acceptable or questionable quality, based on their extant knowledge and experience (participant characteristics are presented in Table 2). Quality assessment ratings for each study based on the published ratings from the source meta-analysis and ratings of each panel member based on their expertise and the data and analysis scripts are available online (https://osf.io/xgy69). Averaged inter-rater agreement for each study across the experts was good (ICC ¼0.75, p<.001), and final consensus-based ratings are also available online (https://osf.io/xgy69). Overall, eleven studies were judged to be of ‘questionable quality’by a majority of the experts (60% agreement), while only four studies were judged to be of ‘acceptable’quality based on the same criterion. The judges were split on their evaluation of the remaining five studies. Next, a different panel of experts (N¼10; age M¼33.33; SD ¼6.04) used the Q-SSP checklist to assess the quality of each of the final studies 3 Original and revised versions of the checklist items and guide are provided online (https://osf.io/xgy69). The finalized version of the checklist and guide is also provided in Appendix A (supplemental materials). C. Protogerou, M.S. Hagger Methods in Psychology 3 (2020) 100031 8