scieee AI-readable full text Open interactive document viewer

Reporting Reliability in Qualitative Engineering Education Research: A Literature Review

Martin, D.; de Andrade, M.; Hunt, S.; Ramachandran, V.; Dokter, L.

Abstract

This study examines the reporting transparency of qualitative coding reliability in engineering education research through a systematic review of articles published in the European Journal of Engineering Education. Using a structured framework derived from qualitative methodology literature, we analysed how researchers document their coding processes, reliability measures, and disagreement resolution procedures. Our findings reveal major variation in reporting practices, with notable gaps in describing coder information, reliability metrics, and disagreement resolution processes. Based on these empirical findings and established methodological guidelines, we offer specific recommendations for enhancing reliability reporting in qualitative EER. This work contributes to the methodological development of EER as the field increasingly addresses complex, interpretative research questions requiring robust qualitative approaches.

Full text

Research Paper Recommended citation: Martin, D., de Andrade, M., Hunt, S., Ramachandran, V., & Dokter, L. (2025). Reporting Reliability in Qualitative Engineering Education Research: A Literature Review. In Kangaslampi, R., Langie, G., Järvinen, H.-M., & Nagy, B. (Eds.), SEFI 53rd Annual Conference. European Society for Engineering Education (SEFI), Tampere, Finland. DOI: 10.5281/zenodo.17631625. This Conference Paper is brought to you for open access by the 53rd Annual Conference of the European Society for Engineering Education (SEFI) at Tampere University in Tampere, Finland. This work is licensed under a Creative Commons Attribution-NonCommercial-Share Alike 4.0 International License. REPORTING RELIABILITY IN QUALITATIVE ENGINEERING EDUCATION RESEARCH: A LITERATURE REVIEW D.A. Martin a, 1 , M. de Andrade a, S. Hunt a, V. Ramachandran a, L. Dokter a, a Centre for Engineering Education, University College London, London, UK Conference Key Areas: Improving higher engineering education through researching engineering education Keywords: qualitative research, data analysis, reliability, methodology ABSTRACT This study examines the reporting transparency of qualitative coding reliability in engineering education research through a systematic review of articles published in the European Journal of Engineering Education. Using a structured framework derived from qualitative methodology literature, we analysed how researchers document their coding processes, reliability measures, and disagreement resolution procedures. Our findings reveal major variation in reporting practices, with notable gaps in describing coder information, reliability metrics, and disagreement resolution processes. Based on these empirical findings and established methodological guidelines, we offer specific recommendations for enhancing reliability reporting in qualitative EER. This work contributes to the methodological development of EER as the field increasingly addresses complex, interpretative research questions requiring robust qualitative approaches. 1 Corresponding Author D.A. Martin [email protected] 1 INTRODUCTION The field of engineering education research (EER) has experienced substantial growth over the past two decades, with researchers increasingly adopting qualitative methodologies to explore the complex social, cultural, and pedagogical dimensions of engineering education (Borrego et al., 2019; Case & Light, 2011). However, as qualitative methods gain prominence in engineering education research, the transparency with which these methods are reported become increasingly important to ensure the field's methodological rigor and the capacity development of engineering education researchers (Walther et al, 2017). Qualitative analysis encompasses a diverse set of methodologies. Braun and Clarke (2006) distinguish between approaches tied to specific theoretical positions (e.g., narrative or discourse analysis) and those that are more flexible and applicable across paradigms (e.g., thematic analysis, grounded theory, framework and content analysis). This paper focuses on the latter, highlighting coding methods that support methodological rigour. The interpretive nature of qualitative research relies heavily on transparent reporting to ensure trustworthiness and transferability (Tracy, 2010). The processes of developing and refining codebooks, coding qualitative data, establishing reliability among researchers, and resolving coding disagreements are fundamental to qualitative analysis (Saldaña, 2021). These methodological details are not mere procedural formalities but rather essential components that enable readers to evaluate research quality, contextualise findings, and potentially build upon established analytical frameworks in subsequent studies (O'Connor & Joffe, 2020). In interdisciplinary fields such as engineering education, where qualitative research traditions from the social sciences intersect with the epistemological norms of technical disciplines, researchers often struggle to translate qualitative methodological principles into the cultural and disciplinary context of engineering (Koro-Ljungberg & Douglas, 2008). This translation challenge appears in empirical studies that either leave out essential details about qualitative processes or present them in ways that fail to fully demonstrate methodological rigour based on established qualitative practices (Koro-Ljungberg & Douglas, 2008; Walther et al., 2013). Additionally, the longstanding quantitative focus of engineering disciplines has shaped a research culture that may not fully recognize the distinct forms of rigour and quality essential to qualitative inquiry. (Borrego & Bernhard, 2011; Malmi et al, 2018; Walther et al., 2013). The European Journal of Engineering Education (EJEE) represents one major publication venue for EER, particularly in the European context, and has published an increasing number of qualitative studies over the past decade (Bernhard, 2018). A previous more general review noted that two-thirds of the EER papers published in this journal in 2009, 2010, and 2013 did not discuss the validity and trustworthiness of their findings (Malmi et al, 2018). To date, no systematic analysis has focused specifically on how qualitative methods of data analysis are reported in EER journals. These include core processes of developing a codebook, coding and interpreting data, determining reliability, identifying interpretative differences, and resolving disagreements. This gap is notable given that previous reviews in more methodologically advanced fields such as STEM education (Cheung & Tai, 2023), higher education (Liao & Hitchcock, 2018), communication studies (Feng, 2014), or human computer interaction (McDonald et al, 2019) have identified inconsistent and often inaccurate and incomplete reporting of the process of data analysis. The aim of this literature review is to build on previous efforts for enhancing the reporting of qualitative methodologies in EER, and by extension, the quality of researchers’ chosen methods (i.e. Walther et al., 2013; Malmi et al, 2018, Koro Ljungberg & Douglas, 2008 a.o.). This review analyses articles from a single issue of the EJEE that use qualitative methods, either independently or as part of mixed methods designs. These articles are examined with attention to their approaches to codebook development, coding processes, and data analysis procedures. Additionally, the review examines how these studies report reliability measures and indices, the management of coding disagreements, and any limitations pertaining to the coding process or the coders themselves. By conducting this focused analysis, this review addresses three significant gaps in current EER understanding: First, it provides insight into the specific qualitative reporting practices within the context of EER disseminated via the major European EER journal, which may differ from practices in other regional contexts (per Wankat, Williams, & Neto, 2014; Williams, Wankat & Neto, 2016). Second, it offers a detailed examination of key methodological elements essential for ensuring the quality of qualitative research yet overlooked in previous reviews (such as Malmi et al, 2016). Third, it establishes a baseline understanding of current reporting practices in a major EER journal that can inform future methodological guidance for engineering education researchers and editorial policies for authors, reviewers and editors alike. This literature review is structured to first outline the methodological approach used to select and analyse articles from the selected EJEE issue. It then presents findings on the reporting of qualitative methodology and methods components. The review concludes with recommendations for improving qualitative methods reporting in engineering education research and suggestions for the next steps. 2 METHODOLOGY This literature review employed a systematic approach to analyse qualitative methods reporting in the EJEE through an interpretive lens. The content analysis focused on how authors describe their coding processes, reliability measures, codebook development, and disagreement resolution strategies. The research team comprised five members with expertise spanning qualitative and quantitative methods. Data was sampled from EJEE Volume 48, Issues 1-6, covering articles published in 2022 or 2023. The inclusion criteria required that articles (1) explicitly employed qualitative methods for data collection and analysis. Mixed-methods studies were included if they contained qualitative components; (2) Were empirical studies rather than theoretical papers or literature reviews. Of the 68 articles in this issue, 46 articles met these criteria and were included in the analysis. 2.1 Analytical process Each article was analysed using a deductive codebook developed based on keywords, descriptors, criteria, and guidelines in the qualitative methodology literature (Bloor & Wood, 2006, Levitt et al., 2018; O'Connor & Joffe, 2020; Saldaña, 2021; Tracy, 2010; Twining et al., 2017). The codebook framework identified key dimensions of qualitative methods reporting alongside the methodology and methods explicitly acknowledged by the authors in the sample articles (Table 1). Table 1. Analytical framework for the codebook Category Items Input type Methodology reported What research methodology the authors explicitly claim to use, as well as the data collection and analysis methods mentioned Multiple choice with more than one choice allowed Coding process documentation Whether the authors describe their approach to coding (deductive, inductive, or hybrid), the number of coders and their role, the details provided about coding procedures and any justification about coding choices. Multiple choice with more than one choice allowed Reliability documentation Whether the authors state they established reliability, including any metrics or indices reported, intercoder/ intracoder reliability measures, limitations purporting to the coding process or coders, positionality or reflexivity statements affecting the coding process. Multiple choice with only one choice allowed Codebook development Whether the authors describe developing and refining their codebook and whether they include code definitions or examples. Multiple choice with more than one choice allowed Disagreement resolution Whether the authors acknowledge coding disagreements and the criteria for resolving discrepancies between coders. Multiple choice with more than one choice allowed Each coding category and subcategory were translated into items for a survey hosted on Qualtrics. This survey was used by the members of the research team to input their analysis of the articles, in three stages: For the pilot stage, three articles were selected by Author 1 and coded by Authors 1-4 based on the framework in Table 1. Each coder conducted their analysis independently, via Qualtrics. The coders reconvened alongside Author 5 to describe and explain their coding choices and align their discrepancy in interpretations via a consensus discussion. This discussion led to the refinement of the codebook and to restructuring the Qualtrics survey for ease of use. For the second stage, Authors 1 and 3 read all articles independently and Authors 2, 4 and 5 read the same 24% of all articles, with the aim of identifying how the aspects listed in the analytical framework in Table 1 are reported. Authors coded each article individually without input from others. Aspects present in Table 1 were typically found in the methodology section of the article, though relevant information sometimes appeared in other sections, such as the introduction or conclusion. In some cases, the results section also described the process of data analysis, with the research design aspects intertwined with the presentation of the research findings. For the third stage, Author 2 conducted a quantitative cross-article analysis based on the inputs of all coders. This synthesised the researchers’ interpretations of the reporting of qualitative data analysis in the sample of articles, while taking note of discrepancies in our interpretations. The cross-article analysis also aimed at estimating inter-coder reliability within articles. Krippendorf’s alpha was chosen as a reliability metric due to its suitability for multi-coder, multi-category qualitative data and its alignment with established criteria for robust agreement measures (Hayes & Krippendorff, 2007). Unlike metrics such as percent agreement or Cohen’s Kappa, which are limited to two coders or nominal data, alpha accommodates varying numbers of coders, handles missing data, and applies across measurement levels. In our review, Alpha was used as a diagnostic metric to highlight specific articles where coding disagreements occurred. Therefore, Krippendorf’s alpha was calculated within each article for the inputs of at least two coders to the items set out in Table 1, resulting in the distribution of alpha values (one per article) shown in Figure 1. The average alpha across all articles was 0.83, where only 6 out of 44 articles (13.6%) did not meet the 0.7 threshold. Articles with an alpha below 0.7 were selected for further analysis by Author 2 towards clarifying the nature of the disagreement. Figure 1: Distribution of Krippendorf's Alphas across 44 of the 46 articles analysed herein. The vertical dashed line represents an alpha of 0.7 which was taken as the cut-off for reliability in the present article. 2.2 Low reliability cases Analysis of the six articles where inter-coder reliability as quantified by Krippendorf’s alpha were lower than 0.7 revealed important sources of disagreement as well as limitations in our implementation of the metric. In general, disagreement occurred due to many degrees of freedom in coding such as being able to choose more than one category in the items present in Table 1 or non-specific options such as “Other” that aimed at capturing categories not listed in the deductive coding scheme. In three out of the six articles, Authors 1 and 3 diverged in deciding between what characterised minimal mention of inter-coder reliability in the article and no mention at all. This could be a highly subjective item since the authors could have been looking for explicit and literal mentions (identifying a specific short text window addressing the item) and diffuse mentions (large text window, possibly implicit or inferred). There was one disagreement with regards to the methodology (“Other” and “No methodology reported explicitly”), and two with regards to the data collection (“Artefact analysis” and “Other”). A significant source of disagreement in terms of the quantification via Krippendorf’s alpha seemed to be the possibility of allowing researchers to choose multiple options in the items outlined in Table 1, which was present in 4 out of the 6 entries with low reliability. Since Krippendorf’s alpha was calculated at the nominal level, entries such as (Code A + Code B + Code C) are interpreted as a different code to (Code B + Code C) because partial agreement was not considered. 2.3 Methodological limitations This review focuses on a single issue of EJEE, limiting generalisability to the broader landscape of engineering education research. Additionally, the analysis examines only what authors explicitly reported about their methodological processes, which may differ from the actual procedures employed. The focus on published descriptions also means that space limitations imposed by the journal may have constrained authors' methodological reporting. Despite these limitations, this focused analysis provides insight into reporting practices and may serve as a foundation for developing thorough reporting guidelines for qualitative EER. 3 RESULTS A total of 44 articles were analysed by both Authors 1 and 3, and 2 additional articles were only analysed by Author 1. Here we report descriptive findings of Author 1’s coding of the 46 articles, of which 27 were qualitative and 19 employed mixed methods. Methodology: 34 articles (74%) did not mention their methodology explicitly, 5 (11%) used case studies, 3 (7%) used ethnography, 2 (4%) used grounded theory, 2 (4%) used action research, 1 (2%) used phenomenology, 1 (2%) used narrative research, and 2 articles (4%) used other methodologies. Four articles mentioned more than one methodology hence the sum of the quantities reported above is 50. Data collection: 18 articles (39%) collected data via interviews, 12 (26%) via document/artefact analysis, 10 (22%) via open-ended surveys, 9 (20%) via observation, 7 (15%) via focus groups, 1 (2%) did not report the method of data collection, and 1 (2%) reported it as a case study. Nine articles mentioned two or more methods of data collection. Number of coders: 25 articles (54%) did not explicitly mention how many coders were involved in the coding process, 3 (7%) reported one coder, 11 (24%) reported two coders, 5 (11%) reported three coders, and 2 (4%) reported more than three coders. Inter-coder reliability: 26 articles (57%) did not mention inter-coder reliability, 17 (37%) made minimal mention of reliability but did not report it alongside formalised metrics, and only 4 (9%) articles mentioned coding reliability in detail and included metrics. Metrics of reliability: 42 articles (91%) did not report any reliability metrics, and 4 articles (9%) quantified reliability using Cohen’s Kappa. Disagreement resolution: 33 articles (72%) did not report a process for resolving coding disagreements. Among those that did, 8 articles (17%) described resolving disagreements through coder discussion until consensus was reached, and 2 articles (4%) reported expert adjudication. Five articles (11%) described using iterative refinement of the codebook to address recurring coding inconsistencies, of which three also included consensus discussions. One article (2%) referred to disagreement resolution without specifying the method employed. 4 DISCUSSION AND CONCLUSIONS This review of qualitative methods reporting in EJEE reveals both strengths and gaps in how researchers document their coding processes, reliability measures, codebook development, and disagreement resolution strategies. By identifying these patterns, this study makes several contributions to advancing methodological rigor in EER. The findings highlight the importance of developing guidelines for qualitative research reporting in the engineering education context. Our recommendations emerge from the integration of two data sources: insights from surveying how qualitative data analysis was reported in issue 48 of EJEE, which revealed inconsistencies and omissions in reliability reporting (Table 3) and guidelines from the broader methodological literature in qualitative research (Table 4). This evidence-based approach ensures the recommendations are both responsive to documented needs in EER and grounded in established methodological principles (see Feng, 2014; Fleiss, 1971;Gwet, 2021; Hayes & Krippendorff, 2007; McPhail et al, 2015; Neuendorf, 2017; O'Connor & Joffe, 2020; Saldaña, 2021) Table 3. Recommendations for conducting and reporting data analysis in qualitative engineering education research Category Recommendations Coder information Document the number of coders – state how many researchers were involved in coding and their role or backgrounds. Coding process Explicitly state the analytical approach – clearly identify whether the coding uses a deductive, inductive, or hybrid approach, with rationale for this choice. Indicate the coding unit – clarify whether codes were applied to sentences, paragraphs, meaning units, or other segments. Include sample coded excerpts – present examples of raw data with applied codes to illustrate the coding process and interpretation. Report calibration activities - If applicable, detail team calibration activities Report coding software or tools used – specify use of data analysis software (e.g., NVivo, ATLAS.ti), including generative AI functions. For deductive approaches: Code source - detail the origins of the codes (e.g., theoretical models, conceptual frameworks, prior studies, literature). For inductive or hybrid approaches: Saturation - report if data was collected until code saturation and how it was assessed. Document coding iterations – provide a description of the iterative process used in developing the coding scheme, including any calibration sessions and practice coding of sample data. Present the development of the coding scheme – explain how codes were created, defined, refined, and organized into themes or categories. This may include inclusion or exclusion criteria for each coding category. If possible, specify the coding unit. Reliability measurement Calculate and report inter-rater or intra-rater reliability metrics – provide appropriate statistics with justification for the chosen metric considering the type of data and number of coders. Note that deductive coding is more amenable to quantifying inter-coder reliability than inductive coding, where intra-coder reliability might be more useful. Specify the sample size used for assessing reliability – mention what portion of the total data was used to calculate reliability measures (typically 15-30%). Disagreement resolution Detail the disagreement resolution protocol – describe the specific process for addressing coding discrepancies, including how decisions were made. The most reported processes are consensus discussion, third party arbitration, the expert decision of the PI a.o. Provide examples of coding disagreements – include representative examples of coding conflicts and their resolution. Document changes to the codebook – track and consider reporting modifications made to the coding scheme as a result of disagreements or emergent insights. Reflexivity and limitations Include positionality statements – reflect on how researchers' characteristics may influence coding decisions and interpretations. Discuss coding limitations transparently – Identify constraints in the coding process and potential impacts on findings (e.g. large time gaps between data collection and analysis, changes to the research team, power relations between coders) Table 4. Recommended data analysis strategies for different number of coders Coders Recommended strategies Optional metrics 1 coder The same coder re-codes a portion of the data and compares coding decisions made at two different time points. Maintain analytic memos to track coding rationale and shifts in interpretation. 2 coders Independent coding, blind review, reciprocal coding. Cohen's Kappa κ (popular for nominal data, limited to 2 coders, sensitive to bias) Percentage agreement (does not account for chance agreement) Weighted Kappa (for ordinal data where degrees of disagreement matter, i.e. Likert scales or ordinal themes, requires