scieee AI-readable full text Open interactive document viewer

A Hybrid Assessment Approach to AI-Enhanced Challenge-Based Learning in Engineering Education

Garcia Huertes, S.; Bragós, R.

Abstract

This paper introduces a hybrid assessment approach integrating Generative AI (GenAI), specifically ChatGPT, into challenge-based engineering education. Students iteratively defined problem statements, brainstormed solutions, and presented their work, receiving feedback through rubric-based GenAI assessments of ChatGPT interactions and direct faculty evaluation. While ChatGPT positively influenced student performance and intellectual engagement, it amplified existing differences in motivation, prior knowledge, skills, and team dynamics, underscoring the importance of comprehensive, human-centered pedagogical approaches. Compared to previous iterations, this approach facilitated robust hypothesis validation and meaningful student reflection on AI usage. Contrary to concerns about reduced effort, most teams leveraged ChatGPT for deeper exploration and critical thinking, resulting in more confident presentations and higher-quality content. These findings reinforce evidence that explicit GenAI adoption, with guided faculty oversight, supports design-thinking strategies, enhances student engagement, and maintains academic integrity. Continued methodological refinement and further empirical research remain essential to maximize GenAI's educational benefits in challenge-based engineering education.

Full text

Research Paper Recommended citation: Garcia Huertes, S., & Bragós, R. (2025). A Hybrid Assessment Approach to AI-Enhanced Challenge-Based Learning in Engineering Education. In Kangaslampi, R., Langie, G., Järvinen, H.-M., & Nagy, B. (Eds.), SEFI 53rd Annual Conference. European Society for Engineering Education (SEFI), Tampere, Finland. DOI: 10.5281/zenodo.17631282. This Conference Paper is brought to you for open access by the 53rd Annual Conference of the European Society for Engineering Education (SEFI) at Tampere University in Tampere, Finland. This work is licensed under a Creative Commons Attribution-NonCommercial-Share Alike 4.0 International License. A HYBRID ASSESSMENT APPROACH TO AI-ENHANCED CHALLENGE-BASED LEARNING IN ENGINEERING EDUCATION S. Garcia-Huertes a,1, R. Bragós b a Telecos-BCN, Universitat Politècnica de Catalunya (UPC), Barcelona, Spain, ORCID 0000-0002-6735-3025 b Telecos-BCN, Universitat Politècnica de Catalunya (UPC), Barcelona, Spain, ORCID 0000-0002-1373-1588 Conference Key Areas: Digital tools and AI in engineering education; Building the capacity and strengthening the educational competences of engineering educators Keywords: Generative AI (GenAI); Challenge-Based Learning; Hybrid Assessment ABSTRACT This paper introduces a hybrid assessment approach integrating Generative AI (GenAI), specifically ChatGPT, into challenge-based engineering education. Students iteratively defined problem statements, brainstormed solutions, and presented their work, receiving feedback through rubric-based GenAI assessments of ChatGPT interactions and direct faculty evaluation. While ChatGPT positively influenced student performance and intellectual engagement, it amplified existing differences in motivation, prior knowledge, skills, and team dynamics, underscoring the importance of comprehensive, human-centered pedagogical approaches. Compared to previous iterations, this approach facilitated robust hypothesis validation and meaningful student reflection on AI usage. Contrary to concerns about reduced effort, most teams leveraged ChatGPT for deeper exploration and critical thinking, resulting in more confident presentations and higher-quality content. These findings reinforce evidence that explicit GenAI adoption, with guided faculty oversight, supports design-thinking strategies, enhances student engagement, and maintains academic integrity. Continued methodological refinement and further empirical research remain essential to maximize GenAI’s educational benefits in challenge-based engineering education. 1 Corresponding Author S. Garcia-Huertes [email protected] 1 INTRODUCTION The rise of Generative AI (GenAI) tools, particularly large language models such as ChatGPT, has attracted considerable scholarly attention and debate within higher education, notably in engineering education. Historically, engineering education continuously adapts to technological advancements, prompting shifts in teaching methods and learning strategies. A clear historical parallel exists with the introduction of calculators into mathematics classrooms during the late 20th century. Initially, calculators faced skepticism from educators regarding their potential negative impact on students' analytical skills. However, research by Thomas et al., (2006) shows calculators eventually gained acceptance, illustrating how initial resistance often precedes integration of transformative educational technologies. Today, engineering educators are similarly challenged by the rapid emergence of GenAI tools, necessitating proactive and open integration into educational practice. Recent studies underscore the growing consensus that, despite initial ethical and academic integrity concerns, the progressive introduction of GenAI AI into higher education curricula is both inevitable and beneficial, significantly enhancing pedagogical strategies, curriculum development, and student engagement (Santos et al., 2024; Stankovski et al., 2024). Furthermore, contemporary case studies reveal how embedding Gen AI within structured educational frameworks, such as design thinking curricula, substantially accelerates innovation processes and improves outcomes for both students and educators (Squalli Houssaini et al., 2024). Therefore, the present study proposes a clear and structured methodology explicitly encouraging faculty and students to clearly adopt GenAI tools within challenge-based engineering courses. By providing shared access to ChatGPT accounts and leveraging the analytical capabilities of these tools themselves for qualitative and quantitative assessment, we aim to enhance students' creativity and iterative design skills, while simultaneously providing faculty with a deeper understanding of student interactions. The overarching research objective addressed in this study is: "How can the transparent integration and shared faculty-student use of GenAI tools like ChatGPT effectively enhance both learning processes and assessment quality in challenge-based engineering education?" 2 METHODOLOGY The activity aims to enhance the Entrepreneurship and Innovation for World Challenges (EIWC) course within the Electronic Engineering Master at Universitat Politècnica de Catalunya (UPC). Structured as a Challenge-Based (Kohn Rådberg et al., 2020) Learning course (5 ECTS, 3 hours / week), it emphasizes societal impact, incorporating specific sessions on sustainability and ethics. Teams engage in three progressively complex challenges. The initial two-week challenge allows students to leverage familiar methods and recognize methodological gaps. In the second six-week challenge, the teaching team introduces Design-Thinking (Leifer & Meinel, 2016) and Lean-Startup (Ries, 2011) frameworks are introduced by the teaching team, each short presentation followed by a hands-on activity performed by the teams. In the third challenge, a disruptive technology, in a low phase of the TRL scale is presented by a startup or research group and the teams should propose alternative applications and business models, intended to have societal impact. The activity described in this paper took place in the second challenge, specifically in the Needfinding and Ideation phases. With the previous approach, a 3 hour session (with the corresponding autonomous preparation work in the precedent week) was devoted to Needfinding, resulting in a challenge brief identifying a specific and validated need, narrower than the original challenge and the second session to ideation, resulting in a solution concept, selected from a set of alternatives. The following sessions are devoted to the refinement of the idea and to the development of the Value Proposition and the Business Model, using the Business Model Canvas (Osterwalder & Pigneur, 2010). In the last two terms, a non-homogeneous use of GenAI by teams has been observed, primarily intended to reduce workload rather than to improve results. To clearly articulate the integration of GenAI within the second challenge, our methodology was structured across three phases: Pre-Challenge Preparation, In-Session Facilitation, and Post-Challenge Evaluation as depicted in Figure 1. Figure 1. Timeline of the different phases to prepare, execute and evaluate all the activities involved in the described methodology. 2.1 Pre-Challenge Preparation In preparation, faculty defined an intentionally broad and open-ended challenge on pedestrian safety risks associated with distracted walking in urban environments. This intentionally broad framing aimed to facilitate exploratory thinking and iterative refinement of problem definitions. Core digital infrastructure was established by creating shared ChatGPT accounts linked to generic emails accessible by both students and faculty. This facilitated transparent GenAI interactions and continuous oversight. The explicit aim was to create an environment where students could openly engage with ChatGPT, while faculty retained visibility and could offer timely formative feedback throughout the course duration to eventually analyse the logs in the evaluation period. The provided ChatGPT account was strongly recommended to facilitate consistent monitoring and comparative assessment across teams. However, its use was not enforced as mandatory, aiming to preserve student autonomy and flexibility in tool selection. This decision had implications for data consistency, occasionally complicating comparative analysis due to varied external tool use, as further addressed in the limitations section. 2.2 In-Session Facilitation Course activities comprised two main classroom sessions complemented by independent student work. The first session began with an introduction to foundational design thinking concepts, clearly differentiating the needfinding (problem exploration) and ideation (solution development) phases. Teams then used ChatGPT iteratively to refine and define their interpretation of the broad challenge, structuring interactions as complete exchanges (student prompts with ChatGPT responses), grouped into sequences called iterations for progressive refinement. The second session expanded these techniques, providing dedicated time for solution refinement. Faculty continuously monitored ChatGPT interactions, offering tailored formative feedback. Final presentations, delivered orally with slide support, were independently evaluated by three faculty members, who then consolidated their assessments into final grades. Advanced outputs (e.g., prototypes, market analyses, business models) were recognized as evidence of deeper GenAI integration aligned with course objectives. 2.3 Post-Challenge Evaluation The evaluation phase combined qualitative and quantitative analyses of GenAI usage. Qualitative evaluation was performed by leveraging one of the OpenAI models, specifically model o1, applying a detailed rubric (see Appendix) to assess iterative prompting, creativity, critical thinking, and reflective engagement from interaction logs. Quantitatively, a Python script systematically analyzed these logs, measuring metrics such as total tokens exchanged, conversation counts, and average token density per iteration, following established analytical practices from recent research (Alves & Cipriano, 2024). This hybrid assessment approach purposefully incorporated GenAI into both learning and evaluation processes, offering faculty valuable additional insights into student interactions. Furthermore, it empowered students to creatively leverage GenAI, while reinforcing critical thinking and reflecting on its usage. 3 RESULTS This section provides an in-depth analysis, presenting both quantitative metrics and qualitative evaluations of student interactions with ChatGPT, alongside detailed assessments of the final presentations delivered by each team. The analysis is structured around clear checkpoints to illustrate changes and development patterns in student engagement, critical thinking, and the complexity of their iterative interactions. Additionally, insights derived from comparative evaluation of team presentations allow for a deeper understanding of the specific contributions of GenAI tools to student performance. 3.1 ChatGPT usage Each team's usage of ChatGPT was analyzed at two checkpoints: an intermediate point (prior to the second working session) and a final point (immediately before the presentation). Table 1 summarizes both quantitative and quantitative metrics. Table 1. ChatGPT Quantitative and Qualitative Usage Assessment (Green indicates an increase; orange indicates a decrease at the final checkpoint.) Intermediate Final Team Iterations Token / iter Usage eval Iterations Token / iter Usage eval A 20 655 2,75 47 708 3,00 B 14 442 2,25 40 369 2,75 C 14 584 2,50 10 499 3,00 D 12 640 2,75 14 647 3,00 E 10 378 3,25 24 312 3,50 F 10 461 3,25 22 600 3,50 G 2 9 1,00 14 672 3,00 The quantitative assessment indicated clear improvement between the intermediate and final evaluation phases across most teams. With the exception of Team C, all teams notably increased their interactions with ChatGPT, highlighting greater familiarity and confidence in utilizing the tool for iterative problem exploration. Teams A and B particularly stood out by substantially increasing their number of iterations, reflecting intensified engagement with the GenAI tool. Meanwhile, the tokens-per-iteration metric revealed that most teams maintained consistent interaction complexity, implying sustained, thoughtful engagement. Qualitatively, teams initially utilized ChatGPT primarily for preliminary brainstorming and problem clarification. However, by the final evaluation checkpoint, interactions exhibited significant evolution, shifting towards more refined solution development. Students revisited earlier conversations to develop increasingly coherent and robust dialogues, effectively employing iterative approaches to problem-solving. This shift was evident through deeper iterative dialogues and more strategic follow-up questioning within the context of the design-thinking methodology, progressively considering user needs, preliminary prototypes, and, occasionally, foundational business model elements. Nevertheless, further enhancements in comprehensive real-world feasibility assessments and detailed market analyses remain possible. Critical thinking improved considerably across most teams, demonstrated by their growing capability to critically assess and challenge the outputs provided by ChatGPT. Teams F and A exhibited exceptional performance by systematically exploring multiple alternative scenarios, integrating external references, and explicitly addressing practical limitations of AI-generated content. Additionally, significant advancements were observed in communication clarity and overall team engagement. Teams E and F consistently delivered interactions characterized by clear, contextually rich prompts and effective tone. Notably, teams initially exhibiting lower qualitative interaction quality (specifically Teams B and G) demonstrated meaningful progress by the conclusion of the study. Team G represented an exceptional case, initially choosing a personal paid account instead of the provided ChatGPT account. Despite this initial choice, they significantly improved their interaction quality and token usage by the final evaluation. 3.2 Presentation evaluation The presentation assessments categorized teams into three distinct groups based on their demonstrated outcomes. The first group, "Developing teams" (A and E), fulfilled basic requirements but lacked specificity regarding target users and in-depth exploration of feasibility or business considerations. Their AI usage remained minimal, mostly limited to basic idea generation and problem definition. The second group, "Proficient teams" (B, C, F, and G), presented clear articulations of their challenges and solutions, effectively leveraging GenAI. However, opportunities remained to enhance technical feasibility validation, market research robustness, and detailed business modeling.The third category included a single "Advanced team" (D), which demonstrated superior depth in problem-solving, strong AI integration, and early validation of technical feasibility and market considerations. This team's presentation indicated a holistic understanding of the design thinking process, substantially enriched by effectively utilizing ChatGPT. Further analysis comparing ChatGPT usage with the final presentation evaluations, as depicted in Figure 2, yields additional insights. The figure illustrates a relatively minor variance (±10%) in usage evaluation among teams compared to their average. However, the presentation evaluations exhibit a notably greater variance, indicating more pronounced differences in teams' performance. Figure 2. ChatGPT usage vs Presentation evaluations with deviations vs the average grade for both Usage and Presentation evaluations Notably, "Developing teams" (A and E) showed significant discrepancies between their ChatGPT usage scores and their presentation evaluations, with presentations graded lower compared to their ChatGPT interactions. Conversely, the "Advanced team" (D) outperformed in their presentation compared to their ChatGPT usage assessment. These observations suggest that ChatGPT usage alone did not equalize team performance, but rather elevated the general standard of outputs, reinforcing existing differences influenced by intrinsic factors such as prior knowledge, skills, motivation, and team dynamics. Thus, while the introduction of ChatGPT positively impacted overall performance and enhanced engagement and depth in team outcomes, it was not a decisive factor in homogenizing the teams' ultimate performance levels. Instead, it amplified existing patterns in team dynamics, highlighting the importance of a comprehensive, human-centered pedagogical approach alongside technological tools. 4 DISCUSSION AND CONCLUSIONS The findings demonstrate that clear and intentional incorporation of GenAI tools like ChatGPT within a structured challenge-based learning framework notably improved student engagement and outcomes. Our results align with emerging literature, showing enhancements particularly in the process of hypothesis validation during the needfinding and ideation phases, driven by GenAI-generated personas offering deeper insights compared to traditional methods. Compared to five prior course iterations, GenAI-generated personas enabled students to more efficiently validate hypotheses during needfinding and ideation, providing deeper insights compared to traditional surveys or interviews. While engaging real stakeholders remains ideal, practical course constraints often limit such interactions. Hence, GenAI provided a pragmatic alternative to optimize validation and facilitate richer student outcomes within limited course duration. Furthermore, introducing GenAI encouraged students to explicitly declare and reflect upon their AI usage, fostering transparency, honesty, and academic integrity. Contrary to initial concerns about reduced effort due to GenAI use, our results showed that teams employed these tools not merely to lessen their workload but to achieve deeper exploration and more comprehensive outcomes. Consequently, student presentations exhibited increased confidence, clearer communication, and higher-quality content. Our study suggests that the transparent and explicit integration of GenAI tools within challenge-based engineering courses significantly benefits both student learning and faculty assessment practices. Continued methodological refinement and further empirical validation remain essential to fully harness these promising technologies in future educational contexts. 4.1 Limitations and future work Several limitations emerged during this study, indicating clear pathways for future methodological improvement. First, data inconsistencies resulted from the optional nature of ChatGPT account use, leading some teams to prefer alternative, non-monitored platforms. To address this, future implementations should mandate standardized GenAI platform usage or establish comprehensive logging systems capable of capturing interactions from all employed tools. Second, accidental deletion of interaction logs by some student teams adversely affected data completeness and accuracy. Future iterations should include explicit guidelines combined with automated backup mechanisms to ensure robust data preservation. Third, the GenAI-driven assessment methodology could be further improved through the adoption of more rigorous evaluation practices, such as majority-voting schemes or aggregated scoring across multiple GenAI-generated rubric evaluations (Garcia Huertes et al., 2024). Finally, additional research employing longitudinal or comparative experimental designs would significantly enhance our understanding of GenAI’s sustained impact on students’ critical thinking skills, problem-solving capabilities, and overall professional competence. Such studies would provide valuable empirical evidence to reinforce the pedagogical efficacy of GenAI integration in long-term educational contexts. 5 ACKNOWLEDGEMENTS This work has received support from the 2025 grant program for participation in conferences and scientific publications in the field of teaching innovation, provided by the Institute of Education Sciences at UPC. Generative AI tools were utilized for language refinement during the drafting of this manuscript. Additionally, as explicitly mentioned within the manuscript, generative AI was employed for the usage assessment of interaction logs. The authors confirm full responsibility for the accuracy, originality, and adherence to ethical and academic standards of all content presented.