Practice Paper Recommended citation: Li, C., Almeida, C., Matthews, S., & Eskelinen, H. (2025). Peer Review Versus Manual Assessment in a Multicultural 3D Modeling Course. In Kangaslampi, R., Langie, G., Järvinen, H.-M., & Nagy, B. (Eds.), SEFI 53rd Annual Conference. European Society for Engineering Education (SEFI), Tampere, Finland. DOI: 10.5281/zenodo.17631325. This Conference Paper is brought to you for open access by the 53rd Annual Conference of the European Society for Engineering Education (SEFI) at Tampere University in Tampere, Finland. This work is licensed under a Creative Commons Attribution-NonCommercial-Share Alike 4.0 International License.
PEER REVIEW VERSUS MANUAL ASSESSMENT IN A MULTICULTURAL 3D MODELING COURSE C. Li a, 1 , C. N. Almeidab, S. Matthews c, H. Eskelinen d a LUT University, Lappeenranta, Finland, ORCID 0000-0001-9400-3998 b LUT University, Lappeenranta, Finland, ORCID 0000-0002-1832-5677 c LUT University, Lappeenranta, Finland d LUT University, Lappeenranta, Finland Conference Key Areas: Diversity, equity and inclusion in our universities and in our teaching; Engineering ethics education Keywords: Peer review; large-scale courses; manual evaluation; international students; engineering education ABSTRACT This paper examines the implementation and effectiveness of peer review evaluation compared to manual assessment in a large-scale 3D modeling course, offered to international and local students at a European university. The study compares student performance data across two academic years, with peer evaluation in 2023 and manual assessment in 2024 for the same course. Results demonstrate that peer review evaluation can be an effective assessment method for large-scale engineering courses, reducing instructor workload from approximately 30 hours to 1 hour per exercise while maintaining educational effectiveness. Additionally, an interesting phenomenon was observed: international students seemed more motivated to pursue higher grades compared to local students. Their performance followed a similar trend, regardless of the evaluation method used. Finally, to better understand this difference, informal discussions were held with students and teachers. These conversations indicated that international students face more challenges after graduation, which may explain their motivation to achieve higher academic results. 1 Corresponding Author C Li
[email protected]
1 INTRODUCTION Large-scale university courses and massive open online courses (MOOCs), which typically operate without a direct instructor involvement, present significant challenges for educators in selecting appropriate evaluation methods. Peer review evaluation, already well established in academic paper publishing, offers a potential solution for managing assessments in high-enrollment courses. In paper publishing process, reviewers provide feedback to authors, who revise their work accordingly. Then revised manuscript is usually re-evaluated and if the reviewers are satisfied, it is accepted for publication. Similarly, peer review has also been adopted in university courses for various reasons, such as encouraging students to think critically, learning from peers, or for MOOC courses (Brkić et al., 2024). In such cases, the teacher takes on a more observational role, overseeing the process rather than directly assessing each student. This can help reduce the teacher’s workload, allowing more time to focus on tasks like developing the course content. In courses where students need to use a software to model something, it can be technically challenging for the teacher to provide written feedback, especially in large-scale courses (Van et al., 2023). Peer review evaluation can support this to some extent, but it requires that both teachers and students are familiar with the necessary technology tools (Badea & Popescu, 2019). For example, teachers may provide a detailed evaluation matrix, and students then review each other’s work based on that matrix to ensure the quality and consistency of the review (Li et al., 2010). Students can learn not only from receiving peer feedback but also from the process of evaluating others’ work (Nicol et al., 2013). However, peer review does not always improve the overall experience compared to traditional assessment methods, and its reliability is questionable. Still, it offers advantages in other areas (Chambers et al., 2014). Students are usually positive when they start providing peer review but may have different opinions when they receive feedback from peers, and sometimes their attitudes might become polarized (Mulder et al., 2014). Studies generally indicate that students want to participate in the assessment process and view peer assessment positively (Wanner & Palmer, 2018; Nieminen et al., 2021). Therefore, it makes sense to include them in the assessment criteria. Furthermore, peer review stands out as a valuable teaching methodology that not only supports student learning but also promotes critical thinking. A study by Guelfi et al. (2021) found moderate agreement between peer-assigned and teacherassigned grades. By implementing peer review systems, instructors can apply this strategy effectively in large classes, reducing their workload while actively engaging students in the assessment process. Partanen et al. (2023) highlight that the fast-evolving nature of engineering requires graduates to become lifelong learners, with an emphasis on deep learning strategies and self-regulation. To support this development, educators are encouraged to implement teaching and assessment methods that cultivate these competencies, including peer and self-assessment. This paper presents a comparative case study examining differences observed when peer review and manual teacher assessments are applied to the same course exercises across consecutive academic years. In 2023, peer review was used to assess selected exercises, while in 2024, the same exercises were evaluated
manually by a single teacher. The study aims to provide practical insights for engineering educators regarding evaluation method selection in large-scale courses. The paper is structured as follows: Chapter 2 introduces the course implementation; Chapter 3 presents student performance data and the teachers’ insights; Chapter 4 provides concluding remarks and implications for practice. 2 COURSE IMPLEMENTATION 2.1 Course description The course “3D Modeling and Technical Documentation” is a mandatory 6-credit, year-long course designed for all first-year engineering students, including both international and local students, at a University in Europe. The course consistently receives a high number of enrollments each year, reaching 600 students in 2023. Upon completion, students should be able to use 3D modeling software to create complex 3D models and 2D drawings, as well as to produce the corresponding technical documentation. The learning objectives include, for 3D modeling: designing, modeling, assembling components using different mates, and creating basic animations; and for 2D drawing: demonstrating an understanding of general tolerances, applicable standards, geometry tolerancing, engineering fit, surface roughness, welding symbols, and the ability to make assembly drawings. The course, in its current form, was first implemented in the academic year 2021. As student enrollment increased significantly, the teachers began to consider the use of peer review evaluation for certain exercises. The goal was to reduce the teachers’ workload without compromising the quality of the course. The grading scheme of the course is presented in Table 1, with peer review evaluation applied for individual exercises 1-3. Table 1. The grading scheme of the course. Assessment task Description Share of the final grade Individual exercises 1-3 Students finish the task using software SolidWorks 30% Individual exercise 4 Students finish the quiz-type question 5% Group exercise 1st phase: make 3D modeling 2nd phase: make 2D drawing 35% Oral examination Group oral discussion with a teacher after group exercise 30% The aim of individual exercises 1-3 is to encourage students to create 3D models based on a given topic. Detailed step-by-step tutorials are provided to guide them, and students who demonstrate innovation in their work can earn additional points. These exercises do not have a unique solution; the only strict requirement is that students submit their exercises correctly. The level of difficulty increases progressively from the first to the third exercise. As the second part of the course focuses on technical documentation in 2D drawing, which may be perceived as less interesting than 3D modeling, it is important to motivate students at the beginning of the course. Therefore, teachers adopt a more flexible approach when grading individual exercises 1-3.
Due to the large number of students, teaching assistants, recruited from the previous year’s students, were hired to provide technical support during the exercise session. Their role is limited to assisting with modeling techniques and helping students to understand peer feedback. Since teaching assistants vary slightly in their technical expertise, they are not involved in evaluating the exercises. In the peer review process, randomized peer review groups are assigned for each assignment to reduce bias and expose students to different perspectives. Supporting student well-being is also a priority: clear communication of deadlines and expectations helps minimize stress and prevent burnout. Feedback is designed to highlight strengths and provide constructive suggestions for improvement. 2.2 Evaluation methods In the academic year 2023-2024, the peer-review assessment was used to evaluate individual exercises 1 to 3. This was implemented using the integrated “Workshop” tool in Moodle. In practice, once the submission deadline is closed, the system automatically assigns each student with three random exercises for evaluation. Students then assess these three exercises using the evaluation matrix integrated into the system, which was designed by the teacher. Because the assessment is subjective, significant variation can occur if the evaluation criteria are too general. To address this, the evaluation questions were designed as presented in column 2 in Table 2. Table 2. Assessment Criteria and Percentage Distribution. Exercise Peer Review Assessment (Moodle Workshop) Teacher Assessment (Single Instructor) 1 •Correct file submission with basic model: 57% • Drawing cleanliness: 14% • Innovation: 29% • Correct basic model: 60% • Innovation: 40% 2 • Correct file submission with basic model and assembly functionality: 29% • Drawing clarity & dimensions: 57% • Innovation: 14% • Correct basic model: 60% • Innovation: 40% 3 • Correct file submission with basic model and component mating: 33% • Technical documentation: 33% • Innovation: 33% • Correct basic model: 67% • Innovation: 33% The assessment criteria offer students flexibility and help reduce the risk of blind evaluation. Since the final score is based on the average of three assessments, the potential for individual bias is expected to be relatively low. During the evaluation process, students not only gain insight by reviewing their peers’ work but also benefit from reading the feedback they receive. In the academic year 2024-2025, the assessment of individual exercises 1 to 3 was realized manually by a teacher. All the exercises were evaluated by the same teacher to ensure consistency and fairness throughout the process. The evaluation
matrix is presented in column 3 of Table 2 and it was provided in the instructions before students began their exercises, allowing them to self-assess their work before receiving a grade. Both assessment methods prioritize core engineering skills, value innovation, maintain consistent professional standards, and are structured to reflect increasing task complexity. To reduce workload and simplify communication, especially when students send emails to multiple teachers, teachers limit responses by email. If students have questions about their final score, they are encouraged to contact the teachers during office hours or attend the exercise sessions. These interactions allow teachers to engage more directly with students, and when students attend the exercise sessions, they usually learn more. 3 RESULTS AND INSIGHTS 3.1 Student’s performance The points for individual exercises 1 to 3 were exported from Moodle and analyzed. Figure 1(a) and 1(b) present bar charts showing the average values and standard deviation distributions for each exercise and the total score across all three exercises. The x-axis represents the item, and the y-axis shows the corresponding point values. Different colors indicate values from different academic years. Fig 1. (a) the average value (b) standard deviation distribution of each exercise and the sum of these exercises. Figure 1 reveals two key observations: the first observation is that the trend of performance from individual exercises 1 to 3 is the same, which is decreasing. This can be reflected in the difficulties of the exercises, and students’ motivation may also have slightly dropped during the course. The second observation is that the total points obtained from these three individual exercises in 2024 are lower than in 2023, with a drop of about 33%. Overall, this result was expected by teachers, because the exercises were adjusted a bit in 2024; if students wanted to get higher points, they needed to be a little bit more creative, instead of blindly following the tutorial. From a pedagogical perspective, the observed results show differences between the peer-reviewed year 2023 and the manually assessed year 2024. Data indicates that manual assessment in 2024 led to lower average scores but a higher standard deviation, partly because exercises required more creativity from students. It seems (a) (b)
that manual assessment by a single teacher allows for finer grading and capturing details like creativity, which the peer-review system might average out across different student reviewers. While peer review in 2023 was functional, reduced teacher workload, and showed less score variation, it differs from the manual evaluation which might be more discriminating but takes more time. This presents a pedagogical trade-off: the efficiency and student learning from peer review versus the potentially more consistent and detailed grading from manual assessment by one expert. It is important to note that, when considering the overall course scores, there is little difference between peer review evaluation and teachers assessment method. Then for each exercise in each academic year, the points of each exercise are divided into six intervals: 1) 0 points; 2) 1-2 points; 3) 3-4 points; 4) 5-6 points; 5) 7-8 points; 6) 9-10 points. All students are divided into two groups, local students and international students. Their performances for exercises 1, 2, and 3 in academic years 2023 and 2024 are illustrated with line charts and shown in Fig. 2 (a)-(f). The results shown in Figure 2 are collected after observing that international students tend to achieve higher grades compared than local students in the course. From the line charts, it can be concluded that most local students got points in the range of 5 to 8, while most international students got points in the range of 9 to 10. When focusing on only one specific group, it can be concluded that the international students’ performance is in the same trend no matter which assessment method is used. However, as for local students, there is an obvious drop in the frequency of high points. Based on the informal discussions with students and teachers, three main reasons may explain the phenomenon. First, international students usually pay tuition fees, and scholarships are often based on performance in courses while for the local students the education is free. Second, international students often seek high grades for applications to other universities or countries, whereas local students usually continue their master’s studies in the same university without grade requirements. Third, international students may have fewer social connections in a foreign country compared to local students; if international students want to invest in themselves for the future, studying and getting higher grades is one of the options that are left. As for local students, due to cultural differences, they may realize that if they are going to work in a company after graduation, grades are not the only aspect that a company will consider. Sometimes, a grade of 3.5 might be even better than a grade of 5, because it may imply that students have a good social life and social skills, that students are creative and do not always follow strict instructions given by the teachers, which are seen as merits by some companies. However, international students form a very diverse group, including individuals from many different backgrounds, and further research is needed to better understand this behavior. 3.2 Teacher’s Insights In the academic year 2023-2024, peer review evaluation was used for individual exercises 1 to 3. In contrast, in the academic year 2024-2025, these exercises were evaluated manually by a single teacher to ensure fairness and quality. During the course, teachers wanted to assess two main aspects: time and quality.
Fig 2. Distribution of students’ performances of exercise 1 in 2023 (a) and 2024 (b), exercise 2 in 2023 (c) and 2024 (d), and exercise 3 in 2023 (e) and 2024 (f). Regarding time requirements, evaluating each individual exercise takes approximately three minutes per student, the total time spent on each exercise is about 30 hours. When peer review is used, teachers only check a sample of submissions, and the total time spent per exercise is reduced to around 1 hour. Only a small number of grades, around 3%, were adjusted across all individual exercises, when peer review was used. The most common reason for these adjustments was the absence of peer evaluations, which required the teacher to evaluate the exercise and assign a grade. The observation that international students are more grade-oriented and motivated is supported by educational research addressing whether this stems from better preparation or higher motivation. While prior preparation is a factor, studies confirm a "self-selection effect" where international students often exhibit higher selfdetermined motivation and a stronger grade-orientation than domestic peers (Chue (a) (b) (c) (d) (e) (f)
& Nie, 2016). In terms of quality, manual assessment tends to be more beneficial for both students and teachers. For students, grades are perceived as more reliable when given by teachers rather than peers. For teachers, manually reviewing the exercises allows them to identify the common mistakes, which can help in improving course materials immediately or for the next academic year. During discussions, teachers feel more confident, are better prepared to answer questions and can provide more effective guidance. On the other hand, peer review enables students to learn from evaluating others’ work and recognize their own mistakes. 4 CONCLUSIONS Analysis of quantitative data and qualitative feedback indicates that peer review is a viable assessment method, demonstrating minimal impact on students' final grades despite potential trade-offs in assessment fidelity. Data analysis confirms that peer review significantly reduces teacher time requirements (approximately 1hrs vs. 30hrs per exercise) compared to manual grading. With the results collected in the past two academic years, teachers are confident that peer-review assessment can be used, and they are thinking of rotating different evaluation methods in different academic years to ensure course quality maintenance. The comparison showed differences between the methods. Manual assessment in 2024 enabled finer grading and the capture of details like creativity, leading to lower average scores but higher variation. Peer review in 2023 was efficient and showed less score variation, presenting a trade-off between efficiency/student learning and detailed grading consistency. Manual grading also provides direct feedback to instructors about common mistakes, helping to improve course materials. An interesting and persistent phenomenon was noticed: local students are less motivated to pursue higher grades compared to international students, regardless of the assessment method. From the line charts and discussions, it seems international students aim for top scores possibly due to factors like fees, scholarships, or mobility goals, while local students' performance clusters in the middle ranges, perhaps reflecting different priorities. However, further research is needed to better understand this behavior. Since the exercises involve 3D modelling rather than mathematics or physics calculation, it is difficult for a system to evaluate then automatically. There are some studies about this topic by Sanna et al., (2012), but implementing them requires some specific skills from teachers. However, the teachers are trying to find a reliable solution to see if there is any artificial intelligence tool that can be used in the evaluation; for example, a certain number of exercises are evaluated in detail by teachers first; then these data will be used in training; and the trained model can be used to evaluate the rest of the submissions. The findings offer practical guidance for engineering educators managing large-scale courses, while also highlighting important considerations for diverse student populations in international educational settings. Future research should investigate the optimal combination of peer review and manual assessment to maximize both efficiency and educational effectiveness.