scieee AI-readable full text Open interactive document viewer

Automated Plagiarism Detection in CAD Modelling Courses: Categorisation of Academic Misconduct

Yu, T.-J.; Renaud-Assemat, I.; Beier, S.; Li, D. D.

Abstract

Proficiency in creating three-dimensional (3D) models using computer-aided design (CAD) software is essential for engineering students. However, plagiarism in CADmodelling assessments is a growing concern in engineering education, as most detection tools are developed for text-based submissions. This paper introduces an automated plagiarism detection tool designed for CAD assignments on the SOLIDWORKS platform (Dassault Systèmes). Results demonstrate that the tool enhances detection performance compared to existing solutions, improving efficiency and accuracy in identifying potential cases of plagiarism. While it streamlines the detection process, human judgment remains crucial for resolving ambiguous cases. Additionally, false positives can slow the review process, and recommendations are made for refining algorithms to improve reliability and accuracy.

Full text

Practice Paper Recommended citation: Yu, T.-J., Renaud-Assemat, I., Beier, S., & Li, D. D. (2025). Automated Plagiarism Detection in CAD Modelling Courses: Categorisation of Academic Misconduct. In Kangaslampi, R., Langie, G., Järvinen, H.-M., & Nagy, B. (Eds.), SEFI 53rd Annual Conference. European Society for Engineering Education (SEFI), Tampere, Finland. DOI: 10.5281/zenodo.17631406. This Conference Paper is brought to you for open access by the 53rd Annual Conference of the European Society for Engineering Education (SEFI) at Tampere University in Tampere, Finland. This work is licensed under a Creative Commons Attribution-NonCommercial-Share Alike 4.0 International License. AUTOMATED PLAGIARISM DETECTION IN CAD-MODELLING COURSES: CATEGORISATION OF ACADEMIC MISCONDUCT T.-J. Yu a, I. Renaud-Assemat b, S. Beier c, D. D. Li d, 1 a School of Mechanical and Manufacturing Engineering, The University of New South Wales (UNSW Sydney), Sydney, Australia, https://orcid.org/0009-0003-1293-218X b School of Mechanical and Manufacturing Engineering, The University of New South Wales (UNSW Sydney), Sydney, Australia, https://orcid.org/0000-0001-9960-7950 c School of Mechanical and Manufacturing Engineering, The University of New South Wales (UNSW Sydney), Sydney, Australia, https://orcid.org/0000-0002-9815-108X d School of Mechanical and Manufacturing Engineering, The University of New South Wales (UNSW Sydney), Sydney, Australia, https://orcid.org/0000-0002-6630-5515 Conference Key Areas: Digital tools and AI in engineering education, Quality assurance and accreditation of engineering educational programs Keywords: academic integrity, CAD, plagiarism detection, SOLIDWORKS ABSTRACT Proficiency in creating three-dimensional (3D) models using computer-aided design (CAD) software is essential for engineering students. However, plagiarism in CADmodelling assessments is a growing concern in engineering education, as most detection tools are developed for text-based submissions. This paper introduces an automated plagiarism detection tool designed for CAD assignments on the SOLIDWORKS platform (Dassault Systèmes). Results demonstrate that the tool enhances detection performance compared to existing solutions, improving efficiency and accuracy in identifying potential cases of plagiarism. While it streamlines the detection process, human judgment remains crucial for resolving ambiguous cases. Additionally, false positives can slow the review process, and recommendations are made for refining algorithms to improve reliability and accuracy. 1 Corresponding Author D. D. Li [email protected] 1 INTRODUCTION Upholding academic integrity is a fundamental priority in higher education, yet maintaining it has become increasingly challenging with the transition to online assessment (Beier et al., 2022). Patterns in plagiarism are unpredictable (Levine & Pazdernik, 2018), especially in the wake of the COVID-19 pandemic (Eshet, 2024); varying across assessment types and learning environments. Computer-aided design (CAD) assessments further amplify these challenges due to the lack of dedicated plagiarism detection tools, making effective monitoring and prevention more difficult. Software-based plagiarism detection methodologies significantly improve both accuracy and efficiency compared to manual processes (Johnson, 2018; Kirstukas, 2018; Morris, 2019). While plagiarism detection tools have long been available, adapting them for non-text-based technical and engineering files presents unique challenges. Successful implementations exist in programming assignments, with tools such as JPlag (Prechelt & Malpohl, 2003) and MOSS (Measurement Of Software Similarity) (Schleimer et al., 2003). Despite these advances, there remains a significant gap in the availability of tools for detecting plagiarism in 3D CAD solid models. Proficiency in CAD is an essential skill in engineering and industrial design, as it enables the creation of precise 3D models for design and manufacture. Recent research highlights the limited availability of plagiarism detection solutions for 3D CAD models, particularly for SOLIDWORKS (Dassault Systèmes), where specialised tools remain scarce (Kirstukas, 2018; Morris, 2019). Many existing approaches exhibit notable limitations. For instance, while Johnson’s MMEM program successfully gathered feature metadata from part and assembly files, the metadata comparison was performed manually. This reliance on manual comparison involved a resource-intensive review of outputs, which could lead to potentially false negative results (Johnson, 2018). Garland Industries’ Graderworks, a CAD-grading program, also provides plagiarism detection functionalities. However, it relies solely on file-level metadata which only includes the username of the student who last saved the file, a data record that changes whenever a new user saves the file, limiting its effectiveness in gauging originality (Garland, 2023). Furthermore, Wang et al. employed feature-level metadata analysis, but their approach relied on a limited set of plagiarism flags, restricting detection sensitivity and making false positive filtering more challenging (Wang et al., 2017). Additionally, unlike Siemens NX (Morris, 2019), SOLIDWORKS files do not possess a unique file ID, which necessitates a more comprehensive examination of feature structures or properties between submissions, rather than merely comparing a singular part ID. This research aims to create an automated tool for detecting plagiarism specifically tailored for CAD-based assignments in SOLIDWORKS, a premier solid modelling CAD software extensively used in both industrial and academic settings. The tool will be designed primarily to detect instances of plagiarism through the direct exchange of partially or wholly complete files, the most common and concerning method of plagiarism. This tool aims to significantly enhance the efficiency of the assessment marking process, improve the accuracy of plagiarism detection, and reinforce academic integrity within engineering education by providing integrated functionalities for markers, including processing assembly files, automated comparisons of file metadata to evaluate the probability of plagiarism, and conducting feature-level metadata comparisons. To maximise accessibility, the tool will be made open source once planned enhancements, including greater customisability for educators, are finalised. This paper first examines the metadata structure of SOLIDWORKS files, which forms the foundation of our plagiarism detection approach. The tool identifies similarities indicative of potential misconduct by comparing metadata across student submissions. Next, the design, logic, and operational principles of the tool are detailed, explaining how it automates and streamlines the detection process. Finally, tool’s performance is evaluated using assignment submissions from a recent student cohort, demonstrating its effectiveness in identifying plagiarism efficiently and accurately. 2 METHODOLOGY 2.1 Model and feature properties in SOLIDWORKS There are two common 3D CAD ‘Model’ types that can be created in SOLIDWORKS: a ‘Part’ model and an ‘Assembly’ model. These models possess three read-only version history attributes: “Created” (date), “Last Saved” (date), and “Last Saved by” (username). The “Created” (date) attribute is read-only and cannot be altered by the user. The “Last Saved” (date), and “Last Saved by” (username) updates upon each file-save operation. A ‘Part’ model created in SOLIDWORKS comprises a collection of defining geometric ‘Features’, such as boss-extrusions, revolve-extrusions, cuts, holes, fillets, and patterns, alongside the two-dimensional sketches utilised to generate these 3D geometries. The ‘Features’ are listed in the Feature Manager sidebar of SOLIDWORKS in the sequence of their creation, as illustrated in Figure 1. Fig. 1. The feature manager sidebar and feature properties dialogue in SOLIDWORKS. Similar to ‘Models’, each of these ‘Features’ also contains attributes including its name, description, and the three most important attributes for plagiarism detection, namely, “Created By” (username), “Date Created”, and “Date Last Modified”. It is essential to emphasise that the “Created By” and “Date Created”, fields are also read-only and cannot be altered, similar to “Created” (date) at the ‘Model’ level. In addition to the geometric ‘Features’ in a ‘Part’ model, an ‘Assembly’ model contains a special type of ‘Feature’ called ‘Mates’, which are user-defined constraints between two parts. Since mates are a subset of the feature type, they share the same three version history attributes as those found in geometric features. 2.2 Program overview The tool was developed as a standalone application independent of SOLIDWORKS, capable of interacting with SOLIDWORKS, Microsoft Excel, and file systems through their corresponding application programming interfaces (API) and libraries. The tool was written in the C# language as a .NET Windows Forms application, complete with a graphical user interface (GUI), as shown in Figure 2. Fig. 2. Program GUI after files have been successfully processed. Using this GUI, the marker of the assessment submissions can intuitively register students’ names and student IDs into the program before file scanning. Each student submission, either a ‘Part’ or an ‘Assembly’, depending on the assessment, is scanned sequentially. All author usernames found in features, parts, assemblies or mates, and a selection of the earliest creation and modification dates, are stored in a comprehensive data structure in memory. After all the files are scanned, the program compares each student’s associated file metadata with other students’ metadata, from the present and past cohorts. The data of past cohorts is accessed via a master folder of JSON files, corresponding to each dataset of previously processed assessment submissions. A set of 8 plagiarism indicator flags is generated for each student, based on specific similarities between their submissions and those of other students. The first two flags are “date match” and “author match”, signifying that timestamps and/or usernames were found to match those in another student’s submission. The next flag is “date before assignment release”, which is followed by the “many authors” flag, meaning that multiple usernames were found within the file data. Next, there are three flags to classify the nature of the usernames present in the submission: “username similar to student’s real name”, “common usernames”, and “foreign username”. The final flag, “Foreign student ID”, is true when a student ID-type username that does not belong to the submitting student is found. This can be used as an indicator of highly likely plagiarism. To use these plagiarism indicator flags for effective plagiarism detection, eight distinct detection criteria were formulated, each represented by a distinct, 8-bit binary sequence, with each bit representing one plagiarism indicator flag mentioned above. By performing a bitwise comparison between a student’s 8 plagiarism indicator flags and each of the 8 plagiarism detection criteria, every student can be assigned a plagiarism risk level, ranging from 0 (indicating ‘minimal risk’ of plagiarism) to 3 (indicative of a ‘high risk’ of plagiarism). Please refer to our previous work for further details of this process (Li et al., 2024). This criteria-based approach, which relies on bitwise comparisons, allows the user to flexibly calibrate the plagiarism detection logic with every new dataset of results generated by the program. Plagiarism indicator flags and criteria can be added or removed, and the algorithms used for flag generation, such as student-to-computer name matching, can be improved or tuned according to the design of the assessment. This adaptability enables the user to reduce false positives further and enhance the identification of plagiarism cases. 2.3 Analysis of results The program automatically outputs Excel workbooks containing the essential information of every submission, including the usernames, dates, 8 indicator flags, and the plagiarism risk level, allowing human markers to review large batches of submissions efficiently. For example, with 300 submissions, this initial screening process takes approximately 15 minutes to identify false positives, clear cases of plagiarism, and submissions requiring further examination. An example of such an Excel output is shown in Figure 3. If required, the marker may further examine submissions by reviewing each submitted file’s full metadata stored by the tool in local text files or by opening the files in SOLIDWORKS. In this study, seven modes or categories of academic misconduct are introduced to offer an insight into the patterns of dishonest behaviour observed among detected students, which are: • 1 – Self-plagiarism; • 2 – Providing files or assistance to peers; • 3a – Submitting work that has been prepared either in part, or in full, with other students in the same cohort; • 3b – Submitting work that has been prepared either in part, or in full, with students not in the course, such as past students of the course; • 3c – Submitting work that has been downloaded from an online CAD file library; • 4 – Contract cheating from an external provider; • 5 – Recreating another student's part without copying. Fig. 3. An example of a typical summary report generated by the program. Through a thorough examination of submitted files and metadata, the behaviour of each student involved in academic misconduct is placed into one of these categories. 3 RESULTS A total of 296 submissions were received from 299 students enrolled in the 2024 Term 3 cohort of the MMAN1130 Design and Manufacturing course at UNSW. The assignment required each student to submit one assembly file containing six custom parts and a few standard bearings and fasteners, resulting in a typical total of 6 to 10 files per student, which were stored in zipped folders. The submissions downloaded from our Learning Management System (LMS) were directly processed by the program automatically, which unzipped files from all submissions and analysed a total of 1612 ‘Parts’, within 301 ‘Assemblies’, in 20 First Name Surname ID Author Names Earliest Date Date Match Author Match Old Date Many Authors Usernames like student name Common usernames Foreign Username Foreign Student IDs Colluder IDs Risk Status Thel Vadam 2552 Arbiter 12.10.2024 13:27 TRUE TRUE John Halo 0117 john117 17.10.2024 14:06 TRUE TRUE TRUE 1234[A][D] 3 Jacob K. C709 jacobk 10.10.2024 23:07 TRUE Avery J. 1234 sarge, john117 17.10.2024 14:06 TRUE TRUE TRUE TRUE 0117[A][D] 3 Jane Citizen 2009 janec 09.10.2024 09:00 TRUE minutes and 48 seconds, including the 16 seconds needed for SOLIDWORKS to launch in the background. The extra assemblies are explained by students mistakenly submitting multiple assembly files. The subsequent creation of a JSON file for raw data storage and summary reports in Excel format took the program around 4 seconds, with the JSON file encompassing all submission data, totalling 667 kilobytes (kB) in size. A summary of the number of plagiarism cases flagged by the program is detailed in Table 1. Table 1. Summary of verdicts on detected cases of plagiarism. Plagiarism risk level (total number of detected cases by program) Number of students Confirmed plagiarism Likely false positive Clear false positive Level 3: “high risk” (18 cases) 11 1 6 Level 2: “moderate risk” (22 cases) 6 3 13 Level 1: “low risk” (2 cases) 2 0 0 Level 0: “minimal risk” (2 students) 2 - - In total, 42 cases were flagged by the program with a plagiarism probability level of 1 or higher, of which 19 were confirmed as clear false positives, primarily due to issues arising from the file author name comparison processes, such as the inability to recognise a username as being similar to the student’s actual name (e.g. Thomas vs. tom_pc) or cases involving students with non-common computer usernames unrelated to their real name (e.g. John vs. masterchief). Four cases warranted more detailed manual examination (lasting approximately 20-30 minutes), which determined them to be false positives as well, and as such they have been classified as "likely false positives". In addition to the remaining 19 cases, two more submissions were confirmed as plagiarism. Although designated as 'minimal risk' (level 0) based on their own file characteristics, the program flagged them for metadata matches with higher-risk submissions, leading to their misconduct being confirmed, making for a total of 21 confirmed cases of plagiarism. These cases were then categorised by the earlier discussed “mode of academic misconduct” involved. Of the 21 confirmed cases, eight were classified as Mode 2, involving providing files or other assistance to fellow students. Nine were Mode 3a, where students submitted work that included contributions from others. Two cases were Mode 3b, involving using files or features created by students from previous cohorts. One case was Mode 4, which involved contract cheating, and one final case was Mode 5, where a student recreated another student's work without directly copying it. The results are summarised in Table 2, along with details of how the program or marker identified the cases. 4 DISCUSSION This innovative plagiarism detection tool for SOLIDWORKS CAD models provides an efficient approach to identifying academic misconduct. The tool enabled a single marker to complete the detection process for 296 submissions within one hour – a significant improvement over previous manual methods, which required a team of four markers working for eight hours to complete the same task. The previous manual approach was limited to inspecting the ‘Last Saved By’ username at the ‘Model’ level, making it time-consuming and ineffective at detecting subtle forms of misconduct. In contrast, this automated tool identified plagiarism through additional indicators, such as matching timestamps and usernames at the ‘Feature’ level, even when students used generic computer names – patterns that could easily be overlooked in a manual review. As such, since the program replicates all the key steps of a manual review, while performing them with significantly greater thoroughness and accuracy, it is unlikely to miss instances of plagiarism that a manual review would detect, except in cases involving human error during the final processing stage. Table 2. Details of all 21 confirmed cases of academic misconduct for this assignment, sorted by their program-assigned plagiarism risk levels. Student No. Risk level Mode of Misconduct Description 1 3 3a Username and timestamp match: Submitted files shared by Student 18. 2 3 3a Username and timestamp match: Submitted files shared by Student 19 3 3 2 Shared all part files, fully assembled, with Student 4. 4 3 3a Username and timestamp match: Submitted files shared by Student 3 and replaced some shared part files with their own. 5 3 3a Username and timestamp match: shared files with and submitted files made by Student 6. 6 3 3a Username and timestamp match: shared files with and submitted files made by Student 5. 7 3 3a Username and timestamp match: Submitted files shared by Student 8, with added assembly mates. 8 3 2 Shared all part files, fully assembled, with Student 7. 9 3 3a Username and timestamp match: Submitted files shared by Student 10, with added assembly mates. 10 3 2 Shared all part files, fully assembled, with Student 9. 11 3 3a One of the submitted parts was created by Student 20. 12 2 3b Username match: Assembly mates done by past student. 13 2 4 All features and assembly mates done by one author, whose username matches a CAD freelancer with a LinkedIn presence. 14 2 2 Username match: created the same part twice and sent one to Student 15, to avoid detection. Created assembly mates for Student 15. 15 2 3a See Student 14. 16 2 2 Suspected of sending files to Student 21 to act as a reference file. 17 2 3b Username match: All parts and features made by past student, while assembly mates are original. 18 1 2 Shared all part files, fully assembled, with Student 1. 19 1 2 Shared all part files, fully assembled, with Student 2. 20 0 2 Username match: created the same part twice and sent one to Student 11, to avoid detection. 21 0 5 Single Username match: Student 16's username was found in the assembly properties. This led to a closer examination of design decisions, which determined that although no files were copied, the design decisions within the two students' submissions were too similar to be independent work. The results of this study highlight several significant advancements over existing approaches, effectively addressing limitations in prior literature. For example, unlike Johnson’s MMEM, which relied on manual metadata comparison (Johnson, 2018), our tool fully automates this process, eliminating labour-intensive review steps and enhancing scalability across large student cohorts. Furthermore, Garland’s Graderworks tool, though capable of detecting plagiarism, relied only on file-level metadata (Garland, 2023), which could be easily altered when a different user saves the file. In contrast, our method effectively addresses this limitation by scanning and storing all critical metadata to the ‘Feature’ level. Finally, while the methodology in Wang et al. utilises feature-level metadata collection, their analysis involved fewer identified plagiarism flags compared to our method, which complicates the filtering of false positives (Wang et al., 2017). Our approach addresses this issue by relying on more plagiarism flags to improve detection rates and reduce false positives. Human evaluation of the program’s outputs remains essential, as 45% of the identified submissions were clear false positives. This aligns with Kirstukas’ findings, where 64% of detected cases were false positives (Kirstukas, 2018). Most misclassifications stemmed from challenges in categorising file author usernames. Since a marker’s investigation is still required to confirm cases of plagiarism, students cannot be wrongly incriminated without human judgment. There is an inherent trade-off between minimising false positives and maintaining sufficient sensitivity to detect actual misconduct. The system’s configuration should reflect the user’s tolerance for risk, balancing accuracy and coverage. Future program versions could incorporate advanced name-matching algorithms, such as fuzzy matching, to improve the accuracy of associating student names and usernames. The current program is designed to detect instances of plagiarism primarily through the direct exchange of partially or wholly complete files. It does not account for cases where students independently reproduce another student’s work without direct copying or file exchange. A solution could involve analysing the sequence and types of features utilised in creating CAD models and the methodologies proposed herein to enhance detection capabilities. Nevertheless, it is essential to note that students imitating one another’s CAD work is a learning experience, and this mode of plagiarism is less significant than other forms. 5 CONCLUSION A comprehensive plagiarism detection tool for CAD-based assignment submissions has been effectively developed and implemented to identify potential academic misconduct for further investigation promptly. The program can efficiently process a significant volume of student submissions with enhanced accuracy, thoroughness, and resource management, significantly reducing the demands on educators' time and resources. The tool provides a structured approach to initial plagiarism identification by utilising metadata analysis. However, human assessment remains essential to filter out false positives and to assess the nature and severity of any detected plagiarism, following the University’s plagiarism policies. The development and implementation of the tool aim to support educators in upholding academic integrity and safeguarding the reputation of educational institutions amidst constantly changing student attitudes towards plagiarism; specifically, collusion within individual assessments. The program can be refined based on user feedback to enhance its effectiveness further.