Full text
RDM 101 Impact Evaluation A Study on Training Effectiveness and Research Practice Changes Authors: Narmin Rzayeva, Francesca Morselli, Carla Strubbia, Nikki Grens, Gargi Kulkarni, Paula Martinez Lavanchy Date: 01.12.25
RDM 101 Impact Evaluation A Study on Training Effectiveness and Research Practice Changes
RDM 101 Impact Evaluation 01 December 2025 4 Contents 1 Executive Summary 6 1.1 Overview 6 1.2 Objectives 6 1.3 Methodology 6 1.4 Key Findings 6 1.4.1 Short-Term Impact 6 1.4.2 Long-Term Impact 7 1.5 Areas for Improvement 7 1.6 Recommendations 7 2 Introduction 8 2.1 About RDM 101 8 2.2 Study Objectives 8 3 Data Source & Methodology 9 4 Short-term impact of RDM 101 10 4.1 Graduate School Feedback Forms 10 4.1.1 Objectives of the Analysis 10 4.1.2 Graduate school Feedback Forms Overview & Structure 10 4.1.3 Methodology & Software 11 4.1.4 Results 12 4.2 Preand Post-training Survey 23 4.2.1 Objectives of the Analysis 23 4.2.2 Survey Overview & Structure 23 4.2.3 Methodology & Software 24 4.2.4 Results 25 5 Long-term impact of RDM 101 38 5.1 Objectives of the Analysis 38 5.2 Interview Overview & Structure 38 5.3 Methodology & Software 39 5.4 Results 40 6 Main Highlights of the Study 45 6.1 Short-Term Impact Summary 45 6.1.1 Course Perception and Engagement 45 6.1.2 Feedback-Driven Improvements 45 6.1.3 Training Effectiveness and Shifts in RDM Practices 45 6.1.4 Online vs In-Person Delivery 46 6.2 Long-Term Impact Summary 47 6.2.1 Motivations and Expectations 47 6.2.2 Knowledge Acquisition and Impact on Practice 47 6.2.3 Course Design and Delivery 47
RDM 101 Impact Evaluation 01 December 2025 5 6.2.4 Emotional and Social Dimensions 47 7 Conclusion 48 Data and Code Availability 49 References 50 Acknowledgments 51 Author Contributions 52
RDM 101 Impact Evaluation 01 December 2025 6 1 Executive Summary 1.1 Overview The RDM 101 course (Martinez-Lavanchy et al., 2024) at TU Delft is a foundational course designed to equip PhD candidates with essential research data management (RDM) skills. This impact study evaluates the course’s effectiveness in achieving its learning objectives and influencing research practices, both in the short and long-term. 1.2 Objectives The study aimed to: 1. Assess short-term knowledge and skill acquisition. 2. Evaluate long-term changes in research practices. 3. Inform improvements to course design and delivery. 1.3 Methodology A mixed-methods approach was used, including: • Analysis of Graduate School feedback forms (2020–2025). • Preand post-training surveys (2024/2025). • Semi-structured interviews with past participants and data stewards. 1.4 Key Findings 1.4.1 Short-Term Impact • High Satisfaction: Participants consistently rated the course highly for content, trainer support, and relevance. • Increased Awareness: Awareness of institutional RDM support rose from 45% to 98% posttraining. • Improved Practices after the course: o 90% changed their data storage strategies. o 81% implemented new backup strategies. o Use of secure institutional storage (e.g., Project Data U: drive) increased significantly. • Enhanced Understanding: o Confidence in RDM skills and familiarity with FAIR principles improved markedly. o 98% felt better equipped to develop or improve their Data Management Plans (DMPs), with the Data Flow Map (DFM) assignment cited as a particularly valuable component.
RDM 101 Impact Evaluation 01 December 2025 7 1.4.2 Long-Term Impact • Sustained Behavioural Change: Participants reported continued use of RDM practices up to 18 months post-training. • Positive Shift in Attitudes: Participants moved from viewing RDM as a compliance task to recognizing its value in research quality and efficiency. • Mindset Shift: Many described a transformation in how they approach data management, emphasizing organization, documentation, and reproducibility. • Emotional and Social Benefits: The course reduced anxiety around RDM and fostered a sense of community and peer learning. • Trainer Influence: Trainers were praised for their approachability and support, contributing significantly to the learning experience. 1.5 Areas for Improvement • Workload vs. Credits: Participants felt the workload was high relative to the credits awarded. • Discipline-Specific Content: There is demand for more tailored examples and case studies. • Delivery Format: While in-person class sessions were preferred for engagement, online and hybrid formats improved accessibility. 1.6 Recommendations • Introduce modular, discipline-specific learning paths. • Rebalance workload and credit allocation. • Keep the blended format of the course, and work on enhancing online delivery with interactive tools and structured sessions. • Investing in expert-led training is key to embedding RDM practices in research at TU Delft. • Continue investing resources on course evaluation to keep it relevant for PhD candidates.
RDM 101 Impact Evaluation 01 December 2025 8 2 Introduction 2.1 About RDM 101 Research Data Management 101 (RDM 101) (Martinez-Lavanchy et al., 2024) is a basic course offered to PhD candidates as part of the Doctoral Education programme at TU Delft. The course aims to equip researchers with essential knowledge and practical skills for managing research data and other relevant research objects in accordance with best practices. Below are the main learning objectives of the course: • realise the important role that good data management plays in research • identify different types of research data and recognize the regulations, policies and/or legal requirements associated with them • list the main components of the FAIR data principles and connect them to your own research workflows • employ the acquired knowledge to design an efficient research data management strategy for your projects according to best practices. The course is delivered in a blended format, combining self-paced online modules in Brightspace LMS with three in-person or online (via Zoom) class sessions. It is taught by four RDM experts from the TU Delft Library's Research Data Services (RDS) training team. The online content includes videos, interactive images, quizzes, and opportunities for peer and trainer interaction through discussion forums. Over the three-week duration of the course, participants use their own research data as case studies to gradually develop a Data Flow Map (DFM) for their projects. This map can later serve as the basis for their project's Data Management Plan (DMP). 2.2 Study Objectives The training team places high value on providing extensive feedback and delivering face-to-face, interactive sessions, as this approach is believed to support the long-term development of data management skills. However, this format is timeand resource-intensive, particularly due to the continuous feedback and detailed review of practical assignments it requires. This often raises the question: does the current course design and time investment of trainers effectively achieve the intended outcomes and impact? To address this, the training team conducted an evaluation of the course, focusing on both its shortterm impact, meaning what knowledge and skills participants gain immediately after the training, and its long-term effects - what changes in RDM practices the training helps researchers to bring about. To achieve this broader aim, three main goals were identified: . Goal 1 (short-term impact). To evaluate the extent to which participants have gained knowledge and developed skills related to RDM best practices. Goal 2 (long-term impact). To explore how participants apply RDM principles to their daily research tasks, such as data collection, analysis, and interpretation. This could involve assessing changes in data organization methods, metadata completeness, documentation quality, version control adherence, data sharing behaviours, etc. Goal 3. To inform improvements to the course design, delivery format, and content, by identifying areas where the course could better align with participants' needs and expectations, and ensuring its continued relevance and effectiveness.
RDM 101 Impact Evaluation 01 December 2025 9 3 Data Source & Methodology The short-term impact of the RDM 101 course was evaluated using two data sources. First, data collected through Graduate School feedback forms from 2020 onwards was analysed. These forms provided insights into participants' immediate reactions to the course content and delivery. Second, for this study a series of preand post-training surveys were conducted across four training sessions during the 2024/2025 academic year. These surveys aimed to measure changes in participants’ knowledge, attitudes, and confidence related to research data management as a result of the training. To assess the long-term impact of the RDM 101 course, interviews were conducted with 11 participants who had completed the course 12 to 18 months earlier. These interviews explored how the training influenced their research practices over time. In addition, two data stewards were interviewed to gather their perspectives on any observable changes in the course participants’ approach to research data management. As seen, we employed a mixed-methods approach, combining both quantitative and qualitative data sources. The figure below summarizes the main data sources used in this study, including feedback forms, surveys, and interviews with PhD candidates and data stewards. Figure 1. Overview of qualitative and quantitative data sources used in the study. To perform this study, we submitted a Human Research Ethics Committee (HREC) application, which included the Data Management Plan (DMP) of the project. The preparation of documentation and the HREC application process was guided by the Data Steward at the Faculty of Technology, Policy and Management. Due to ethical and privacy considerations, the raw research data containing potentially confidential and identifiable personal information are not publicly available and are securely stored on a university network drive. To support transparency and reproducibility of the study, the anonymised quantitative datasets and the high-level qualitative categories, together with their conceptual descriptions and containing no personally identifiable information about participants, are published in the 4TU.ResearchData repository (Rzayeva et al., 2025). The scripts used for the analysis are openly available in the 4TU.ResearchData repository (Grens, 2025) and are linked to a public GitHub repository to ensure the reproducibility of the analysis. For more information, please see the Data and Code Availability section of the report.
RDM 101 Impact Evaluation 01 December 2025 16 Figure 5. Course engagement level rating by participants. When comparing 2024 with 2023 data, we observe that the average engagement level increased for a much larger number of participants, and the deviation decreased. The data from one run in 2025 can also serve as an initial supporting indicator of this trend. We believe this suggests that the recent modifications to the training delivery design and format may have had a positive effect on participant engagement. Feedback on Trainers Participants also evaluated the trainers from various perspectives. For instance, they were asked to rate their level of agreement with the statement “I felt safe to share personal experiences/opinions” on a scale from 1 (fully disagree) to 5 (fully agree). Across all years, responses were highly positive, with the vast majority selecting 5. In 2020, the average score was 4.8; in 2021, it rose slightly to 4.87. Similarly high levels of agreement continued in subsequent years: 4.91 in 2022, 4.58 in 2023, 4.57 in 2024, and again 4.57 in 2025. This is a particularly relevant indicator for us, as one of our main goals is to foster the exchange of knowledge and best practices between peers and trainers. We analysed their responses from 2020 regarding the trainers’ ability to present course materials in a clear and structured way (Figure 6), as well as the clarity and value of feedback provided by trainers (Figure 7). This data was available only for the recent years.
RDM 101 Impact Evaluation 01 December 2025 17 Figure 6. Participants’ agreement level with the clarity and structure of the trainers' material presentation.
RDM 101 Impact Evaluation 01 December 2025 18 Figure 7. Participants’ agreement level with the clear and valuable feedback provided by trainers. As observed in the figures above, the vast majority of participants agreed or fully agreed that the trainers performed well across the given aspects in all observed years. This suggests that the course setup—including class sessions with trainer’s explanations and interactions, feedback provided during sessions and on assignments, as well as other interactions with trainers via Brightspace is of high value to participants. However, the slight decrease in average scores observed in 2023 and 2024 may indicate the effect of the training team still “landing and gaining experience”. It also suggests that there is room for improvement for the training team, possibly through exploring new delivery methods and pedagogical approaches, feedback practices, as well as testing new session dynamics to further enhance the participant experience. Course relevance 93% of participants agreed or strongly agreed that the skills provided in the course are relevant to their PhD project and/or trajectory, and 87% reported that for their future career. The data for these questions is available only for recent years.
RDM 101 Impact Evaluation 01 December 2025 19 Figure 8. Participants’ agreement level with the relevance of the skills provided by the course to PhD projects and trajectories.
RDM 101 Impact Evaluation 01 December 2025 20 Figure 9. Participants’ agreement level with the relevance of the skills provided by the course to their future career. These results suggest that, upon completing the course, PhD candidates feel equipped with the necessary skills for effective research data management during their PhD project. The results presented in Figure 9 are the first indication that participating in the course is considered to provide skills that can be useful in the long-term. We consider these two indicators as strong evidence supporting the importance of continuing the course, while also recognizing the need for further improvements to ensure the course remains relevant, with specific areas for enhancement to be identified through this study. Participants Highlights and Suggestions As previously mentioned, Section 8 of the feedback form included an open-text response designed to gather participants’ insights on the most valuable aspects of the training and suggestions for improvement. Both these questions were optional. The responses on the question about the most valuable parts of the course were available from the first edition of the training in 2020 and are illustrated in Figure 10.
RDM 101 Impact Evaluation 01 December 2025 21 Figure 10. Aspects of the course participants found most useful and valuable for learning (n = 85). This plot was generated based on an open-ended question indicated in the title. The responses of participants were grouped and categorised manually. Participants were also asked to indicate suggestions that would make the course a better learning experience. We grouped the main responses of the participants into five broad categories: “Workload”, “Content and Relevance”, “Class design”, “Interactivity and Format”; “General positives.” The table below demonstrates the summary of the responses as well as main highlights extracted from them. Table 1. Participant suggestions to improve the course (n=71) Category name Summary and Highlights Workload Participants highlighted that the workload is too high relative to the course credits. Many reported that the amount of self-study, homework, and preparation time exceeds their expectations for a PhDlevel course. Specific suggestions include: 1. Reducing the volume of assignments. 2. Extending deadlines or allowing more preparation time. 3. Assigning more credits to reflect the actual time commitment.
RDM 101 Impact Evaluation 01 December 2025 22 Content and Relevance Some participants felt that the course content lacked disciplinespecific case studies and practicality in terms of examples and templates: 1. Include more discipline-specific case studies, tools and templates. 2. Provide examples of real data management plans and data preservation. Class design Participants had mixed opinions about the number and format of class sessions. While some advocated for more meetings to allow deeper discussion, others preferred fewer sessions or even a fully selfpaced format. Few participants also suggested improving the structure of the classes, particularly by reducing the emphasis on group discussions in favour of more direct instruction (lectures) and practical exercises: Reduce discussion time in favour of more focused lectures or hands-on exercises. Interactivity and Format Several participants found class sessions crowded and preferred inperson sessions over online classes: 1.Suggest smaller class and more structured breakout sessions for better discussion. 2.Prefer to have an option of offline (in-person) classes. General positives Substantial number of respondents found the course valuable and essential with strong recommendation on: 1. Making the course mandatory for all the PhD candidates. 2. Emphasizing its role in the early stages of PhD study. Participant comments related to “Content and Relevance” provide additional context for the results presented in Figure 4. Respondents expressed a need for content, use cases, feedback, and examples that are more closely aligned with their specific research topics. While the course already offers personalized feedback through discussion forums and extensive feedback on their assignments by the trainers and peers, recent cohorts, who often enter the course more aware about RDM and its importance, appear to seek more in-depth knowledge that supports direct application to their own research projects. Therefore, as a training team, we would like to explore the possibility of introducing tailored learning paths within RDM 101. That seems to be the most suitable approach, given the existing resource limitations and concerns about the time required from trainers to provide intensive feedback. It may also help balance participants’ growing expectations for advanced content with the practical limitations of the current course structure.
RDM 101 Impact Evaluation 01 December 2025 23 4.2 Preand Post-training Survey 4.2.1 Objectives of the Analysis Preand post-training surveys were designed as part of this research, to collect data aligned with the study’s objectives. The primary objective of preand post-training surveys was to assess the immediate knowledge, skills, and changes in perspective that participants gained through the RDM 101 course. To guide this evaluation, the following research questions were formulated: • To what extent does students' understanding of research data management change after taking the course? • What are the expectations of the participants before the course? And are they met during the training? • Which parts of the training students find more effective for gaining knowledge and developing skills? 4.2.2 Survey Overview & Structure The survey was designed around four main sections: 1. General information, including information about participants (e.g., PhD year, academic discipline), as well as questions assessing participants’ confidence and familiarity with RDM 2. Data practices - aimed at evaluating participants’ knowledge, practices, and satisfaction in the following areas: a. Data Storage and Backup b. FAIR Data c. Data Management Plan d. Data Publication 3. Learning expectations - questions assessing participants’ expectations prior to the course and perceived learning gains and training effectiveness after the course The surveys were distributed to participants across four runs of the RDM 101 course: three delivered in the first semester and one in the second semester of the academic year 2024/2025. They were conducted at two key time points: the pre-training survey was conducted four days prior to the start of the course, and the post-training survey was shared at the end of the final class session, and a reminder was sent the next day. No incentives were offered to participants, and participation was entirely voluntary, with no impact on course credits. No personal data of participants was collected. The privacy statement and consent text added at the beginning of the survey was drafted with the support of the data steward from the Faculty Technology, Policy and Management to ensure alignment with GDPR requirements. For convenience, from here onward, RDM 101 runs that were delivered with class sessions in online format will be referred to as "online runs," and those with in-person class sessions will be referred to as "in-person runs". For certain questions where it is relevant, we distinguish results of participants from online and in-person runs. In total 115 participants attended the course across the four considered runs, with 58 participants joining the online runs and 57 participants in the in-person runs. Of these, 95 completed the pretraining survey and 101 completed the post-training survey, resulting in response rates of 83% and 88%, respectively. To ensure the robustness of the results, we included in the analysis only responses of participants who completed both the preand post-training surveys (here and after
RDM 101 Impact Evaluation 01 December 2025 24 “matched responses”). The table below shows the distribution of participants who took the preand post-training the surveys. Table 2. Number of RDM 101 participants who completed preand post-training surveys. Mode Pre-training Survey Post-training Survey Matched Responses Online runs 49 49 33 In-person runs 46 52 32 Three participants (two from online runs and one from in-person runs) did not complete the posttraining survey, and their responses were therefore excluded from the analysis. As a result, the total number of matched participants included in the analysis is 62. Therefore, we consider that the effective response rate is 54%. 4.2.3 Methodology & Software Both, preand posttraining surveys include quantitative and qualitative mixture of questions. To match preand post-training survey responses from the same participant while maintaining anonymity, we included a question allowing respondents to generate a unique code. This code was based on information known only to them—that is, the first letter of the city where they were born, the number of siblings they have, and the first two letters of their current street. This three-elements ID code showed to be the most effective to avoid confusion between the participants. All surveys were designed and distributed using the TU Delft-licensed Qualtrics survey platform. The analysis of the results was conducted using the R programming language. The scripts used for the analysis are openly available on 4TUResearch Data repository (Grens, 2025) integrated with public GitHub repository to ensure the reproducibility of analysis. For open-text questions, an initial round of manual coding of a sample of open-text responses was conducted by one principal investigator (NR). Subsequently, the coding scheme was refined, consolidated, and validated in collaboration with a second investigator (PML). Categories were developed separately for the pre-training and post-training sample datasets, which resulted in distinct category sets where this was methodologically appropriate. Once categories were defined the remaining survey datasets were cleaned and coded following the same scheme (as described in the accompanying codebooks). Categories were then assigned either through direct command-line input in RStudio or via pre-generated Excel files that allowed iterative revision. This workflow ensured consistent, transparent categorisation across datasets, and the final categorised responses were subsequently summarised and visualised to support interpretation. To enhance transparency and reproducibility, the resulting codebooks, containing the final category structures and their conceptual descriptions, have been published in the 4TU.ResearchData repository (Rzayeva et al., 2025).
RDM 101 Impact Evaluation 01 December 2025 25 4.2.4 Results General Information Figure 11a and Figure 11b provide an overview of the survey participants. Figure 11a shows their faculty distribution, while Figure 11b represents their distribution by year in the PhD trajectory. Figure 11a. Faculty distribution of survey participants. Figure 11b. PhD trajectory year distribution of survey participants. To start with, we asked participants if they were aware about dedicated support available at the university/faculty for research data/software-related questions. Before the training 55% (n=34) of respondents reported they were not aware about any support services/resources available at TU Delft, and this percentage dropped to 2% (n=1) after taking the course. This was an open-ended question, allowing participants to indicate multiple options. Figure 12a and Figure 12b show the distribution of institutional support reported by participants before and after taking the course, respectively.
RDM 101 Impact Evaluation 01 December 2025 32 FAIR Data Both the preand post-training surveys included a question asking participants to provide at least three keywords that they think are related to the FAIR principles for data and/or software. The word clouds below visualize the most common keywords from both preand post-training surveys. Figure 17a. Keywords related with the concept of ‘FAIR principles for data and software’ indicated by participants in the pre-training survey. Participants were asked to provide at least three keywords. Figure 17b. Keywords related with the concept of ‘FAIR principles for data and software’ indicated by participants in the post-training survey. Participants were asked to provide at least three keywords. While the main aspects of FAIR were identified at approximately similar rates in both the preand post-training surveys, a closer look reveals a shift in the specificity of responses. After the training, the spread of words was greatly reduced with participants more frequently mentioning concrete methods and tools for making research data and/or software FAIR, rather than just referring to general concepts. This can be seen from the appearance of terms like “documentation,” “metadata,” and “README.” The combination of the word clouds and the results presented in Figure 15 suggests
RDM 101 Impact Evaluation 01 December 2025 33 that while participants seemed to be already familiar with the general definitions of FAIR, the course helped them gain more practical knowledge and skills for implementing these principles. After completing the course, all participants, except one, reported that they had planned or implemented specific measures/actions to make their research data and/or code (scripts) more FAIR, based on the knowledge and skills gained during the training. Participants’ responses were manually grouped into thematic categories, which are illustrated in the figure below. Figure 18. Examples of measures/actions to make research data/code more FAIR, based on knowledge and skills gained during the training. This plot was generated based on an open-ended question indicated in its heading. The responses of participants were grouped and categorised by the training team members. The only participant who reported not having planned or implemented any specific measures explained that their software already adhered to the FAIR principles, thanks to the prior knowledge before the training. Data publication Before the training 76% of participants planned to publish the data/code (scripts) of their research projects. This number grew to 95% after completing the course. The figure below illustrates which repositories participants indicated to publish their data/code (scripts) in preand posttraining surveys. Participants were allowed to indicate multiple repositories.
RDM 101 Impact Evaluation 01 December 2025 34 Figure 19a. Planned repositories for publishing research data/code (scripts) indicated by participants in the pre-training survey. This plot was generated based on an open-ended question indicated the title. The responses of participants were grouped and categorised manually. Figure 19b. Planned repositories for publishing research data/code (scripts) indicated by participants in the post-training survey. This plot was generated based on an open-ended question indicated the title. The responses of participants were grouped and categorised manually. Although the vast majority of participants already intended to publish research data or code (scripts) of their projects before the training, Figure 19a and Figure 19b show a shift in their choice of repositories. For instance, we see a twice doubled increase in preference for the 4TU.ResearchData repository after completing the course. This repository is specifically recommended during the training as a reliable option that meets key requirements related to institutional governance, data management best practices, and technical infrastructure, while also ensuring long-term data preservation. Zenodo, another recommended repository highlighted in the course materials, appeared in the post-training responses only, though at a low rate.
RDM 101 Impact Evaluation 01 December 2025 35 As also discussed in the section on data storage and backup, the relevance of platforms like GitHub and GitLab for participants remained relatively stable before and after the training. This suggests that most participants were already familiar with these platforms prior to RDM 101. We also see some confusion in the survey responses regarding the use of the Project Data (U:) drive and the TU Delft Repository for data publication, both of which are not intended for this purpose. Although this misunderstanding decreased after the course as indicated in the post-training survey, it still persisted among some participants. Before the training 20 participants reported that they do not plan to publish the data/code (scripts) of their research project. When asking participants about their reasons for not planning to publish their research data/code (scripts), the most common themes included concerns around privacy and confidentiality, especially when dealing with confidential data from external partners or data-type specific databases, which restricts public sharing due to contractual or legal obligations. Another frequent reason was uncertainty or indecision when participants indicated that they had not yet discussed data publishing plans with their supervisors or advisors. Some respondents explained that the data/code (scripts), in their view, would not be broadly beneficial to the scientific community. Several respondents mentioned working with personal data. After the training, the number of participants who reported that they do not plan to publish any data or code (scripts) from their research project decreased to 3. The reasons behind these responses matched the same as before the training: confidentiality concerns and uncertainty or indecision. Learning Expectations As a team of trainers, we were interested in understanding how different training components performed across various delivery formats — specifically during online and in-person runs. The figure below presents participants’ ratings of the helpfulness of specific training components in both formats. Figure 20. Participants’ ratings of how helpful each training component was in facilitating effective learning, presented separately for in-person and online formats. Results are shown as percentages on a 1–5 scale, where 1 = Not helpful at all and 5 = Extremely helpful. Participants could rate each component independently. Most course components are considered effective for facilitating learning, with the majority of ratings falling between 3 and 5. Looking separately at the delivery formats, we see that the effectiveness of assignments and the feedback provided, as well as the materials and activities provided through the
RDM 101 Impact Evaluation 01 December 2025 36 LMS appeared to be equally effective across both formats, with the balanced distribution of ratings between “4” and “5” points. Notably, in-class activities and interaction with peers and trainers received higher ratings from almost twice as many participants in in-person sessions compared to online sessions. These results provide us with valuable insights on which components we should prioritise and invest resources in when updating the course. Before the training we asked participants if they can provide an example of how they think this course could help them with their research. This question was optional, the number of participants replied was 33. It was an open-ended question and the responses of participants were grouped and categorised by the training team members. The top five categories we could identify among responses include efficient data/code organization/structure/management, guidance in writing/finalising a DMP, data sharing/publication, data storage/backup, and accessibility measures. After the course, 74% of participants reported that the training helped them address specific challenges or problems related to data and/or code (scripts) management in their research projects. This question was open-ended and the responses of participants were grouped and categorised by the training team members. The top five categories that emerged were: folder structure/data and code organization, use (standard) metadata, data/code documentation (readme, data dictionaries, etc.), file naming conventions, and research data/software management practices. Before the training, participants were also asked what they expected to learn most from the course. The question was open ended and participants’ responses were again manually categorised. The aspects that participants expected to learn more about from the course include efficient data/code organisation/structure/management, storage/backup practices, DMP writing and guidance, TU Delft RDM guidelines/regulations, and accessibility measures. To finalise the post-training survey, and get a global impression about satisfaction of the participants with the course we asked participants to rate how the course met their expectations. The results are illustrated in Figure 21 combined overall for all participants (top part), and separately for the online and in-person runs (bottom part).
RDM 101 Impact Evaluation 01 December 2025 37 Figure 21. Extent to which the course met participants’ expectations. Despite the differences in the evaluation of certain course components between the in-person and online formats observed in Figure 20, the results in Figure 21 show that the course meets or exceeds learners' expectations, regardless of the delivery mode.
RDM 101 Impact Evaluation 01 December 2025 38 5 Long-term impact of RDM 101 5.1 Objectives of the Analysis To assess the long-term impact of the RDM 101 course on participants’ knowledge, attitudes, skills, and behaviours related to research data management, we conducted semi-structured interviews with PhD candidates who had completed the course RDM 101 between 1 and 1.5 years prior to the interview. Additionally, two data stewards were interviewed to provide external perspectives on any observable changes in participants’ approaches to RDM. 5.2 Interview Overview & Structure Each interview lasted up to one hour and followed a semi-structured format. We first asked participants general questions about when they attended the RDM 101 course and whether they took it in person or online. The conversation then explored their motivations and expectations for joining the course, as well as their prior knowledge and experience with research data management practices. We then invited participants to reflect on what they learned during the course, how they interacted with trainers, data stewards, and peers, and how they experienced the training overall. They were also asked to describe any changes in their RDM practices after completing the course, how they applied the knowledge gained, and whether they continued to use the skills they had learned. The interviews concluded with a request for suggestions and recommendations on how the course could be improved. Below is a copy of the interview guide we used as a prompt for the discussions. While the guide provided a structure, additional ad-hoc follow-up questions were asked to the interviewed participants as the conversation developed. These spontaneous prompts were tailored to each interview and are not included in the guide. PART A - INTRODUCTION • Can you introduce yourself? • Which year of the PhD are you? • Which faculty? PART B - Experience before or during the course • When did you attend the RDM 101 course? • Did you attend in person or the online course? • Why did you choose to attend the RDM 101 course? • What was the level of the RDM knowledge before you took the course? • How and how much do you think the course helped you to organize the data management of your PhD study? Can you please recall an event/ episode when you thought that what you were learning was crucially important? Wakeup call PART C - Experience after the course • Can you describe what you have learned from this course? What would be different in your research if you didn’t take the course? • Were you in touch with the course educators or the data steward at your faculty after the RDM 101 course? Do you remember for which reason? • [if they took the course online]: how would you describe the online course experience?
RDM 101 Impact Evaluation 01 December 2025 39 • [if they took the course in person]: how would you describe the in-person course experience • Do you think you will use the knowledge acquired during the course in the future? If so, how? • Do you have any recommendations/ advice, or feedback for trainers of the course? On the design of the course? 5.3 Methodology & Software We used a qualitative descriptive study design, informed by Braun and Clarke (2006). For the thematic analysis framework and reporting the Consolidated Criteria for Reporting Qualitative Research (COREQ) (Tong, Sainsbury, & Craig, 2007) was used. Interviews were conducted with 11 PhD candidates who had completed the course 12 to 18 months earlier and two data stewards from different faculties at TU Delft. Participants represented diverse disciplinary backgrounds and worked with various data types, including personal and sensitive data, source code, large-scale simulations, and large language models (LLMs). This variation reflects the interdisciplinary reach of the RDM 101 course and its relevance across different research contexts. Participants were selected through purposeful sampling, a technique used to identify and recruit individuals who can provide rich, relevant, and diverse information. This approach ensured variation across faculties, types of research projects, and data types (e.g., personal data, code, large language models). Importantly, selection was not based on participants’ performance in the RDM 101 course. All participants were recruited voluntarily and provided informed consent. Participants were not remunerated for their time. We pseudo-anonymised data collected from participants to protect their identities and ensure confidentiality, particularly because direct quotes are used throughout the report. Throughout the report, we removed or generalised any identifiable information to ensure that participants could not be recognised based on their words, faculty, or project context. We conducted the data analysis using TU Delft–licensed Atlas.ti (version 23), a commercial qualitative data analysis software used to facilitate coding and organise qualitative data. Data was analysed thematically using an inductive approach. Coding followed a constant comparative method, a process of continuously comparing new data with existing codes to refine patterns and categories. Participants’ own words were used to guide code development, strengthening the credibility of the analysis. Two principal investigators (CS and GK) independently reviewed and manually coded the transcripts to enhance familiarity with the data and to strengthen the trustworthiness of the findings. A third researcher (NR) joined later to review sections and contribute to refining the themes. Discrepancies in coding were resolved through discussion among the team until consensus was reached. Initially, we analysed the responses from PhD candidates and data stewards separately; however, as the analysis progressed, both commonalities and differences across the two groups were examined. We identified six main themes that correspond closely to the research goals and objectives and the participants’ experiences. Reflexivity was applied throughout the process to remain aware of researcher influence and to prioritise participant voices. Trustworthiness of the findings was supported through triangulation (multiple coders), rich description of context (the impact of RDM 101 over time), and a transparent audit trail from protocol to reporting. Due to ethical and privacy considerations, we do not share verbatim transcripts or detailed coding linked to individual participants, in order to maintain confidentiality. The interview protocol, a blank version of the participant information sheet and consent form, and documentation explaining the methodology are openly available via the 4TU.ResearchData repository (Rzayeva et al., 2025).
RDM 101 Impact Evaluation 01 December 2025 40 5.4 Results In the analysis we followed an inductive approach. This allowed themes to emerge directly from the participants’ narratives. This method ensured that the findings remained grounded in their lived experiences and perspectives. The findings are organized into four key themes: • Motivations and Expectations – exploring why participants enrolled in the course and how well it aligned with their needs. • Knowledge Acquisition and Impact – examining what participants learned and how it influenced their data management behaviours. • Course Experience and Feedback – capturing participants’ impressions of the course structure, delivery, and learning tools. • Emotional and Social Dimensions – highlighting the role of emotions and peer interactions in shaping the learning experience. We believe that, taken together, these themes offer a deeper understanding of the course’s effectiveness from the participants’ perspectives and highlight areas for potential improvement. Theme 1: Meeting Motivations and Expectations and Navigating Challenges Participants enrolled in RDM 101 with a wide range of motivations, shaped by both institutional requirements and personal interests. These motivations generally fell into three categories: • External Motivators: Many participants attended the course to fulfill Graduate School requirements, such as earning credits or completing a Data Management Plan. For some, this was perceived as a bureaucratic necessity. As participant C noted, “At first, it felt like a box-ticking exercise; I needed to complete my Data Management Plan.” • Internal Motivators: Others were driven by a genuine desire and aspiration to improve their research practices. Participant A shared, “I was keen to learn how to make my data FAIR and reusable in the future.” Similarly, participant M explained, “It's something that I wanted to learn early on in the research process because I know that data management is important at every stage, so I enrolled.” • Practical Needs: Several participants were already working with large or complex datasets and needed immediate, applicable strategies. Participant B, for example, reflected: “As soon as I started, I had to do a lot of simulation and there was a lot of data. I was already missing some simulations, some codes, so I said, ‘OK, it's better to follow a course and see what they suggest and how to organize the data.” Participant E added “One important aspect is how to work with companies. If you focus only on data management that can be done at TU Delft, but don’t talk about security protocols or how to ensure data is secure; or working with commercially constrained data, that’s a gap.” This diversity of motivations shaped how PhD candidates engaged with the RDM 101 course. Some approached it with minimal expectations, while others hoped for a transformative learning experience. Data stewards also reflected on this range of engagement among PhD candidates, noting, “We see the course as both a necessity and an opportunity” [Participant E]. Overall, participants appreciated the course’s broad and flexible design, which offered both foundational knowledge and practical tools. However, a notable difference in expectations emerged between those seeking valuing general foundational knowledge and those looking for disciplinespecific data management practices and guidance. Particularly participants from the social sciences and humanities expressed a desire for more tailored content. One participant, who worked with sensitive survey and interview data, shared: “As a PhD student, you have no colleagues to rely on. I
RDM 101 Impact Evaluation 01 December 2025 41 thought this course would teach me everything I need to know to be compliant [with ethical and privacy requirements]” [Participant N]. Despite receiving individual feedback on assignments, some participants found it unclear whether the course addressed social science data needs or focused primarily on technical software domains. And while trainers were described as helpful and motivated some social science-specific questions were often redirected to faculty data stewards. As one participant noted, “Some questions they could answer and some they could not. I guess they were also more from a technical background” [Participant C]. This feedback highlights a key tension: while RDM 101 is designed as a general introduction to research data management, it cannot fully meet the specific needs of every discipline within its limited scope. Nevertheless, several participants found value in the interdisciplinary nature of the course, which encouraged reflection on their own practices through exposure to diverse research contexts. One participant remarked, “Even though it wasn’t about my specific field, hearing about other disciplines’ data challenges made me rethink my own practices” [Participant H]. Importantly, this broader perspective seemed to encourage participants to continue engaging with research data management even after the course ended. Many reached out to their faculty data stewards and supervisors with follow-up questions, suggesting a lasting shift in awareness and a more proactive attitude toward managing their data. At the same time, the variation in participants’ expectations highlights an ongoing challenge: balancing a general, foundational approach with the specific needs of different disciplines. This is especially relevant given that RDM 101 is designed to provide broad principles rather than specialised, discipline-specific instruction. Theme 2: Knowledge Acquisition, Impact on Practice and Cognitive Shifts. The RDM 101 course appeared to enhance participants’ theoretical understanding and practical application of research data management principles, as reflected in their descriptions of changes in attitude and behaviour. Many reported gaining new competencies in areas such as version control (e.g., Git), data storage strategies, file naming conventions, and documentation standards like READMEs. For instance, the in-person session on data documentation tools helped participants better record experiments and their versions and ensure compliance with repository requirements like 4TU.ResearchData. Participants noted how these tools streamlined their workflow and improved data traceability. Beyond acquiring technical skills, the course also encouraged a deeper mindset shift, participants began to approach data management more thoughtfully and deliberately in their day-today research. One participant shared: “Before the course, I never really thought about how I stored my data or named files. Now [after one year] I’m much more deliberate and consistent.” [Participant M] Others noted how the course helped initiate an iterative process of revisiting data and refining methods. Participant C explained “That kind of feedback between thinking about your data and then your methods, and then back to your data, really helped me. It was a process that was started by this course… ” One of the most appreciated parts of the course was the introduction to the FAIR data principles (Findable, Accessible, Interoperable, and Reusable). Learning about these concepts encouraged participants to reflect critically on their own data practices. Participant A admitted, “I definitely think it's useful because I wasn't fully aware of data management before. Even if I was struggling with poor data management from someone who worked on my part of the project before me.” Several mentioned that the frameworks and tools introduced during the course helped them save time and work more efficiently, especially when preparing to publish their research. "My second project [I] already had a ready a folder called data management. So every time I collect more data I immediately put it the right way in there, so I know that as soon as it's finished and I want to publish, I can just basically do copy paste and put it in the for TU directory, which is really helpful because it saves you a lot of time if you have to do that last minute - when you find out that that is required." [Participant N]
RDM 101 Impact Evaluation 01 December 2025 48 7 Conclusion Overall, the results of this study demonstrate that RDM 101 is achieving its core objectives - improving participants’ knowledge, confidence with, and application of best practices in research data management. The study evidences that the course not only introduces foundational theoretical concepts but also equips participants with practical knowledge and skills that can be directly applied to their research workflows. We also observed a lasting shift in mindset, with many participants coming to see RDM as an integral part of the research process. The report also provides concrete evidence of the importance of investing resources and having an expert training team at the library to design and deliver the course (and other courses) in an interactive and practical way to make a long-term impact on the integration of RDM practices in research at TU Delft. Because of the abstract nature of concepts like RDM, FAIR data, and FAIR software, there is a need to provide opportunities for participants to learn from each other, receive guided feedback, and actively share challenges and solutions. This approach helps translate abstract concepts into practical relevance for participants' research, motivating them to apply what they have learned. Additionally, having sufficient resources within the training team allows us to dedicate time for evaluating our courses in detail and updating them in a timely manner, which is essential for fast-evolving topics such as RDM, FAIR data and software, and to meet the needs of an evolving community of learners. At the same time, there is room for further refinement, especially in managing the course workload and credit balance, implementing a more modular, and discipline-tailored learning paths within the course, and adopting a customizable format. The report also provides input for thinking about the format of course delivery. The feedback received through the Graduate School feedback forms showed that in-person delivery was generally preferred by participants, while during the interviews, participants acknowledged that formats such as online classes and a completely self-paced format could increase accessibility. The results of this report confirm that a blended format of the course is the best way to maintain the short-term and long-term impact of the course. However, as a team, we acknowledge that further exploration is needed to improve the blended format with online classes through improved session design, delivery strategies, and the use of interactive tools. The findings of this study provide strong support for the current design of RDM 101 while highlighting promising directions for future development of the course. We think that the results of this study, along with the ongoing analysis of new feedback data and continued collaboration with stakeholders, such as researchers, data stewards, and faculties, will provide valuable insights for future updates to the course, ensuring it remains relevant and effective. With this, we believe that newly structured and well-supported RDM 101 training will enhance its role as means of fostering Open Science and responsible research practices at TU Delft.
RDM 101 Impact Evaluation 01 December 2025 49 Data and Code Availability Due to ethical and privacy considerations, the raw research data containing potentially confidential and identifiable personal information is securely stored on the university-curated network storage, the Project Data (U:) drive, and are available upon request. Researchers can request access to the raw data by contacting the TU Delft Library Research Data Services Training Team at rdmtraining- [email protected]. To support transparency and reproducibility of the study, the anonymized quantitative datasets and the high-level qualitative categories, together with their conceptual descriptions and containing no personally identifiable information about participants, are published in the 4TU.ResearchData repository (Rzayeva et al., 2025). These materials include: • Preand post-training survey questionnaires and opening statement • Cleaned preand post-training qualitative survey data • Codebooks for qualitative survey data • Cleaned GS feedback form data • Codebooks for qualitative GS feedback data • Interview protocol form • Interview methodology • The participant information sheet and consent form • Code group data The scripts used for the analysis are openly available on 4TUResearch Data repository (Grens, 2025) integrated with public GitHub repository to ensure reproducibility of analysis.
RDM 101 Impact Evaluation 01 December 2025 50 References 1. Ahlers, I., Andrews, H., Cruz, M., van Dijck, J., Dintzner, N., Dunning, A., Eggermont, R., den Heijer, K., Ilamparuthi, S., van der Kruyk, M., van der Kuil, A., Kurapati, S., Love, J., Martinez Lavanchy, P., Plomp-Petersen, E., de Smaele, M., Teperek, M., Turkyilmaz-van der Velden, Y., Versteeg, A., & Wang, Y. (2020). TU Delft Research Data Framework Policy (V2). Zenodo. https://doi.org/10.5281/zenodo.4088123 2. Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. https://doi.org/10.1191/1478088706qp063oa 3. Delft University of Technology. (n.d.). The PhD development cycle. https://www.tudelft.nl/en/ae/research/ae-graduate-school/the-phd-development-cycle 4. Grens, N.(2025). Analysis Code Supporting "RDM 101 Impact Evaluation: A Study on Training Effectiveness and Research Practice Changes". 4TU.ResearchData. https://doi.org/10.4121/f7e9b088-cf91-464d-a6a9-772fde5d8bff 5. Martinez-Lavanchy, P. M., Rzayeva, N., Dauwels, S., Barrera-Coronel, M., Utrilla Guerrero, C., Dace, H., Strubbia, C., van Schöll, P., & Zormpa, E. (2024). TU Delft Research Data Management 101 course (Version v2) [Lesson]. Zenodo. https://doi.org/10.5281/zenodo.10732095 6. Rzayeva, N., Morselli, F., Strubbia, C., Grens, N., Kulkarni, G., & Martinez-Lavanchy, P. M. (2025). Data supporting "RDM 101 impact evaluation: A study on training effectiveness and research practice changes" [Data set]. 4TU.ResearchData. https://doi.org/10.4121/90a1d562-8496-44f3-b8a55610ddc871d5. 7. Tong, A., Sainsbury, P., & Craig, J. (2007). Consolidated criteria for reporting qualitative research (COREQ): A 32-item checklist for interviews and focus groups. International Journal for Quality in Health Care, 19(6), 349–357. https://doi.org/10.1093/intqhc/mzm042
RDM 101 Impact Evaluation 01 December 2025 51 Acknowledgments We would like to express our gratitude to Selin Kubilay and Elviss Dvinskis (Data Managers at the Digital Competence Center - DCC), for their great support, expertise and enthusiasm. Their guidance in ensuring that the code for analysing the data were developed according to best practices and properly documented was essential to ensure reusability and reproducibility of our analysis. Our other big gratitude is to Nicolas Dintzner (TPM Data steward) for his invaluable support throughout the entire project. His expertise was crucial in developing the Data Management Plan for this project, preparing the opening statement for surveys and informed consent for interviews, and navigating the HREC application process. We are very thankful for his flexibility and guidance in addressing several RDM tasks during the study. A big thank you goes to Judith Linnemann-Karel (Central Graduate School) who took the time to provide us with the feedback forms data in the format we needed for this study.
RDM 101 Impact Evaluation 01 December 2025 52 Author Contributions Narmin Rzayeva (NR) (RDS trainer) was the project lead of the RDM 101 Impact study. She oversaw the design of the project plan, and also kept the coordination of all the activities of the different contributors. She contributed to the design of the methodology with a particular focus on quantitative aspects of the research, designing and analysing preand post-training surveys, as well as supporting the analysis of the Graduate School feedback forms. Francesca Morselli (FM) (Postdoc VU Amsterdam) contributed to this investigation with a particular focus on the qualitative research component, supporting the analysis of the theoretical framework, the design of the methodology, as well as the development of interview protocols, conducting interviews and designing of preand post-training surveys. Carla Strubbia (CS) (RDS trainer) contributed to the project as a research data management training expert. She was responsible for guiding the analysis of qualitative data and shaping the overall evaluation approach. Her experience in research training, qualitative methodologies and commitment to improving data practices ensured that the study was both rigorous and impactful. Nikki Grens (NG) (RDS student assistant) contributed to the project as a student assistant from its start. She was responsible for preprocessing the data, developing the codes in R, generating plots, and documenting the analysis. Her prior knowledge, combined with her enthusiasm and commitment to learning, enabled the implementation of best practices acquired during the project, ensuring a smooth and efficient analysis process.
RDM 101 Impact Evaluation 01 December 2025 53 Gargi Kulkarni (GK) (RDS student assistant) contributed to the project as a student assistant to support us with the analysis of the interviews. She was part of the team working on the qualitative part of the study. Her curiosity and commitment to learning supported the implementation of thoughtful and rigorous analysis throughout the research project. Paula Martinez Lavanchy (PML) (RDS training coordinator) is the initiator and facilitator of this research study. She also contributed to the development of the methodology for quantitative data analysis and provided subject-matter expertise for validating the qualitative results.
RDM 101 Impact Evaluation 01 December 2025 54