scieee AI-readable full text Open interactive document viewer

Das Marie Kondō Prinzip für Forschungsdaten im Bildungswesen

Seidel, Niels

Abstract

Abstract:Am Forschungszentrum CATALPA beschäftigen wir uns mit einer stark datengetriebenen Forschung an den Schnittstellen von Hochschule, Bildungstechnologie und Künstlicher Intelligenz. In dem geplanten Vortrag möchte ich Erfahrungen aus mehreren interdisziplinären Forschungsprojekten weitergeben und mich dabei auf die praktischen Methoden, Prozesse und Werkzeuge für den Umgang mit Forschungsdaten konzentrieren, die unsere Arbeit vorantreiben. In Anlehnung an Marie Kondo liegt der Fokus in diesem Jahr auf dem Wegwerfen und Aufräumen von Forschungsdaten.

Full text

FernUniversität in Hagen CATALPA – Center of Advanced Technology for Assisted Learning and Predictive Analytics Das Marie Kondō Prinzip für Forschungsdaten im Bildungswesen Niels Seidel Hagen, 2025/11/18 2 Image source: https://learn.konmari.com/ 3 Books METADATA AND PRESENTATION SENTIMENTAL DATA PUBLISH RESEARCH DATA DATA SHARING LIVING WITH RDM Image source: https://learn.konmari.com/ DATA KONMARI 101 MAKE A DATA MANAGEMENT PLAN FAIR DATA FASHION DATA USE AND PROTECTION RESEARCH DATA POLICIES WHAT IS THE PROBLEM WITH RESEARCH DATA? RDM TOOLS 4 Books doi: 10.5281/zenodo.17668427 METADATA AND PRESENTATION SENTIMENTAL DATA PUBLISH RESEARCH DATA DATA SHARING LIVING WITH RDM Image source: https://learn.konmari.com/ DATA KONMARI 101 MAKE A DATA MANAGEMENT PLAN FAIR DATA FASHION DATA USE AND PROTECTION RESEARCH DATA POLICIES WHAT IS THE PROBLEM WITH RESEARCH DATA? RDM TOOLS 5 WHAT IS THE PROBLEM WITH RESEARCH DATA? Image source: https://learn.konmari.com/ Where I’ve started … 6 Where we are … 7 Melkamu Jate, Wabi; Striewe, Michael (2023): Umgang der DELFI-Community mit Forschungsdaten und Softwareartefakten - Eine Erhebung auf Basis der Tagungsbände im Zeitraum 2018-2022. 21. Fachtagung Bildungstechnologien (DELFI). DOI: 10.18420/delfi2023-27. Bonn: Gesellschaft für Informatik e.V.. pp. 167-172. Where we are … 8 Haim, A., Shaw, S., & Heffernan, N. (2023). How to Open Science: A Principle and Reproducibility Review of the Learning Analytics and Knowledge Conference. LAK23: 13th International Learning Analytics and Knowledge Conference, 156–164. https://doi.org/10.1145/3576050.3576071 LAK’22 + LAK’23 2. Phase The Data-Driven Revolution in Education 9 Phases of digitization in education (Wahlster, 2017) 2018 BERT … 2019 GPT-2 … 2025 GPT-5.1 Source: S. Joksimović, V. Kovanović, S. Dawson. The Journey of Learning Analytics. HERDSA Review of Higher Education 1. Phase Developing Comprehensive Data Management Plans (DMPs) 16 How? - make data management a work package in your project - make DMP creation a joint task - use a living document - don’t use any forms and DMP tools - blend structures and questions of multiple DMP templates to find a suitable structure for your project - update and revise the document regularly 17 DATA KONMARI 101 Image source: https://learn.konmari.com/ Signs of Data Chaos 18 Research Data ● poor structure ○ inappropriate folder structure ○ inconsistent naming schemes ○ proprietary file formats ○ compressed data ● cluttered ○ unclear separation between raw data and processed data ○ work-in-progress data ○ secondary data such as paper drafts ○ redundant data, e.g., for multiple platforms (e.g., SPSS, R, Python) ○ conflicting file versions ○ overly bulky tables ○ unrelated comments ● lack of documentation ○ no general overview ○ no metadata or codebooks Special Case: Research Software ● dedicated for personal use only ● multi-purpose tools ● install instructions missing ● no lisence declared ● unclear maturity level ● files/folders missing ● passwords mentioned 19 Data Konmari 101 1. Tidy everything up at once, in a short time and perfectly. 2. All data to be tidied up is collected in one place. 3. Decide what to keep based on the question: Does it make me happy when this data is used in this way? 4. Every file that is kept is assigned its place. 5. All files must be stored there correctly. -Reserve time exclusively for data management -Radical inventory instead of continuous chaos - Establish a clean, sustainable system once and for all 20 Data Konmari 101 -Gather all data: local hard drives, external storage devices, cloud services, email attachments, USB sticks -Create a central overview: What actually exists? - Make data visible that is lying dormant in forgotten folders -Create catalogues: What data records, analyses and raw data are available? 1. Tidy everything up at once, in a short time and perfectly. 2. All data to be tidied up is collected in one place. 3. Decide what to keep based on the question: Does it make me happy when this data is used in this way? 4. Every file that is kept is assigned its place. 5. All files must be stored there correctly. 21 Data Konmari 101 ✅ Keep if data … - ensures reproducibility (for published results) - must be retained for legal/contractual reasons - is unique and cannot be re-collected - has potential for future research questions ❌ Remove if data is … - duplicate or outdated version - test data or result of failed analyses - (raw data already available in processed form) - scientifically/methodological incorrect 1. Tidy everything up at once, in a short time and perfectly. 2. All data to be tidied up is collected in one place. 3. Decide what to keep based on the question: Does it make me happy when this data is used in this way? 4. Every file that is kept is assigned its place. 5. All files must be stored there correctly. 22 Data Konmari 101 - Define a clear folder structure and storage locations - Storage locations by function: -Active work: locally on the computer -Collaboration: Coscine, GitLab, GitHub -Long-term archiving: (institutional repository) -Publication: OSF, Zenodo, subject-specific repositories 1. Tidy everything up at once, in a short time and perfectly. 2. All data to be tidied up is collected in one place. 3. Decide what to keep based on the question: Does it make me happy when this data is used in this way? 4. Every file that is kept is assigned its place. 5. All files must be stored there correctly. 23 Data Konmari 101 -File names: Descriptive and consistent (YYYY-MM-DD_project name_description_v01.csv) -README files: In each project folder (what, when, who, how) -Metadata: Use standards (Dublin Core, DataCite) -Version control: Git for code, documented versions for data -Backup rule: 3-2-1 (3 copies, 2 media, 1 external) -Access control: Who is allowed to view/edit what? -Data management plan (DMP) - Software management plan (SMP) 1. Tidy everything up at once, in a short time and perfectly. 2. All data to be tidied up is collected in one place. 3. Decide what to keep based on the question: Does it make me happy when this data is used in this way? 4. Every file that is kept is assigned its place. 5. All files must be stored there correctly. 24 DATA PROTECTION FOR DATA USE Image source: https://learn.konmari.com/ Data Protection for Data Use 25 Record of Processing Activities (VVT) - purpose, data/categories, process, persons/parties, deletion - technical-organizational measure Informed consent - OPT-IN - Example: https://aple.fernuni-hagen.de/admin/tool/policy/viewall.php#policy-6 Without informed consent? (see fulltext) Data Protection for Data Use 32 Ensure k-anonymity Simulated data - generate data with similar distribution as the original dataset - = truly anonymous Data Protection for Data Use 33 Replacing first and last names from names_dataset import NameDataset print(NameDataset().search('Frodo')) Video image distortion (see manual for VLC) Disguise voice - Audacity: Pitch -20% > Distortion Pixelate faces - Gimp: select > filters/blur > pixelate 34 METADATA AND PRESENTATION Image source: https://learn.konmari.com/ Metadata, presentation, use 35 Datacite metadata standard Software - e.g. mod_longpage readme.md incl. SMP Data and data analysis - notebooks: Jupyter notebooks, RMarkDown - web app: R Shiny Apps (gallery) - sandboxes: learnr (R) FAIR principles for AI Models - open datasets incl. notebooks to explore the data - baseline model + trained model for multiple platforms incl. notebooks - HDF5 or ROOT file format 36 PUBLISHING RESEARCH DATA Image source: https://learn.konmari.com/ Publishing Research Data 37 Service Data types DOI File size limit API Zenodo files, tables, software, models yes 50 GB yes CMU Data Shop * no - yes Open Science Framework * yes - no GitHub free source code no 2 GB yes re3data - registry for research data reporitories - should have a declared plan to enable permanent accessibility Data Journals - List of general data journals, e.g. for software: Software Impacts, J. Open Source Software, JORS anonymous.4open.science Summary & Outlook 38 Summary - Unique RDM challenges in educational research - Efforts with DMP, Record of Processing Activities, informed consent, and metadata descriptions are feasible - Clean research data is a prerequisite for collaboration -Data sharing is possible despite of privacy obligations - RDM works best as a team Outlook - NFDI via Verbund Forschungsdaten Bildung - COSCINE.nrw - RDM policy Q&A 39 Thank you! doi: 10.5281/zenodo.17668427 Dr. Niels Seidel [email protected] Notes and references 40 Wilkinson, M. D. et al. The FAIR guiding principles for scientific data management and stewardship. Sci. Data 3, 160018, https://doi.org/10.1038/sdata.2016.18 (2016). Wilkinson, M. D. et al. A design framework and exemplar metrics for FAIRness. Scientific Data 5, 180118, https://doi.org/10.1038/sdata.2018.118 (2018). Tso, et al., "The R Journal: Advancing Reproducible Research by Publishing R Markdown Notebooks as Interactive Sandboxes Using the learnr Package", The R Journal, 2022 NFDI: https://www.nfdi.de/konsortien/ Khalil, M., & Prinsloo, P. (2025). The lack of generalisability in learning analytics research: why, how does it matter, and where to? Proceedings of the 15th International Learning Analytics and Knowledge Conference, 170–180. https://doi.org/10.1145/3706468.3706489 Liu, Q., & Khalil, M. (2023). Understanding privacy and data protection issues in learning analytics using a systematic review. British Journal of Educational Technology, 54(6), 1715–1747. https://doi.org/https://doi.org/10.1111/bjet.13388 Wahlster, W. (2017, June). Künstliche Intelligenz als Treiber der zweiten Digitalisierungswelle. IM+io, 4.