Mapování správy výzkumných dat na ČZU – Komplexní zpráva
Abstract
Tato komplexní zpráva představuje výsledky prvního celouniverzitního šetření zaměřeného na správu výzkumných a studentských dat na ČZU. Průzkum z roku 2025 mezi 227 respondenty analyzuje oblasti plánování, ukládání, sdílení, dokumentace a bezpečnosti dat. Dokument shrnuje současnou praxi, bariéry a potřeby a formuluje doporučení pro posílení otevřené vědy a správy dat podle principů FAIR na ČZU.
Full text
Mapping of Data Management at CZU Comprehensive Report Author: Data Management Competence Centre (DMCC), Library, Czech University of Life Sciences Prague (CZU) Date: 1. 10. 2025 Document version: 1.0 Contact: openscienc[email protected]
2 Title: Mapping Research Data Management at CZU – Comprehensive Document Author: Data Management Competence Centre (DMCC), Library Institution: Czech University of Life Sciences Prague (CZU) Abstract: The document summarises the results of a questionnaire survey conducted in 2025 among CZU employees and academics, focused on the management of research and student data. The output is an analysis of current procedures, needs, and barriers to data planning, storage, sharing, and documentation. 227 respondents participated in the survey. Keywords: Research data, student data, data management, FAIR principles, open science, RDM, DMP, CZU, library, academic environment Language: English Format: PDF Date: 1. 10. 2025 Contact: openscienc[email protected] Version: 1.0 Tools used: Google Forms, MS Excel, Power BI, MS Word, ChatGPT, Adobe InDesign License: Note on updates: The document may be updated in the future. Minor changes, such as clarifications of wording, corrections or additions of minor information, will be indicated by incrementing the decimal number of the version (e.g. v1.1, v1.2). Major modifications – in particular the addition of interpretations, the development of measures or the extension of the analysis – will lead to the release of a new major version (e.g. v2.0, v3.0). Repository: https://osf.io/mp9h2 Supporting materials • Anonymised dataset (.csv): without identification, contains answers from the questionnaire survey: 20250730_survey-data.csv (https://osf.io/z5esa) • Methodological description (.txt): description of the purpose of the document, objective, target group, data collection, questionnaire structure, dataset format, overview of variables (name, data type, number of values): Methodology_v1.0.txt (https://osf.io/sh2zj) • Tools used (.txt): list of software used in data processing and document production: Tools_used_v1.0.txt (https://osf.io/9ptbw) • License (.txt): terms of use of this document and dataset: LICENSE.txt (https://osf.io/pqy4x) • Readme file (.txt): Description of the contents of the dataset: README.txt (https://osf.io/x9ywe)
3 Version history Title Version Processed by Date Revision First Edition 1.0 DMCC 1. 10. 2025 Acknowledgements On behalf of the authors, we would like to express our sincere gratitude to the network of faculty data stewards who played a key role in preparing and implementing the Mapping Research Data Management at CZU study. Their support and feedback were crucial in questionnaire design, pilot testing, dissemination across faculties, identification of key stakeholders, and revision of the resulting documents. We also thank all colleagues who contributed their time, expertise, and practical suggestions. Without their willingness to share experiences and collaborate, this document could not have been produced in its present form.
4 Content Introduction of The Issue............................................................................................................................... 5 Survey Overview and Methodology ............................................................................................................... 7 Summarised Results of The Data Survey ..................................................................................................... 8 Identification questions ................................................................................................................................. 8 Research Data .............................................................................................................................................. 10 Student Data ................................................................................................................................................. 37 Conclusion .................................................................................................................................................... 43 Key Findings ................................................................................................................................................. 44 Recommendations for Implementing Open Science and Data Management ........................................ 47 List of Images ............................................................................................................................................... 50 Abbreviations and Glossary ........................................................................................................................ 51
5 Introduction of The Issue Open science is a priority of both European and Czech science policy. In the European context, the requirements for open access to research data are mainly embedded in the Horizon Europe 1 2 programme, which sets out the obligation to process and continuously update the DMP 3 and emphasises adherence to FAIR principles. These principles do not lead to the mandatory opening of all data, but to its quality management. This is also related to the European Commission's principle "as open as possible, as closed as necessary"1, which emphasises data openness only to the extent that it does not jeopardise, for example, the protection of personal data or intellectual property. The European strategy is followed by the Czech legal framework, in particular the amendment to Act No. 130/2002 Coll. 4 , on the Support of Research, which introduces a legal definition of research data, the obligation to plan their management, record them in the Research, Development and Innovation Information System (IS of R&D&I Council) and, under certain conditions, make them available free of charge after the end of the project 5 . The obligation to publish research data is supplemented by exceptions that consider the 1 European Commission: EU Grants: Horizon Europe Programme Guide (version 5.0, 15.05.2025), part 17. Open science – open science obligations (OA publications, DMP, FAIR) and the principle of "as open as possible, as closed as necessary". Available from: https://ec.europa.eu/info/fundingtenders/opportunities/docs/20212027/horizon/guidance/programmeguide_horizon_en.pdf (Accessed: 26/08/2025). 2 European Research Executive Agency (REA): Open Science / Research Data Management Guidance. Overview of obligations and recommendations for open science and RDM in Horizon Europe. Available from: https://rea.ec.europa.eu/openscience_en (Accessed: 26/08/2025). 3 A list of relevant abbreviations and a brief explanation of the terms is provided at the end of the document in the chapter Abbreviations and Glossary. 4 Act No. 130/2002 Coll., on the Support of Research, Development and Innovation from Public Funds (consolidated protection of trade secrets, security or personal data. The Ministry of Education, Youth and Sports (MEYS) supports implementing these principles through the Open Science Framework, which formulates recommendations and standards for institutions and research teams, and is also implemented under the conditions of the Jan Amos Comenius Operational Programme (OP JAK). 6 This framework is based on international documents, particularly the 2021 UNESCO Recommendation on Open Science, which defines Open Science as an inclusive concept linking open access to scientific publications, open data, public engagement and international cooperation. Given the above context, research and student data management have become key topics in the university environment in recent years, not only in connection with open science and digitisation of research, but also regarding new legislative and grant requirements. The Czech University of Life Sciences Prague (CZU) also sees this area as a strategic priority. As part of the project "Increasing the Quality and Efficiency of the Educational Process and version). Definition of "research data" (Section 2 (2) (o), obligation to provide information on results and research data through the R&D&I IS (Section 12), access to research data and exceptions (Section 12a). Available from: https://www.zakonyprolidi.cz/cs/2002-130 (Accessed: 26/08/2025). 5 Research, Development and Innovation Council: Research, Development and Innovation Information System (R&D&I) — basic information and legislation. Available from: https://vyzkum.gov.cz/FrontClanek.aspx?idsekce=610 (Visited: 26/08/2025). Additionally: Public interface of the R&D&I IS (CEA, CEP, RIV, VES): https://www.isvavai.cz (Accessed: 26.08.2025). 6 OP JAK / MEYS: Handbook of Open Science Practices, version 1.1, effective from 28.06.2023. Available from: https://opjak.cz/prirucka-postupu-otevrene-vedy/ (Accessed: 26/08/2025). Optional: OP JAK / MEYS: Open Science – Documents (Handbook, Data Management Plan, etc.)tag. Available from: https://opjak.cz/dokumenty/otevrena-veda/ (Accessed: 26/08/2025).
6 Strategic Management at CZU", the Data Management Competence Centre (DMCC) was established, whose main task is to develop CZU system approaches and rules in the field of research and student data management, to provide methodological support and to help students (undergraduate and postgraduate), scientific and academic staff and employees in orienting themselves in the new rules, tools and procedures. In addition to methodological and consulting support, DMCC is also involved in creating rules and standards for data management, building the university's data infrastructure, and coordinating a network of experts at faculties and education in data management. It provides awareness and dissemination of FAIR principles, open access and emphasis on data quality and security. Its goal is to create an environment where data – whether from research or teaching – is considered a valuable output of academic activity that needs to be handled responsibly and effectively. To ensure that the proposed measures are in line with the real needs of the university community, it was necessary first to map the baseline in detail – i.e. to find out what types of data staff and students work with, what rules, tools and support they have at their disposal, what barriers they feel and what they might need to change. So far, no systematic analysis of this area has been conducted at CZU. A questionnaire survey was conducted at CZU in 2025 to gain practical insight into current practice. The survey aimed to provide adequate recommendations based on the answers and formulate future steps and measures in the university's strategic data management field.
7 Survey Overview and Methodology In April 2025, a questionnaire survey was conducted at the CZU. The aim was to obtain specific information about everyday work with research and student data, identify commonly used tools and procedures, and determine the level of institutional support and requirements for future services. It was the first-ever comprehensive probe into this area across the university. The questionnaire was created in cooperation with DMCC and faculty data methodologists, who participated in commenting on the content, testing and distributing the survey. Data was collected via Microsoft Forms from April 8 to May 4, 2025. The questionnaire was not anonymous – the identification of respondents was intended solely to supplement or clarify the answers. All data was processed exclusively by DMCC employees, and the results are presented in an aggregated and anonymised form. The questionnaire was structured into three thematic areas and included 58 questions, some conditional on previous answers and further branched out according to context. It combined closed-ended questions with various answers, optional comments, and several entirely open-ended questions that allowed respondents to share specific experiences, opinions, and suggestions. The thematic structure of the questionnaire (apart from the initial identification) was based on the logic of the life cycle of research data – i.e. from planning and data creation, through their storage and management, to sharing, archiving and reuse. The questionnaire also considered the institutional framework and available support in data management. The closing section of the questionnaire was devoted to student data and related practice in using student data. A total of 227 respondents 7 from all six faculties of CZU and the Institute of Education and Counselling participated in the survey. The survey was distributed to scientific, academic, and technical-economic workers and doctoral students – i.e. groups that actively work with data. Although the survey is not quantitatively representative, and it is impossible to generalise the results, it provides a valuable qualitative overview of current practice, habits, and barriers employees and doctoral students perceive when working with data. The findings serve as a basis for designing specific measures, developing university infrastructure and targeted support services in data management. 7 The text uses generic masculinity as a linguistic means to simplify expression; this designation refers to people of all gender identities. The research was attended by scientific, academic and technical-economic workers and doctoral students, as well as doctoral students.
8 Summarised Results of The Data Survey The following sections of the document focus on the individual questions of the questionnaire. Each section is accompanied by a commentary that helps put the findings into context and identify the primary areas of interest. The aim is not only to record what respondents answered, but also to understand what their answers say about the state and needs of the university in research and student data management. Identification questions At which faculty/institute do you primarily work? Figure 1: Respondents' professional background Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k The highest proportion of responses came from the Faculty of Forestry and Wood Sciences/FLD (25%, 57), followed by the Faculty of Engineering/TF (22%, 50). Approximately one sixth of the participants work at the Faculty of Environmental Sciences/FZP (16%, 37), while the Faculty of Agrobiology, Food and Natural Resources/FAPPZ is represented by 34 respondents (15%). The Faculty of Tropical AgriSciences/FTZ is represented by 29 people (13%), the Faculty of Economics and Management/FEM 19 (8%) and the Institute of Education and Counselling/IEC 3 (1%). Two respondents stated their competence at several faculties (a combination of FZP, TF & FTZ, FZP). Although the primary goal of the survey was to obtain an overall overview of access to research and student data across the entire university, it is also helpful to know the professional background of the respondents. This can help identify field specifics, but does not serve as a basis for comparing faculties. The highest proportion of responses came from the Faculty of Forestry and Wood Sciences (25%, 57), followed by the Faculty of Engineering (22%, 50). Approximately one sixth of the participants work at the Faculty of Environmental Sciences (16%, 37), while the Faculty of Agrobiology, Food and Natural Resources is represented by 15% (34) of respondents. The Faculty of Tropical AgriSciences is represented by 13% (29) of people, the Faculty of Economics and Management by 8% (19) and the Institute of Education and Counselling by 1% (3). Two respondents stated their competence at several faculties (a combination of FZP, TF & FTZ, FZP). The total number of responses is 229. What is your primary job position? Most respondents are research, academic or technical-economic professionals (77%, 176), and about a fifth of respondents are PhD students (22%, 49). In the "Other" category, the technician or data methodologist role was listed as the primary job position. These roles are not numerically significantly represented (1%, 2). 3 19 29 34 37 50 57 IVP/IEC PEF/FEM FTZ FAPPZ F P TF FLD Number of responses Professional workplace
9 Figure 2: Respondents' job positions Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k What is your primary field of work (according to OECD classification)? The internationally recognized OECD classification (OECD Ford – Fields of Research and Development (Frascati Manual 2015) was used to classify the respondents' field of expertise, distinguishing six areas: Natural Figure 3: Representation of industries (according to OECD FORD) Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k Sciences; Engineering and Technology; Medical and Health Sciences; Agricultural and Veterinary Sciences; Social Sciences; Humanities and Arts), which allows for later comparison across disciplines and between institutions. Field representation can help identify specific challenges in data management in individual scientific domains. Most respondents work in the Natural Sciences (47%, 106), followed by Engineering and Technology (31%, 71) and Agricultural and Veterinary sciences (25%, 56). Social Sciences (9%, 21), Multidisciplinary research (8%, 18) and other categories are less represented. Within the "Other" category, respondents mentioned, for example, Food Science, Geoinformatics and Remote Sensing in the context of forestry, or work at the Ministry of the Environment (Department of Circular Economy and Waste). Humanities and Arts (1%, 3) and Medical and Health Sciences (1%, 2) are minimally represented. 2 3 4 18 21 56 71 106 Medical and Health Sciences Humanities and the Arts Other Multidisciplinary Research Social Sciences Agricultural and veterinary sciences Engineering and Technology Natural Sciences Number of responses Primary field of work (acc. to OECD ORD) Doctoral student 22% Research/acade mic/technical staff 77% Other 1% Primary ob position
16 Figure 10: Ways to ensure the quality of the collected data Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k “The accuracy of data collection is to some extent related to a given country, a given culture, a given university, faculty, team... the scientist himself decides (based on the clearest conscience) what can bring such noise that the data would not be credible, when working with living organisms, it is impossible to capture all factors...” “I don't do specific quality control because I use partially processed data.” “Due to the nature of my research, there is no need for (quality control).: How do you secure your data during the collection process? The most common way to secure data is to restrict access to authorised persons (45%, 103). About four out of ten respondents (39%, 89) said they do not implement any specific measures, which is worrying regarding data security and compliance with legislation. A third of respondents (31%, 71) store data on secure storage such as institutional servers. Only 6% of respondents (13) anonymised or pseudonymised data at the time of collection, which was in line with the risk minimisation principle. Other measures, such as instant data encryption, are reported by only 3% (7) of respondents, signalling a low level of technicalprotection for some surveys. In the "Other" category, respondents mention that they do not consider any measures necessary (they do not work with sensitive data), the use of external storage (e.g. physical disks) and the case where the respondent does not actively work with the data himself. 3 25 55 79 115 174 Other I do not carry out specific quality control Data collection in compliance with ethical guidelines and legislation Regular data checking and cleaning Calibrated/validated measuring instruments/tools Standardised data collection methodologies within the discipline Number of responses Methods of ensuring quality of collected data
17 Figure 11: Methods of data security during data collection Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k “I think the measure (to secure data) does not apply to entomological research.” “I do not take any special measures (to secure data). I do not consider measuring technical data of physical quantities to be sensitive data.” “The data does not contain sensitive data and is therefore not treated by any of the above procedures.” The answers of some respondents show that the absence of sensitive or personal data is perceived as a reason for not implementing any security. However, this assumption is problematic. Data security isn't designed exclusively for working with sensitive data – it's an essential part of responsible research practice that protects the integrity, trustworthiness, and availability of data in any form. Even so-called "non-sensitive" data deserves basic standards of protection – e.g. limited access, secure storage or backup. These steps serve not only to prevent loss or unauthorised access, but also to protect institutional property and meet grant or legislative requirements. Have you ever used research data from other researchers, institutions, etc.? + If so, how did you access these data? + Why have you not yet used data from other researchers, institutions, etc.? 17 More than half of the respondents (58%, 131) have used data from other researchers or institutions. On the other hand, 42% (96) have not yet worked with such data. Figure 12: Sources of research data used from other scientists Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k The most common mode of accessing others’ data is direct provision upon request (48%, 109), followed by the use of public databases (29%, 65)—typically those of national or international institutions (e.g., Eurostat, the Czech Statistical Office, OECD, NASA Earthdata). A smaller share of respondents rely on general-purpose repositories such as Zenodo, Figshare, Dryad, or OSF (11%, 25), discipline-specific repositories (7%, 16)—e.g., HEPData, ioChem-BD, PANGAEA, CAIDA—or on purchasing data from commercial providers (10%, 23). In the “Other” category, respondents most often mentioned data provided as supplementary materials to publications or remotely sensed data. 7 7 13 71 89 103 Data are encrypted immediately Other Anonymization/ Pseudonymization Data are stored only on secure storage No specific measures taken Restricted access to data (authorised Number of responses Methods of data security during collection 2 16 23 25 65 109 Other Field-specific repository Purchased from commercial General repository Public databases Individiual request Number of responses Sources of used research data from other scientists
18 Figure 13: Reasons for not using data from other scientists Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k Among the main reasons for not using data from other sources are a preference for working with primary data the collected/ generated themselves—often influenced by concerns about the quality of external data (7%, 15) and limited awareness of suitable repositories (7%, 15). 24% (54) have not yet used data from other researchers, but do not rule out doing so in the future. 7% (15) reported that the data were unavailable or could not locate them, and 3% (6) expressed concerns about metadata quality. “My biggest fear is that shared data will be of inadequate quality. In my practice, I have often encountered data being altered and the like. For that reason, I do not trust others’ data. The mere presence of metadata does not mean they are not—pardon the expression—lies. I can vouch only for myself and for data produced under my control.” Do your research data contain sensitive information according to the GDPR? + How do you protect sensitive data during the collection phase? Figure 14: Sensitive information in research data Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k Most respondents (71%, 161) said their research data contained no sensitive information. However, about one in six (16%, 36) works with personal data – i.e., data subject to protection under the GDPR. Other sensitive information includes special categories of data (e.g. health, biometric) (7%, 1 13 14 15 20 36 161 Other Yes, data subjected to contractual restrictions Yes, geospatial data can identify individuals Yes, special categories of data I am not sure Yes, personal data No, no sensitive information Number of responses Sensitive information in research data 3 6 14 15 15 54 69 I don't collect/generate data I don't trust the metadata of available data Required data are not available, or I was unable to find them I have concerns about the quality of third-party data I don t know where to look for suitable data I haven't yet tried, but I may do so in the future I work exclusively with primary data I collect or generate myself Number of responses Reasons for not using other scientists data
19 15), geospatial data allowing identification of persons (6%, 14) and contractually protected data (e.g. trade secrets) (6%, 13). Notably, 9 % of respondents (20) are unsure whether their data contains sensitive information – indicating uncertainty about the obligations of the GDPR. One respondent in the "Other" category states that they do not collect, generate, or process data. Figure 15: Ways to protect data during data collection Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k One in three respondents protect their data when collecting it by limiting access (32%, 73), followed by requiring informed consent (20%, 45) and anonymising data at collection (12%, 27). The responses also show that even some researchers who do not consider their data to be sensitive still apply basic protection measures (23% of 161 people). One (<1%) respondent with a declared collection of sensitive data does not exercise any protection, which is problematic from the point of view of legal compliance and ethics. The answers show that some respondents do not have sufficient information about what is considered sensitive or protected data. This includes, for example, data that allow indirect identification of persons, location data (which, in combination with other data, lead to identification), or data protected by contract, not just by law. Conscious and unconscious underestimation of risks can have serious consequences – both legal and ethical. 3 8 10 27 45 73 135 I don't take any measures Data encryption Pseudonymisation Anonymisation during collection Informed consent from respondents Limited access (to authorized persons) I don't work with sensitive data Number of responses Methods of data protection during data collection
20 Data Processing Phase What types of data or data files do you work with during your research? Figure 16: Types of data or datasets Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k Nearly all respondents (90 %, 205) stated that they work with tabular data (e.g., CSV, XLSX, SPSS…). A significant share also processes texts and documents (82 %, 187), e.g., PDF, DOCX, TXT, XML…, images (70 %, 160), e.g., JPEG, PNG, TIFF…, and archive containers (e.g., ZIP, TAR, GZIP…, 41 %, 92). Approximately one third of respondents work with geographic data (e.g., SHP, GeoTIFF, NetCDF…, 33 %, 75), statistical data (e.g., DTA, SAV, SAS…, 33 %, 74), and video recordings (30 %, 69). One quarter of respondents use structured data and databases (25 %, 57), e.g., SQL, JSON, relational databases, cloud DB…, and source code and outputs from scientific software (23 %, 52). Other data types include audio recordings (17 %, 39) and executable files, configuration data (13 %, 29), and specific formats such as NetCDF, LAS/LAZ, RData, STEP, or ROCKY, which were listed in the responses under the “Other” category. One respondent (<1 %) does not personally collect or process data; another works only with textual documents such as journal articles. Roughly how much research data have you generated in the past year? Figure 17: Estimated data volume Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k A quarter of respondents (25%, 56) estimate they have worked with 1-20 GB of data in the last year. However, a considerable proportion (19%, 42) cannot estimate the volume of their 6 29 39 52 52 57 69 74 75 92 160 187 205 Other Executable files and configuration data Audio recordings Specific outputs from instruments or scientific software Source code Structured data and databases Videos Statistical data Geospatial data and coordinates Archive containers Images Texts and documents Tabular data Number of responses Types of data or data files 1 2 3 8 23 26 31 35 42 56 More than 500 TB I don't have el. data 101 500 TB 11 100 TB 1 10 TB 21 100 GB 101 GB 1 TB Less than 1 GB I don't know 1 20 GB Number of responses Estimated data volume
21 data, which may indicate a lack of awareness of the basic characteristics of their own data. This can lead to insufficient capacity planning, backups, and archiving. A further 15% (35) of respondents said they had worked with less than 1 GB of data last year. A volume between 21 and 100 GB is estimated by 11 % (26) of research participants, while 14 % (31) of respondents generate data in the 101 GB and 1 TB range. Approximately one in ten respondents (10%, 23) state a 1-10 TB volume. A volume between 11-100 TB was mentioned by 4% (8) of respondents, 101500 TB was indicated by 1% (3) of respondents, and more than 500 TB was reported by one (<1%) respondent. This is also essential information, as it is necessary to address infrastructure and performance requirements for such data volumes, e.g., data transfer, backup or long-term data retention. Two of the research participants said that they do not work with any electronic data. Does your research involve any ethical or legal issues? Figure 18: Specific information contained in surveys Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k More than half of the respondents (52%, 118) said their research does not contain any information requiring specific legal or ethical protection. Nevertheless, a considerable number of respondents are researching areas that carry legal or ethical risks. Approximately 22% (48) of researchers work with copyrighted works – i.e. outputs (texts, illustrations, photographs, software, etc.) that are subject to license terms and rules for further distribution. The same proportion (22%, 48) say that their research includes personal data of participants that is subject to protection under the GDPR. One in six respondents (17%, 38) work with human participants in research, even though personal data is not always collected. Other sensitive areas of research include intellectual property protected by licenses or patents (15%, 34), protected data associated with planned commercialization or contractual conditions (14%, 32), sensitive data in the sense of special categories under the GDPR (e.g. health status (2%, 4)), and information related to national security (1%, 3), the handling of which is governed by special legal regimes (e.g. Act No. 412/2005 Coll.). In the "Other" category, respondents reported, for example, geolocations of protected species or samples from animals, which may affect the availability of data or require a special regime. Researchers in natural sciences, agricultural and veterinary sciences, engineering, and technology mainly manage this information. The responses show that a significant proportion of researchers are not fully aware of the legal and ethical obligations associated with different types of data, especially in the areas of copyright, personal data protection, handling of data from human or animal subjects, and research outputs intended for commercialisation or restricted access. In particular, combining research with human participants and the lack of awareness of the need for special protection is risky in compliance with codes of ethics and legal regulations. 3 4 4 32 34 38 48 49 118 Information related to national security Sensitive data Other Other protected data Content protected by intellectual property rights Research involving human participants Personal data of research participants Copyrighted works Data don't require specific protection Number of responses Specific information in research
22 Where or how do you store active data, you are currently working with (editing, analysing, etc.)? Most respondents use personal and local storage when working with currently Figure 19: Types of Storage When Working with Live Data Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k processed data. The most common form is storage on portable media (66%, 149), followed by university computers (64%, 146) and personal computers (57%, 129). A relatively high proportion of research participants (56%, 128) also use the Office 365 cloud storage, which the university provides under an institutional license. Other cloud services include Google Drive (36%, 81). Less common, but also used storage facilities are local storage of devices or sensors generating data (22%, 50), physical data storage (e.g. printed documents, physical samples in the warehouse (17%, 38), commercial storage such as Dropbox or Amazon Cloud (14%, 32), university network storage with the possibility of backup (14%, 32), computing clusters or academic grids (e.g. MetaCentrum, CESNET – 6%, 14). Only 2 (<1%) respondents use the CZU Data Platform (DaMP), which is designed for the management and sharing of research data. The vast majority of respondents (91%, 206) stated that they combine multiple types of storage when working with live data, while only 9% (20) rely on a single storage method, of which 60% (12) of respondents store their data only on a computer (personal or university) and 1 (<1%) respondent only on portable media (USB, external HDD, CD/DVD...). One (<1%) respondent said they were unsure. It follows that a decentralised approach to the storage of live data prevails. At the same time, the predominance of portable media and personal devices poses a significant risk in data security and protection (loss, access by unauthorised persons), insufficient backup and the risk of noncompliance with institutional information security policy. How do you version your data? How do you respond to changes in the content or form of research data (input, interim or outcome)? Figure 20: Data versioning methods Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k 3 13 24 41 50 155 Other SW with automatic versioning Version control tools Versioning in cloud services I don't version datasets I manually create copies of files Number of responses Methods of data versioning 1 2 4 14 32 32 38 50 81 128 129 146 149 I don't know CZU Data-Platform Other Computing cluster or academic grid University network storage Dropbox, Amazon Cloud and other commercial storage Physical storage Local storage of a device Google Drive Office 365 Personal computer University computer Portable media Number of responses Storage types when working with live data
23 The results show that the most common data versioning strategy is manual file copying, which was reported by 68% (155) of respondents. While this practice allows for some control over change, it also carries significant risks of confusion, inefficiency, and error, especially when working in a team or during long-term research. More than a fifth of respondents (22%, 50) do not create data versions at all; they only work with the current file version, with any previous changes being overwritten without possibly tracing them back. This approach is considered a high risk of information loss, undocumented development, or misinterpretation of results from the point of view of research data management. About one in five respondents (18%, 41) use the versioning features of cloud storage such as OneDrive or Google Drive, which can be partly effective, especially for individual work, but less for systematic team collaboration. Only a small proportion of respondents (11%, 24) use specialised versioning tools (e.g. Git, GitHub, GitLab, Bitbucket...), and only 6% (13) use software with automatic versioning. In the "Other" category, tools such as pg_dump or cron were listed, pointing to the presence of technically proficient respondents, especially in computer science or database technologies. Do you create metadata or structured documentation for your data? + Do you use a recognised metadata standard or schema to describe your data? + Which metadata standard do you use? Over half of the respondents (52%, 118) do not create metadata for their research. Around one third (33%, 74) said that metadata is created occasionally, while only 7% (16) of respondents always create metadata. Eight percent (19) of respondents said they were unsure, which may indicate a lack of awareness of what metadata represents and its role in documenting the data. Regarding metadata standards, 61 % (67) of respondents state that they do not use such a standard to describe their data, and 35 % (38) are unsure whether they use it. Only 4 % (4) of respondents reported specific uses of established metadata schemes, while Darwin Core (DwC) and Dublin Core were mentioned, for example, to describe biological data. The findings show a low level of awareness and practice in the field of metadata creation and the use of established field-specific standards. Low awareness in this area can significantly limit the fulfilment of FAIR principles (especially data interoperability) and their future usability and sharing. Figure 21: Creating metadata as part of documentation Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k How do you anonymise sensitive data before further processing? + Do you use specialised tools or software for data anonymisation? Almost two-thirds of respondents (65%, 148) said they do not work with sensitive data, and another 8% (19) believe that anonymisation is unnecessary in their research. Although in question 16. "Does your research data contain sensitive information according to the GDPR?" 8 respondents answered that they come into contact with sensitive information as part of the research, in question 26. "How do you ensure the anonymisation of sensitive data before its further processing?" stated that they do not work with sensitive data or that anonymisation is unnecessary. On the contrary, around 15 % (33) of respondents use pseudonymisation, while 10 % use complete anonymisation (23). A comparable proportion (10%, 22) ensures anonymity through data aggregation. Only 4% (9) of respondents say they do not anonymise because they fear losing key information. One (<1%) respondent stated that they were not sure whether this was necessary for their data, but that they did not anonymise. Another do not process research data. Yes 40% No 52% I don't know 8% Creation of metadata as part of documentation
24 Figure 22: Ensuring anonymisation before further data processing Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k 79% (178) of respondents do not perform any anonymisation of their data, even though some respondents who indicated in Question 16 that they work with sensitive data state here that they either do not anonymise or do not consider anonymisation necessary. This inconsistency may indicate a misunderstanding of the GDPR definition of sensitive data, which is further supported by the fact that 4% (9) of respondents are unsure about the appropriate anonymisation tool. Manual anonymisation is performed by 14% (33) of respondents, while automated tools are used by only 3% (6). The results suggest that most people working with sensitive data rely on manual anonymisation or none at all. Figure 23: Use of anonymisation tools Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k Yes, automated anonymisation tools I don t know which tools to use No, I anonymise data manually 14% No, I don't perform anonymisation 79% Use of tools for anonymi ation 9 19 22 23 33 148 I don't anonymise because it would lead to loss of information Anonymisation isn't necessary Data aggregation Full anonymisation Pseudonymisation I don't work with sensitive data Number of responses Ensuring anonymi ation before further data processing
25 Data Analysis Phase What tools do you use for data processing and computation? Figure 24: Calculation and data processing tools used Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k The most commonly used tool for calculating and processing research data is a spreadsheet – e.g., Microsoft Excel, used by most respondents (86%, 196). The second most frequently used tool is statistical software (e.g. SPSS, STATA, Statistica, SAS, Gretl...), which is used by 45% (103) of participants. The R programming language, which is used by 37% (84) of respondents, and to a lesser extent Python (15%, 35), also has a significant presence. Other tools mentioned include integrated measuring instrument software (16%, 37), Power BI (7%, 17), Matlab (5%, 11), MathCad (2%, 4), GIS (2%, 4), SciLab (1%, 3), and minority (<1%) use of other specialized tools such as C++, Solvers, Canoco, .NET, Fortran 90, Bash, Orange Data Mining, Jupyter, ChatGPT, Ansys, Grapher. One (<1%) respondent creates software tools for calculations and data processing. Interestingly, 5 % (12) of respondents stated that they perform calculations manually, and 2 % (5) do not perform any calculations. In connection with the question mentioned above, we provide a brief overview of centrally procured software and the method of its availability at CZU (as of 26/08/2025). The high proportion of spreadsheet users indicates the need for education in more advanced data analysis tools, especially in statistical software and programming languages that are more in line with the principles of reproducible research. 5 11 12 17 18 35 37 84 103 196 I don t use calculations Matlab I calculate manually Power BI Other Python SW of measuring device Program. language R Statistical software Spreadsheet software Number of responses Used tools for calculations and data processing Centrally acquired software by the Department of Information and Communication Technologies: • Statistica (download https://czuvpraze.sharepoint.com/sites/S W-DOWNLOAD) • SAS (download: https://czuvpraze.sharepoint.com/sites/S W-DOWNLOAD) • IBM SPSS Statistics/IBM SPSS Modeler (download: https://czuvpraze.sharepoint.com/sites/S W-DOWNLOAD) • ArcGIS – information and how-to: https://gis.czu.cz/portal/apps/storymaps/ stories/2d8ba92885e84326b8674325d4d c6a27 • Autodesk (e.g. AutoCAD and others) – students handle it independently through Autodesk Education: https://www.autodesk.com/education/edu -software/overview • Microsoft applications (paid, such as MS Project, Visio, Access, etc.): For students, available after account verification in Azure for Students: https://azure.microsoft.com/cscz/free/students; employees, according to internally assigned licenses in Microsoft Entra ID (formerly Azure AD). • Installation (incl. Open-source software such as R, Python, QGIS): Installation occurs based on a request via HelpDesk; regular users do not have administrator privileges by default.
32 Figure 34: Ways of sharing inactive data outside the CZU Source: DMCC, 2025, own processing based on data research, datahttps://osf.io/e3u6k Of those who share data, email is the most common channel (60%, 37%), followed by Office 365 (53%, 33) and Google Drive (42%, 26), followed by disposable services such as Uschovna or Letecka posta (19%, 12) and portable media (18%, 11). Dropbox (13%, 8), CESNET (10%, 6) or general repositories (11%, 7) have a relatively low representation. Fieldspecific repositories were not mentioned at all. As additional tools in addition to the options offered, individual respondents choose GitHub (<1%, 1) and commercial database storage (<1%, 1). The CZU data platform has not yet been used for this purpose, although one of the respondents declares interest in using it. “I would like to use the data platform at CZU (to share data within and outside CZU).” Why have you shared your data in the past? Researchers share their data primarily for pragmatic reasons, often directly related to a specific requirement or purpose. 60% (137) share data as a basis for a published article, 50% (113) based on an individual request of a colleague, 47% (106) share data for the purpose of collaborating with other researchers and 37% (83) for fulfilling project outputs. Other less common reasons include that 15 % (34) share data due to the terms and conditions of financial providers, 14 % (32) do so out of their own belief in the meaningfulness of sharing, and 11 % (25) see sharing as an opportunity to increase visibility and citation. Only 7% (15) consider the public interest (efficient use of public funds), and 22% (49) of respondents have not yet shared data at all. Figure 35: Past motivations for data sharing Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k 1 2 3 6 7 8 11 12 26 33 37 Local storage of a device Shared network drive Other CESNET General repository Dropbox Portable media File transfer services Google Drive Office 365 E-mail Number of responses ays of sharing inactive data outside CZU 15 25 32 34 49 83 106 113 137 To ensure effective use of public funding To increase the visibility and citation of my research I believe in the value of data sharing Due to funder requirements I have not shared data before To meet project deliverables To support collaboration with other researchers Upon individual request from a colleague As supplementary material for a published article Number of responses Previous motivations for data sharing
33 Do you use a licence when sharing your research data (datasets/code/etc.)? Figure 36: Licenses used when sharing data Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k The results show that 83% (188) of respondents do not use any licensing when sharing their research data. 20% (46) use some form of Creative Commons (CC-BY; CCBY-SA; CC-BY-NC; CC-BY-NC-SA; CC-BY-NCND), with CC-BY being the most common (54% of these 46 responses, i.e. 25 people). Respondents also reported the use of other open licenses such as Apache, BSD, GNU GPL (17%, 8), GNU Free Documentation License, Open Government License (13%, 6), Open Data Commons (11%, 5), Public Domain Dedication (7%, 3), and Eclipse Public, Mozilla Public, Zlib License (4%, 2). Some respondents said "Other" (11%, 5) or that they do not share data (11%, 5). 7% (3) of respondents are unsure about the license, 4% (2) state that they choose the license as needed or after agreement with partners. In which format do you usually share your data? Respondents most often share data in tabular format (81%, 185). Also as texts and documents (75%, 171), images (54%, 122), archive containers (ZIP, TAR, GZIP...) (30%, 69),statistical data (DTA, SAV, SAS...) (20 %,46), Geographic data/coordinates (SHP, GeoTIFF, NetCDF...) (18%, 41), video (15%, 35), structured data and databases (12%, 28), specific/proprietary formats (10%, 22), audio (8%, 19) and rarely as executables, configuration data, Rdata, netcdf, 3D point clouds (laz, las, etc.). 2% (5) of respondents do not share data at all. Figure 37: Shared data formats Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k 4 5 5 19 22 28 35 41 46 69 122 171 185 Executable files, I don't share data Other Audio Specific formats Structured data Video Geospatial data/coordinates Statistical data Archive containers Images Texts and documents Tables Number of responses Shared data formats
34 In your opinion, what are the benefits of sharing research data? Figure 38: Benefits of sharing research data Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k The responses show that researchers at CZU most often perceive the sharing of research data as a means of strengthening the credibility and transparency of research. This benefit was indicated by 56% (127) of respondents. An almost equally important advantage is the ability to use data that someone else has already collected (52%, 118), without the need for repeated collection. The third most frequently mentioned advantage is the support for replicability and robustness of research, which was mentioned by 50% (113) of respondents. A smaller but still significant share of researchers cite the acceleration of the research process in the field (45%, 102) as an advantage of sharing. At the same time, it is assumed that the availability of already processed data can lead to faster follow-up research and reduce duplication of efforts. Approximately two-fifths of respondents (39%, 88) see the higher citation of related publications as an advantage, i.e., making data available can increase the academic impact of a given study. A lower share of respondents (22%, 49) identify the fulfilment of the requirements of financial providers as a motivation for sharing. 22% (49) consider this important because shared data can be cited independently and thus represent a separate research output. On the other hand, 10% (22) of respondents do not see any benefits in data sharing. In open-ended responses, some researchers expressed concerns about the potential misuse of data, especially in a political or media context, or its poor quality. Others remind that sharing makes sense, especially for larger and more complex atasets, while for smaller, they perceive the added value as questionable. “The ability to easily work with data (metaanalyses), the spread of the subconscious about our research among other researchers... And of course, there are many disadvantages in using data with human error.” “It's nonsense, see the situation in the world, the use of data in countries that are considered hostile.” “(The advantage is) consultation and the possibility of further development.” “Sharing data only makes sense for large datasets; sharing small local datasets is a waste of time and energy.” 9 22 49 49 88 102 113 118 127 Other I don't see any benefits Possibility of independent data citation Meeting funder or publisher requirements Increased citation of the associated publication Acceleration of research Support for reproducibility and robustness Use of data collected by others Greater credibility and transparency Number of responses enefits of sharing research data
35 What prevents you from sharing your research data? Figure 39: Barriers to data sharing Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k About a third of the researchers surveyed (31%, 71) said nothing could prevent them from sharing data. Nevertheless, there are several challenges that the rest of the respondents perceive as obstacles. The most common are limited experience sharing data (30%, 69) and concerns about data misuse (30%, 66). There is also the absence of clear instructions from departments/faculties/university (18%, 41), technical or capacity limits in the preparation of metadata (17%, 39) or lack of motivation to share data (16%, 37). Other significant barriers include concerns about misinterpretation of data (16%, 36), ethical or legal uncertainty (15%, 35), and the complexity of preparing data for sharing (15%, 33). The sensitive nature of the data, the lack of support from the university, financial constraints, and difficulties in ensuring FAIR principles were also mentioned. Only 5% (11) of respondents declared their unequivocal disagreement with sharing their data. One of the respondents added that he had not yet generated any data to share as part of the research. “It was not 'fashionable' to take data as the most important article in the research, but the conclusions from the data, i.e. the actual interpretation of the data, always told us that what we see in the data is important, not the data itself.” "Some of the data cannot be easily shared due to the size, typically 3D scans that my coworkers and I send to each other on external drives.” 9 11 19 21 23 24 33 35 36 37 39 39 41 66 69 71 Other I don t want to share data I can't ensure FAIR principles Lack of financial resources Inadequate support from the university/workplace Data contain sensitive/legally protected information Preparing data for sharing is demanding Concerns about legal/ethical risks Concerns about misinterpretation Lack of motivation Technical barriers I can't prepare the metadata Missing clear guidance from university/workplace Concerns about of data misuse Lack of experience I perceive no barriers Number of responses arriers to data sharing
36 Would you like to share any comments regarding research data (e.g. collection, processing, sharing, storage, support needed...)? In the final open question, respondents shared several vital observations and suggestions for improving research data management at CZU. The most common was introducing a unified university repository that would enable efficient and secure storage, sharing and identification of data (e.g. through DOIs). Uploading data to this repository should be voluntary but professionally managed – including regular audits, clearly defined responsibilities, and user support. The need for personnel support, especially in the form of data stewards or data management coordinators – ideally at each faculty, also resonated significantly. They would provide practical support, conduct training, provide consultations and help navigate technical and legal issues related to data management. Respondents also call for standardisation of processes and infrastructure, including the ability to use ready-made templates for working with metadata, technical support for processing larger volumes of data, and unification of workflows in data analysis and archiving across faculties. One of the most common proposals was the introduction of a university research data management guideline that would set minimum requirements, clearly define responsibilities, and promote a uniform approach across the institution. At the same time, the guideline should reflect the different needs of individual fields. ”… Based on what I found out from the questionnaire, I feel that I should educate myself and improve my skills in this area, especially in project data management and storage.”
37 Student Data This part deals with the data that arise during the study, by which it is meant during the Bachelor's and Master's studies. These are data collected, generated, edited, etc. by students as part of work on theses, or as part of other projects, studies, or assignments. These can therefore be, for example, theses, software, scripts, codes, surveys, interviews, data from theses, fieldwork data, experiments, etc. Such data may later be used in teaching, research, or other academic activities. Do you encounter any types of student data in your work? + What types of student-generated data do you encounter? Figure 40: Types of Student Data Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k Two-thirds of respondents (67 %, 152) said they encounter student data in their activities. Most commonly, these are theses (79 %, 120), followed by underlying data for these theses (60 %, 91) and data generated within teaching activities, such as projects and assignments from classes (59 %, 89). Approximately half of the respondents who work with student data also deal with data from seminar papers and experimental data (72 respondents). Less frequently, they encounter records and protocols from laboratory classes (32 %, 49) or data from field exercises (26 %, 39). Do you use student data in follow-up work/research/projects or as a basis for publications, etc.? + How do you use student data? Of the 152 respondents who work with student data, two-thirds (64%, 97) confirm their use in follow-up activities. This data is most often used as part of a larger research or project (85%, 82 from this group), or as a basis for publications (78%, 76). Many respondents also use them to create teaching materials (38%, 37) or methodologies (32%, 31). Some respondents also mention the use of student data for validation (22%, 21), data analysis training (20%, 19) or in cooperation with practice and external institutions (11%, 11). Respondents mentioned other uses of student data, such as comparative studies or a pilot experiment. 1 39 49 72 72 89 91 120 Other Fieldwork data Laboratory Experimental data Seminar paper data Course project data Underlying data for theses Theses Number of responses Student data types
38 Figure 41: Student Data Usage Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k How do students share or submit their data to you? Figure 42: Ways to share student data Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k The predominant way of sharing or transmitting this data is e-mail communication. This method was indicated by 87% (84) of respondents. Platforms within the Office 365 suite (66%, 64) and physical portable media such as USB drives or external hard drives (56%) are also relatively strongly represented. Cloud services such as Google Drive (35%, 34) or tools such as Uschovna or Letecka posta (26%, 25) serve as additional channels. Only 13% (13) of respondents use Dropbox to transfer data. A smaller proportion of respondents mention alternative data transfer paths, e.g. through local sensor storage (3%, 6), CESNET systems (2%, 4), physical (printed) versions of documents (1%, 3) or shared network drives (<1%, 2). Only one (<1%) respondent reported using the CZU university data platform (<1%, 1). Individual answers in the "Other" category point to the use of university systems (UIS, Moodle) or situations where students do not pass on data because the researcher works with them directly within the team. 1 2 3 4 6 6 13 25 34 54 64 84 CZU Data-Platform Shared network drive Printed version CESNET Local storage of a device Other Dropbox File transfer services Google Drive Portable media Office 365 E-mail Number of responses ays to share student data 2 11 19 21 31 37 76 82 Other Sharing with external parties or partners Teaching and data analysis training As control/validation data Development of methodologies Development of teaching materials Creation of publications As part of larger research/project Number of responses Use of student data
39 How do you address legal or ethical issues when using student data? Figure 43: Legal and ethical aspects of the use of student data Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k Solving legal and ethical issues using student data among researchers at CZU is very diverse and often lacks systemic anchoring. Around a third of respondents (35%, 34) rely on verbal student consent, while 31% (30) only inform about further use of the data but do not require consent. Although moderately transparent, this procedure does not meet the standards of informed consent defined in research codes of ethics. Only 10% (10) of respondents require written consent from students, which is clearly the preferred form from the view of legal certainty and good practice. Another 7% (7) follow the internal guidelines of the faculty, which points to partial institutional support, but also to the absence of a uniform university guideline. The most problematic group is the 6% (6) of respondents who do not inform students about the further use of the data and do not require their consent. This procedure is not only ethically questionable, but also potentially legally challengeable. Some respondents in the "Other" category report that students are included in a broader research project or are automatically taken as coauthors. Others work only with publicly available data from final theses or modify the method of obtaining consent depending on the nature of the data. “After a personal consultation with the students, we agree on their data output; if they are used in an English-language article, we address authorship with the students (lead author versus co-author). The students’ level of awareness is, pardon the expression, at the faculty as well as across the whole of the Czech Republic, dismal; when you, for example, offer students co-authorship on an IF article, they are glad and agree, whereas abroad they would discuss with you the order of coauthorship, possibly the corresponding order if the data are theirs. Here (as elsewhere) we do not even have it determined who is the owner of the data—whether the student or the supervisor, who is (mostly) the author of the idea of the experiment…’’ “Students are informed in advance that their work is part of the project (they often receive a scholarship from the project), and we create the data (measure, evaluate in the field) together.” “Students usually work with data created under my supervision, with my intellectual property. It's not "their data". However, students are considered part of a team, and if there is a publication, they are in the author's team (positions according to ethical principles.'' 6 7 10 10 30 34 I don't inform students about further use I follow faculty guidelines Consent is provided in writing Other I inform students about further use but don't require consent Consent is provided verbally Number of responses Legal and ethical aspects of the use of student data
40 How do you anonymise student data before further use? Figure 44: Ways to anonymise student data before using it Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k The methods of anonymising student data before further use show significant variability among respondents, with some of the chosen procedures being contrary to the recommended ethical and legal standards. Researchers often anonymise only student data containing personal or sensitive information (31%, 30). While this practice is partially compliant with data protection requirements, it may be insufficient in cases where it is possible to identify a specific individual retrospectively from non-explicit data (so-called indirect identifiers). Worryingly, 28% (27) of respondents do not anonymise at all. Together with another 14% (14) who use data in its original form, this is more than twofifths of researchers who potentially violate the principles of privacy and expose students to the risk of misuse of their data. Only 13 % (13) of respondents stated that they always anonymise all student data, which is an ideal— albeit sometimes technically demanding—yet entirely safe approach. Only a small share of respondents use the method of pseudonymisation (7 %, 7), in which, for example, names are replaced with codes. Several researchers stated that anonymisation is unnecessary because they work exclusively with numerical values or cite the data as authorial outputs (e.g., in methodologies). “These are numerical laboratory data; there is no need for anonymisation.”’ “Since I use them for methodology, I quote where I got the information.” 6 7 13 14 27 30 Other Pseudonymization Anonymization of all student data I use original data I don't anonymise student data Anonymization of data containing personal or sensitive information Number of responses Methods of anonymi ing student data before their use
41 Are you interested in using student data? + What access methods to student data would you find useful? Figure 45: Preferred Options for Accessing Student Data Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k Of 130 respondents who do not currently work with student data, 24% (31 respondents) expressed interest in using them. A higher proportion (45%, 59) were unsure whether they would be interested in using student data, and 31% (40) said they were not. The largest group comprises those undecided on this issue, which may indicate a lack of information about the possibilities or benefits of such use. As for preferences in accessing student data (187 respondents answered), respondents would most often welcome sharing via Office 365 tools (61%, 114), especially via OneDrive, SharePoint or MS Teams. E-mail (44%, 82) and the CZU data platform (30%, 57) are also relatively common choices, confirming the earlier interest trend in centralised infrastructure. Other preferred tools include Google Drive (27%, 50) or portable media (20%, 38). A more significant proportion of respondents mentioned Dropbox (7%, 14%), CESNET (7%, 13) or local storage of the device/sensor (7%, 13). Only a few participants were unsure of their preference (3%, 6) or indicated another option (e.g. personal handover, UIS, Moodle). "I don't know any local storage from where it will be easy to download, and it will clearly specify what kind of data it is. They also need to be collected correctly; otherwise, the student data is useless.“ “I would like to use the data platform at CZU.” “I don't need it or it's by agreement. For example, a student gives me a sample from the field for identification.” Do you think it would be beneficial to have information about the data generated during the study? It can be stated that most respondents (64%, 146) consider it useful to have an overview of the data generated within the study, even if they do not currently use it. This result highlights the potential of student data as a valuable resource for the wider academic community. On the other hand, only a minority of respondents (7%, 15) do not consider this information beneficial. A considerable proportion of participants (29%, 66) chose the answer "I don't know", which may indicate that the use or existence of such data is not sufficiently known or grasped in the institutional framework. Figure 46: Evaluating the Benefits of Student Data Information Source: DMCC, 2025, own processing based on data research, data: https://osf.io/e3u6k 5 6 13 13 14 38 50 57 82 114 Other I don't know CESNET Local storage of a device Dropbox Portable media Google Drive CZU Data-Platform E-mail Office 365 Number of responses Preferred options for accessing student data Yes, it could be useful 64% I don't think it's useful 7% I'm not sure 29% Evaluation of the contribution of student data information
48 Strategic ramework The university's data governance guidelines should create a binding framework setting minimum standards across the institution. Faculties will be able to develop these standards with regard to the specifics of the field. Rules should be defined for secure, ethical and transparent data management, including encryption, versioning and long-term preservation. The strategic framework should take into account: ▪ DMP requirements in all projects ▪ Recommendations for safe data handling, selection of appropriate repositories, and data transfer ▪ Ethical and legal aspects, informed consents and procedures ▪ Documentation standards and metadata creation according to FAIR principles ▪ Support for open formats and interoperability Strategic measures should include the design of a university data platform that enables secure central data storage, automated backups, and versioning support. Administrative Support and Methodological Materials Administrative support should include a comprehensive set of documents, templates, and tools to simplify and automate data management processes. A key component is standardised templates tailored to different disciplines and types of research, supplemented by practical guidance on how to complete them. In cooperation with the legal department, templates of informed consents and consents to processing personal data should be prepared, ensuring the protection of sensitive data. To aid orientation in the issue of sensitive data and their identification, checklists and recommended measures for anonymisation and pseudonymisation should be available. Staffing DMCC should play a central role in providing personnel support for data management, together with faculty data methodologies. This support network should provide consultation, methodological support and feedback on data management issues. The common goal should be developing the university environment of data management and implementing the principles of open science into CZU's research practice. The recommendation also includes creating data steward positions at the level of individual pro ects, who should be responsible for regularly monitoring the DMP and providing specific support to the pro ect team Possible supporting methodological materials include: ▪ Informed consents, consents to the processing of personal data, including university forms for informed consent of students to the use of their data ▪ Templates for creating metadata according to industry standards ▪ Dataset description patterns for repositories and instructions for correct data citation ▪ Clear materials explaining the different types of licenses (e.g. Creative Commons) and their suitability, supplemented by decision trees for selection ▪ Recommended formats for exporting documents ▪ Practices for Secure Data Sharing ▪ Checklists for long-term data retention ▪ Clear instructions, templates and recommendations for the use of student data, including examples of good practice
49 regarding the partial specifics of the financial providers, industry and project goals. Infrastructure and Technical Tools The University should actively develop a data infrastructure that fully complies with the requirements for FAIR data management. This includes comprehensive support for sub ect and general repositories, ensuring the data is easily found and accessible to the broad scientific community. A key element should be implementing a robust backup system that minimises the risk of data loss and supports data versioning through user-friendly tools. Making anonymisation and data security tools available could contribute to protecting sensitive information and upholding ethical standards. Awareness-raising and The Development of an Open Science Culture Support for open science and data sharing is essential for the development of modern research and should therefore be actively promoted through awareness-raising and educational activities. The aim should be to create an environment where data sharing is perceived as a standard and beneficial part of research work, rather than as an additional burden. This includes not only disseminating awareness of the principles of open science but also providing practical tools and methodologies for the effective sharing and reuse of data. Integrating the principles of open science into project methodology and evaluation criteria can ensure systemic support and motivation for researchers. Possible activities in the field of infrastructure and technical tools: ▪ Support and recommendations for the use of disciplinary (e.g., disciplinespecific databases) and general repositories (e.g., Zenodo, Figshare) ▪ Implementation of data backup, e.g., according to the 3-2-1 rule—make three copies of the data, distribute them across two media (e.g., two external drives), and place 1 of them at a different location ▪ Tools for data versioning with a low entry barrier (e.g., plugins for Office 365, Git) ▪ Provision of tools for anonymisation and data security (e.g., pseudonymisation software, encryption tools) ▪ Design of a university data platform for secure central data storage Examples of activities in the field of education: ▪ Awareness-raising campaigns (e.g. workshops, seminars, leaflets) on the importance and benefits of open science and data sharing ▪ Informal events such as "data coffee", "open grill", etc. ▪ Case studies and examples of good practice in data management and sharing ▪ Conditional on funded projects on a data sharing plan ▪ Integration of open science principles into internal project schemes and evaluation criteria (e.g. criteria for evaluating grant applications)
List of Images Figure 1: Respondents' professional background .................................................................................... 8 Figure 2: Respondents' job positions......................................................................................................... 9 Figure 3: Representation of industries (according to OECD FORD) ................................................................. 9 Figure 4: Tools used to create a data management plan ....................................................................... 11 Figure 5: Rules for data management in the research team/workplace ............................................... 11 Figure 6: Specific assistance during research in the past ..................................................................... 12 Figure 7: Areas of preferred support ....................................................................................................... 13 Figure 8: Form of preferred support ........................................................................................................ 14 Figure 9: Methods of data collection or generation ............................................................................... 15 Figure 10: Ways to ensure the quality of the collected data .................................................................. 16 Figure 11: Methods of data security during data collection .................................................................. 17 Figure 12: Sources of research data used from other scientists .......................................................... 17 Figure 13: Reasons for not using data from other scientists ................................................................ 18 Figure 14: Sensitive information in research data .................................................................................. 18 Figure 15: Ways to protect data during data collection ......................................................................... 19 Figure 16: Types of data or datasets ....................................................................................................... 20 Figure 17: Estimated data volume ........................................................................................................... 20 Figure 18: Specific information contained in surveys ................................................................................... 21 Figure 19: Types of Storage When Working with Live Data ........................................................................... 22 Figure 20: Data versioning methods ........................................................................................................ 22 Figure 21: Creating metadata as part of documentation ....................................................................... 23 Figure 22: Ensuring anonymisation before further data processing ..................................................... 24 Figure 23: Use of anonymisation tools .................................................................................................... 24 Figure 24: Calculation and data processing tools used ......................................................................... 25 Figure 25: Preferred Data Visualisation Tools ........................................................................................ 26 Figure 26: Ways to verify data quality and integrity ................................................................................ 26 Figure 27: Final data store after the end of active work ......................................................................... 27 Figure 28: Preferences for data storage in CZU repositories ................................................................ 28 Figure 29: Repositories in use .................................................................................................................. 28 Figure 30: Storing metadata along with research data .......................................................................... 29 Figure 31: Ways of data sharing within CZU ........................................................................................... 30 Figure 32: Ways of sharing data outside the CZU .................................................................................. 30 Figure 33: Sharing of inactive data with researchers within CZU .......................................................... 31 Figure 34: Ways of sharing inactive data outside the CZU .................................................................... 32 Figure 35: Past motivations for data sharing .......................................................................................... 32 Figure 36: Licenses used when sharing data .......................................................................................... 33 Figure 37: Shared data formats ............................................................................................................... 33 Figure 38: Benefits of sharing research data .......................................................................................... 34 Figure 39: Barriers to data sharing .......................................................................................................... 35 Figure 40: Types of Student Data ............................................................................................................. 37 Figure 41: Student Data Usage ................................................................................................................ 38 Figure 42: Ways to share student data .................................................................................................... 38 Figure 43: Legal and ethical aspects of the use of student data ................................................................... 39 Figure 44: Ways to anonymise student data before using it .................................................................. 40 Figure 45: Preferred Options for Accessing Student Data ..................................................................... 41 Figure 46: Evaluating the Benefits of Student Data Information ........................................................... 41
51 Abbreviations and Glossary A Academic grid also called grid computing A system that interconnects large numbers of computers— typically in scientific or academic settings—to solve large-scale computational tasks collaboratively; e.g., MetaCentrum (Czechia), San Diego Supercomputer Centre (SDSC) EGI (European Grid Infrastructure), San Diego Supercomputer Centre (SDSC) AgEcon Search Agricultural and Applied Economics Search A free web repository that provides open access to full-text scholarly works in agricultural and applied economics Apache License An open-source software license created by the Apache Software Foundation (ASF) Arctic Data Center The primary repository for data and software for the Arctic section of the U.S. National Science Foundation’s Office of Polar Programs (NSF OPP) Arxiv An online repository of preprints of scientific articles, primarily in mathematics, physics, computer science, and biology B Bitbucket A platform for version control and team collaboration on code (Git), offering pull requests, code reviews, and integrations BSD License Berkeley Software Distribution License A permissive open-source license with minimal restrictions; allows modification, redistribution, and commercial use without any obligation to disclose source code (no copyleft) C CAIDA Center for Applied Internet Data Analysis A research organisation focused on studying internet architecture and behaviour. CD/DVD Compact Disc/Digital Versatile Disc Portable media (optical discs) for data storage CESNET Czech Education and Scientific NETwork The Czech research and education e-infrastructure providing computing capacity, data storage, and other services to universities and research institutions CGSpace A repository that preserves and provides access to research outputs and knowledge products from CGIAR (Consultative Group for International Agricultural Research) Cloud databases Databases operated/hosted in the cloud, e.g., Firebase, AWS RDS etc. Creative Commons A suite of public licenses that enable the legal sharing and reuse of copyrighted works
52 C Cron A time-based job scheduler (daemon) for planning and automating tasks on Unix-like systems CSO Czech Statistical Office The central government authority of the Czech Republic that collects, processes, and publishes statistical data on social, economic, demographic, and environmental developments .csv Comma Separated Values – a *.csv file A plain-text file format (*.csv) for storing tabular data, where values are separated by commas CUAHSI HydroShare Consortium of Universities for the Advancement of Hydrologic Science HydroShare A web repository and hydrologic information system operated by the Consortium of Universities for the Advancement of Hydrologic Science (CUAHSI) CZU Czech University of Life Sciences Prague A public university in Prague, founded in 1952, consisting of six faculties and one institute D DaMP CZU Data Platform The original internal CZU platform for storing and sharing research data; a legacy service currently operating in a limited mode Dataset A collection of data grouped for analysis or research .docx A *.docx file A text-document format natively used by Microsoft Word; supports text, formatting, images, and styles8 DMCC Data Management Competence Centre A specialised university unit at the CZU Library tasked with supporting and developing the research data management environment at CZU DMP Data Management Plan A plan describing how research data will be created, stored, shared, and archived DMPOnline An online tool for creating data management plans Dropbox A cloud storage and file-sharing service that lets users store and synchronise files across devices and share them with others Dublin Core A metadata standard for describing digital objects 8 Part of the Office Open XML (ISO/IEC 29500) open standard, but in practice its behavior depends on Microsoft Office’s implementation.
53 D Dryad An open repository for storing, sharing, and publishing research data—especially in biology, ecology, and medicine—often linked to journal publications .dta A *.dta file A data file format used by, e.g., the Stata software for statistical analyses DwC Darwin Core A metadata standard for describing and sharing biodiversity, biological, and ecological data E Eclipse Public An open-source software license Eurostat The statistical office of the European Union, providing statistics to EU institutions and member states F FAIR principles Principles for data management to make data Findable, Accessible, Interoperable, and Reusable APPZ Faculty of Agrobiology, Food and Natural Resources, Czech University of Life Sciences Prague A faculty focusing on agriculture, food science, ecology, and the natural sciences. FEM Faculty of Economics and Management, Czech University of Life Sciences Prague A faculty focused on economics, management, informatics, and systems engineering in agriculture and rural development Figshare An open, general-purpose repository for publishing research data and outputs LD Faculty of Forestry and Wood Sciences, Czech University of Life Sciences Prague A faculty specialising in forestry, wood science, game management, and landscape engineering. ORD Fields of Research and Development (OECD) An OECD classification of research & development fields (Frascati Manual 2015). It enables comparable categorisation across disciplines. It uses six broad areas: Natural Sciences; Engineering and Technology; Medical and Health Sciences; Agricultural and Veterinary Sciences; Social Sciences; and Humanities and the Arts OS Fields of Science and Technology (OECD) An earlier OECD field classification (2007). Functionally equivalent to today’s FORD; in newer OECD terminology, it has been replaced by “Fields of Research and Development.” rascati Manual Frascati Manual 2015: Guidelines for Collecting and Reporting Data on Research and Experimental Development
54 F The OECD methodology for collecting and reporting R&D data. It also defines the framework for the FORD field classification and related terminology TZ Faculty of Tropical AgriSciences, Czech University of Life Sciences Prague A faculty focusing on tropical and subtropical agriculture, development cooperation, and global nutrition ZP Faculty of Environmental Sciences, Czech University of Life Sciences Prague A faculty focused on environmental protection, landscape, ecology, and sustainable development. G GB Gigabyte A unit of data capacity; 1 GB = 1,000 MB. Roughly fits about 250 photos or one larger presentation with hundreds of slides GDPR General Data Protection Regulation The EU regulation on personal data protection is important when handling sensitive data GenBank A repository of nucleotide sequences—a database that stores DNA and RNA sequence data. GeoTIFF Geographic Tagged Image File Format – A *.tif / *.tiff file A raster image file format that, in addition to image data, contains geographic information (georeferencing) Git A free, open-source, distributed version control system designed for fast and efficient management of projects of any size GitHub A web platform that uses Git to host and manage software projects GitLab An alternative to GitHub, commonly used in companies, for hosting repositories and managing software projects GBIF Global Biodiversity Information Facility An international platform for sharing biodiversity data GNU GPL GNU General Public License A free-software license that lets users run, study, modify, and distribute software (a copyleft license) GNU FDL GNU Free Documentation Licence A license for freely available documentation—similar in spirit to the GPL but for manuals, textbooks, and other documents; it guarantees the freedom to copy and redistribute with or without modifications, commercially or non-commercially Google Drive Google Drive
55 G A Google service that allows storing, sharing, and synchronising files in the cloud .gzip / .gz A *.gzip / *.gz file A compression format commonly used on Unix/Linux. It uses the DEFLATE algorithm and compresses single files (for multiple files, they’re typically archived with tar first, producing .tar.gz). H HEPData High Energy Physics Database A repository for experimental high-energy-physics data HDD A fixed/internal hard drive or an external hard drive; a magnetic data-storage device I IEC Institute of Education and Counselling, Czech University of Life Sciences Prague An institute focused on lifelong learning, pedagogy, advisory services, and cooperation with professional practice. ioChem-BD Computational Chemistry Data Repository A platform for managing, archiving, and sharing data in computational chemistry and materials science J . peg Joint Photographic Experts Group – a *.jpeg file A lossy compression format for raster images, widely used for photography and web graphics; technically identical to *.jpg . son JavaScript Object Notation – a *.json fileA text format for structured data, commonly used for data exchange between systems and applications9 L .las/.la A *.las / *.laz file 3D formats for point clouds (e.g., from airborne LiDAR scanning); *.las is uncompressed, *.laz is compressed M MATLAB A software environment and programming language used primarily for numerical computing, modelling, simulation, data analysis, visualisation, and algorithm development MathCad Software for engineering, physics, and mathematics calculations and their documentation MetaVO MetaCentrum 9 he format is easy to process in common programming languages such as Python, R, or JavaScript.
56 M The Czech national grid e-infrastructure for computing (operated by CESNET) Blank measurements (using blank samples) Control measurements used to verify accuracy Moodle An open-source platform for building learning management systems and online courses (LMS) Mozilla Public An open-source software license MS Teams Microsoft’s communication and collaboration platform for teams N NASA Earthdata NASA’s portal for access to remotely sensed data about planet Earth NC I National Center for Biotechnology Information The U.S. National Center for Biotechnology Information that operates public databases of genetic, genomic, and biomedical data (e.g., GenBank, Gene Expression Omnibus, PubMed) NetCDF Network Common Data Form A file format (and data model) for scientific multidimensional data, such as meteorological or climate data O OneDrive Microsoft’s cloud service that lets you store, share, and access files from any device OGL Open Government Licence A license that allows the public to freely use and share information created by the public sector—such as data, text, images, and other content OECD Organisation for Economic Co-operation and Development An international organisation bringing together 38 economically advanced member countries Open Data Commons A set of legal tools and licenses designed for publishing, sharing, and using open data—especially for databases Open Formats Data formats with publicly available technical specifications that allow reading, writing, and processing without dependence on a specific vendor; support long-term accessibility. Examples: *.csv, *.xml, *.json, *.ods (LibreOffice), *.svg, *.txt10 Open Source Software whose source code is openly available and may be used, modified, and distributed OSF Open Science Framework 10 Although formats such as .docx and .xlsx have been standardized as open formats (ISO/IEC 29500), their full functionality is often available only in Microsoft Office. Consequently, they are sometimes regarded in practice as partly proprietary.
57 O An online platform for managing research projects and supporting open research practices P PANGEA A repository for geoscience/earth & environmental data PC Personal computer A personal computer .pdf Portable Document Format – A *.pdf file A universal document format by Adobe that preserves text, images, and layout across devices and systems Personal Data Any information relating to an identified or identifiable natural person (e.g., name, contact details, photos, education, family, employment, etc.) pg_dump A utility for backing up a PostgreSQL database .png Portable Network Graphic – A *.png file A lossless raster image format suited to graphics with transparency and sharp edges (e.g., logos, icons, charts) PostgreSQL An advanced open-source object-relational database management system (ORDBMS) Power BI Power Business Intelligence Microsoft’s business analytics platform that lets you visualise and share data Portable Media Devices or media used to store and transfer data between computers; typically USB flash drives, external hard drives (HDD/SSD), SD cards, or CDs/DVDs Proprietary / Vendorspecific Formats Data formats whose specifications are secret or restricted by the owner (e.g., a company). Data in such formats can usually be read and written only with the vendor’s software. Examples: .doc (Microsoft Word), .psd (Adobe Photoshop), .dwg (AutoCAD), .sav (SPSS), .ai (Adobe Illustrator), .accdb (MS Access database), etc.11 Protocol Exchange The original platform for sharing research protocols under the Nature brand (part of Springer Nature). Since 2024, its content has been migrating to the unified protocols.io platform, which is now also run by Springer Nature Protocols An open repository for sharing research protocols, operated by Springer Nature 11 Some formats may have a partially open specification, but because they depend on a specific implementation, they are still regarded as proprietary in practice.