scieee AI-readable full text Open interactive document viewer

DMP: Genre Mobility Analysis: Impact on Reception for Specialized Actors switching Genras

Dancea, Vlad

Abstract

Data Management Plan (DMP) for the project "Genre Mobility Analysis: Impact on Critical Reception for Specialized Actors". This document outlines the data management strategy for the experiment, including: Handling of input data from IMDb and Kaggle. Production of aggregated analysis results and source code. Licensing, storage, and long-term preservation strategies. The actual datasets and code described in this plan are deposited in the TU Wien Research Data Repository.

Full text

Data management plan (DMP) DMP: Genre Mobility Analysis: Impact on Reception for Specialized Actors switching Genras GMA Version Effective date Description of document/changes 1.0 28/11/2025 First version of the DMP – created for the start of the project Level of distribution This DMP is licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0). It is publicly available under https://doi.org/10.70124/81hv1-h4r21 2 GMA DMP version 1.0 Project details Project Coordinator Principal Investigator Contact person (responsible for data management and DMP) Vlad Dancea, [email protected]en.ac.at Contributors Vlad Dancea, [email protected]en.ac.at Start date 2025-10-01 End date 2026-01-31 Funder Funding programme, grant number Internal project number List of acronyms DMP data management plan RDM research data management … … … … … … … … … … … … GMA DMP version 1.0 3 Content INHALTSVERZEICHNIS" You$are$almost$there!$Error!%Bookmark%not%defined. INTRODUCTION) 4 Science$Europe$practical$guide,$FAIR$data$ 4 Relevant$Policies$and$Guidelines$ 4 1. DATA)DESCRIPTION) 5 1a Lists$of$datasets$that$will$be$reused$or$produced$ 5 1b Data$generation$and$reuse$ 6 2. DOCUMENTATION)AND)DATA)QUALITY) 6 2a Data$organisation,$metadata$and$documentation$ 6 2b Data$quality$control$ 7 3. STORAGE)AND)BACKUP)DURING)RESEARCH)PROCESS) 7 3a Storage$and$backup$facilities$ 7 3b Data$security$and$protection$of$sensitive$data$ 7 4. LEGAL)AND)ETHICAL)REQUIREMENTS) 7 4a Personal$data$ 7 4b Intellectual$property$rights$and$rights$of$use$ 8 4c Ethical$issues$ 8 5. DATA)SHARING)AND)LONG-TERM)PRESERVATION) 8 5a Data$publication$and$access$conditions$ 8 5b Long-term$preservation$and$deletion$of$data$ 9 6. RDM)RESPONSIBILITIES)AND)RESOURCES) 9 6a RDM-roles$and$responsibilities$ 9 6b Resources$ 9 ! 4 GMA DMP version 1.0 Introduction Science Europe practical guide, FAIR data A DMP is a structured document that keeps record of what research data is created and what happens to that data during and after a project. It helps with planning the research process and defining responsibilities in a research project involving several researchers or institutions. For writing this DMP, we followed the recommendations of Science Europe as they reflect the guidelines agreed upon by the major funders in Europe. To make our data FAIR, they generally will be treated according to the following criteria: § We will make our data findable, by uploading it to a data repository that provides a persistent identifier and adding relevant metadata. § We will make our data accessible by providing open access to data, wherever possible. In cases, where open access is not possible, we will provide meaningful metadata plus contact information for access requests. § We will make our data interoperable by providing and describing data in a way that is common within our domain by using the same file formats, schemas and vocabularies. We will provide good documentation for all our datasets. § We will make our data reusable by adding metadata and comprehensive Readme files to all published datasets. The descriptions include details on the methodology used, analytical and procedural information. In case of publication, licenses for code and data will always be assigned and clearly marked. Relevant Policies and Guidelines § European Commission’s document on Ethics and Data Protection: https://ec.europa.eu/info/funding-tenders/opportunities/docs/20212027/horizon/guidance/ethics-and-data-protection_he_en.pdf § Other (e.g. from a project partner) GMA DMP version 1.0 5 1. Data description 1a Lists of datasets that will be reused or produced Produced datasets dataset ID title type format estimated volume contains sensitive data P1 Aggregated Genre Transitions Data Structured text csv 100 - 1000 MB no P2 Analysis Code Source code py 100 - 1000 MB no P3 Result chart Images png 100 - 1000 MB no P4 IMDb NonCommercial Datasets Structured text tsv 100 - 1000 MB no P5 Rotten Tomatoes movies and critic reviews dataset Structured text csv 1 - 5 GB no Description for "Aggregated Genre Transitions Data": Machine-actionable table containing the calculated score deltas for actors moving between genres. Description for "Analysis Code": Python scripts used for data cleaning, merging and chart generation. Description for "Result chart": Resulting visualization chart 6 GMA DMP version 1.0 Reused datasets dataset ID title type format estimated volume contains sensitive data R1 IMDb NonCommercial Datasets Structured text tsv 100 - 1000 MB no R2 Rotten Tomatoes movies and critic reviews dataset Structured text csv 1 - 5 GB no Description for "IMDb Non-Commercial Datasets": Machine-actionable tables containing the raw data. Description for "Rotten Tomatoes movies and critic reviews dataset": Machine-actionable table containing the raw data. 1b Data generation and reuse Methods and software used for data generation and reuse 2. Documentation and data quality 2a Data organisation, metadata and documentation The project follows a standardized directory structure to ensure reproducibility: - src/ for source code - data/ (not synced) for input files - results/ for generated outputs - docs/ for documentation. File Naming: - Files will follow the convention YYYYMMDD_Project_Content_Version.ext - (e.g., 20251130_GenreMobility_Transitions_v1.0.csv) Versioning: - Source code is version-controlled using Git. - The final published dataset will use Semantic Versioning (v1.0.0) Comprehensive metadata will be provided to ensure FAIR, including: - Descriptive: Title, Description, Creator, Publication Date. - Administrative: Rights. - Keywords: "Entertainment Analytics", "IMDb", "Rotten Tomatoes". - Persistent Identifier: A DOI will be automatically assigned by the repository. This will help others to identify, discover and reuse our data. Data quality is controlled programmatically via the Python processing pipeline: GMA DMP version 1.0 7 - Filtering: Records with missing critical data are automatically excluded 2b Data quality control The following data quality checks will be done: standardised data capture, data entry validation and representation with controlled vocabularies. 3. Storage and backup during research process 3a Storage and backup facilities For the duration of the project, storage and backup of data will be ensured by Vlad Dancea (acting as the person responsible for data management and DMP) in cooperation with the system operator. External and internal storage options will be used. R2 (Rotten Tomatoes movies and critic reviews dataset), P1 (Aggregated Genre Transitions Data), P2 (Analysis Code), P3 (Result chart), R1 (IMDb Non-Commercial Datasets) will be stored on Local Laptop. P1 (Aggregated Genre Transitions Data), P2 (Analysis Code), P3 (Result chart) will be stored on Institutional Cloud: Institutional Cloud is your institutions cloud service. It runs on your IT department's servers and has every standard cloud functionality. 3b Data security and protection of sensitive data We pay strict attention to compliance with the relevant institutional and national data protection policies listed in the introduction of this document. At this stage, it is not foreseen to process any sensitive data in the project. If this changes, advice will be sought from the data protection specialist at our institution, and the DMP will be updated. Access to data during research: dataset ID selected project members all other project members the public P1 writing reading only no access P2 writing reading only no access P3 writing reading only no access R1 writing reading only no access R2 writing reading only no access 4. Legal and ethical requirements 4a Personal data At this stage, it is not foreseen to process any personal data in the project. If this changes, advice will be sought from the data protection specialist at our institution, and the DMP will be updated. 8 GMA DMP version 1.0 4b Intellectual property rights and rights of use Information regarding rights and control of access to data has yet to be identified. 4c Ethical issues No particular ethical issue is foreseen with the data to be used or produced by the project. This section will be updated if issues arise. 5. Data sharing and long-term preservation 5a Data publication and access conditions As far as possible, obtained datasets will be published in repositories. Details on access conditions, reuse licenses, reasons for restrictions, etc. are collected in the table below. dataset ID access conditions estimated publication date location for publication (repository) PID license P1 Open 2026-01-31 TU Wien Research Data https://doi.o rg/10.7012 4/81hv1h4r21 CC-BY-4.0 P2 Open 2026-01-31 TU Wien Research Data https://doi.o rg/10.7012 4/81hv1h4r21 MIT P3 Open 2026-01-31 TU Wien Research Data https://doi.o rg/10.7012 4/81hv1h4r21 CC-BY-4.0 R1 Closed R2 Closed Closed Dataset Reasons: IMDb strictly says "must not be altered/republished". Kaggle dataset is already online. Repository description: TU Wien Research Data is an institutional repository of TU Wien to enable storing, sharing and publishing of digital objects, in particular research data. It facilitates the funders' requirements for open access to research data and the FAIR principles by making research output findable, accessible, interoperable, and reusable. A DOI is assigned to each dataset published in TU Wien Research Data. This service is developed by the TU Wien Center for Research Data Management and hosted by TU.it. https://researchdata.tuwien.at/ GMA DMP version 1.0 9 Methods or software needed to access and use data: Results are provided in standard CSV format and as a PNG graph. 5b Long-term preservation and deletion of data dataset ID location for long-term storage minimum retention period (≥ 10 years) foreseeable research uses and/or users P1 TU Wien Research Data 10 years Researchers in film studies and data scientists in entertainment analytics. P2 TU Wien Research Data 10 years P3 TU Wien Research Data 10 years 6. RDM responsibilities and resources 6a RDM-roles and responsibilities The Principal Investigator (Vlad Dancea) will direct the data management process overall, with the research assistants responsible for ensuring metadata production, day-to-day cross-checks, back-up and other quality control activities are maintained. 6b Resources There are no costs dedicated to data management and ensuring that data will be FAIR.