scieee AI-readable full text Open interactive document viewer

DMP: Temporal Context in Computer Vision Detection Model

Zideck, Moritz Anton

Abstract

This dataset contains evaluation results, model artefacts and analysis outputs from experiments investigating the impact of temporal context on robustness to visual perturbations in video object detection. Models based on YOLOV-Swin and YOLOX architectures were trained using PyTorch on the VisDrone dataset (public, CC BY-NC-SA 3.0) and evaluated under both clean conditions and simulated noise levels ranging from 10 % to 60 %. In addition, the XS-VID dataset (MIT license) was used to assess model performance in scenarios dominated by extremely small objects, enabling targeted evaluation of temporal modelling under challenging visual conditions. The dataset includes performance metrics (mAP, mAR, inference time) in CSV format, visual performance plots, qualitative prediction examples, trained model checkpoints and summary reports. All scripts, model configurations and environment specifications required to reproduce or extend the experiments are provided in the accompanying code repository:https://github.com/mozi30/TemporalAttentionPlayground Noise-based robustness results were generated programmatically using performance degradation models, rather than full retraining on corrupted data. Further evaluation using additional architectures such as MSDA and TransVOD is planned as future work. This dataset is intended to support reproducible experimentation and comparative benchmarking in drone-based perception research, temporal attention modelling and robust computer vision development.

Full text

Data management plan (DMP) Temporal Context in Computer Vision Detection Model TCR-VOD Version Effective date Description of document/changes 1.0 30/11/2025 First version of the DMP – created for the start of the project Level of distribution This DMP is licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0). It is publicly available under 10.5281/zenodo.17771932. 3 FWF Data Management Plan (DMP) I General Information I.1 Administrative information PI: FWF project number: Internal Project ID: DMP version: 1.0 Contributors: Moritz Zideck, [email protected], TU Wien, ROR: ror.org/04d836q62, Data Manager I.2 Data management responsibilities and resources Person responsible for data management and DMP: Moritz Zideck, [email protected], TU Wien, ROR: ror.org/04d836q62 Co-ordination of data management responsibilities across partners: Moritz Zideck, [email protected]en.ac.at, TU Wien, ROR: ror.org/04d836q62 Resources costed in for RDM: There are no costs dedicated to data management and ensuring that data will be FAIR. II Data Characteristics II.1 Data description and collection or reuse of existing data Produced datasets: dataset ID title type format estimated volume contains sensitive data description P1 Temporal Context in Computer Vision Detection Model - Results Standard office documents, Structured text, Images, Configuration data application/pdf, image/JPEG, text/x-python 5 - 20 GB no This dataset contains the outputs generated during the robustness evaluation of temporal attention models for video object detection. It includes: Quantitative evaluation results stored as CSV files for each dataset (VisDrone and XS-VID) and each level of noise applied (0%, 10%, 30%, and 60%). These files contain performance metrics such as precision (mAP), recall (mAR), and inference time. Graphs and visualizations illustrating metric trends across models and noise levels. Example output images that show qualitative prediction results for different architectures. Model checkpoints and training logs, including configuration files and TensorBoard logs used during model development and evaluation. Metadata, including experiment timestamps, code version, dataset references, license information, and a note indicating that noise robustness data was synthetically derived from base metrics. The included results are suitable for transparency, reproducibility, and comparative model analysis but should be interpreted with caution, as robustness values under noise are simulated rather than measured through separate full training runs. Reused datasets: dataset ID title source rights (e.g. license) contains sensitive data description R1 Visdrone-VID https://github.com/VisDrone/VisDroneDataset CC BY-NC 4.0 (nonno This dataset contains the outputs generated during the 5 commercial use only) robustness evaluation of temporal attention models for video object detection. It includes: Quantitative evaluation results stored as CSV files for each dataset (VisDrone and XS-VID) and each level of noise applied (0%, 10%, 30%, and 60%). These files contain performance metrics such as precision (mAP), recall (mAR), and inference time. Graphs and visualizations illustrating metric trends across models and noise levels. Example output images that show qualitative prediction results for different architectures. Model checkpoints and training logs, including configuration files and TensorBoard logs used during model development and evaluation. Metadata, including experiment timestamps, code version, dataset references, license information, and a note indicating that noise robustness data was synthetically derived from base metrics. The included results are suitable for transparency, reproducibility, and comparative model analysis but should be interpreted with caution, as robustness values under noise are simulated rather than measured through separate full training runs. R2 XS-VID https://github.com/gjhhust/XS-VID MIT no XS-VID is a synthetic videobased dataset designed to evaluate object detection performance in scenarios dominated by very small objects. It simulates challenging visual conditions typically observed in high-altitude or lowresolution aerial imagery, where objects occupy only a few pixels. The dataset is used to benchmark robustness and temporal consistency of detection models, particularly under perturbations such as noise and visual degradation. XS-VID serves as a complementary evaluation benchmark to VisDrone, focusing specifically on sensitivity to object scale. The results derived from this dataset highlight model limitations in detecting and tracking extremely small targets. License: MIT Methods and software used for data generation and reuse 7 The research data is generated through training and evaluation experiments using video object detection architectures. Two datasets are used for this purpose: the publicly available VisDrone dataset, which is reused for evaluation under realistic aerial conditions, and the synthetically generated XS-VID dataset, which simulates scenarios where objects are extremely small. Models based on the YOLOV-Swin and YOLOX architectures are trained using PyTorch, and temporal configurations, such as the use of multiple input frames, are assessed to analyse their impact on robustness. After training, each model is evaluated using standard COCO metrics (mAP and mAR) and inference time, under both clean input conditions and simulated noise levels of 10 %, 30 % and 60 %. The noise robustness values are generated programmatically using performance degradation models rather than retraining the models on physically corrupted data, allowing expected robustness behaviour to be illustrated efficiently. The software used for data generation includes PyTorch for model development and evaluation, custom Python scripts for metric extraction, statistical analysis and CSV creation, Matplotlib and Pandas for visualisation and processing, and TensorBoard for logging during training. All scripts, configuration files and evaluation tools required to regenerate or extend the results are available in the code repository at https://github.com/mozi30/TemporalAttentionPlayground . The repository also contains environment configuration files (requirements.txt and setup-specific dependency files) to allow reconstruction of the original software environment and support reproducibility and reuse. The research data produced consists of performance metric CSV files, visual plots, example prediction images, model checkpoints and summary documents (e.g. PDF reports). For reuse, these files can be accessed either individually or via the structured results package, and can be reprocessed using the provided Python scripts and model configurations. Due to time constraints, two additional architectures—MSDA (Multi-Scale Deformable Attention) and TransVOD (Transformer-based Video Object Detection)—were not evaluated within the current experimental phase. These models were identified as promising candidates for extended robustness benchmarking, particularly concerning multi-frame temporal inference, and are planned for inclusion in future work to enhance the comprehensiveness of the analysis. III Documentation and Data Quality III.1 Metadata and documentation Data organisation, metadata, and documentation: The research data is structured in a clearly organised hierarchical folder system that separates input datasets, derived results, model files, and visual outputs. Within the project repository (https://github.com/mozi30/TemporalAttentionPlayground), all generated artefacts are stored inside the top-level directory results/. This includes a raw_data/ subfolder containing machine-readable CSV files for evaluation results. Separate subdirectories exist for each input dataset, such as visdrone/ with files like base_results.csv and noise-variant metrics (base_results_noise10.csv, base_results_noise30.csv, base_results_noise60.csv), and an equivalent structure for the XS-VID dataset. Visualisations of the performance metrics (mAP, mAR, inference time under different noise levels) are stored in the graphs/ directory as PNG files. Qualitative prediction results are collected in the examples/ folder. Model checkpoints, training logs, TensorBoard data and configuration scripts are stored under models/. Additionally, a metadata.json file is included to provide machine-readable documentation that describes the datasets, evaluation metrics, licensing conditions, and links to external repository versions or DOIs. All CSV files follow a consistent column schema (including fields such as timestamp, dataset, model, noise_level, map50-95, mAP50, mAR50-95, and inference_time_ms) and are documented in the accompanying README to ensure that the meaning and units of all variables are transparent. File names encode dataset identity and noise level (e.g. base_results_noise30.csv) to support self-descriptive storage. Raw numerical data is stored separately from derived artefacts such as plots and PDF summaries to enable users to regenerate figures or carry out additional analyses. Versioning is implemented at both code and data level. All modelrelated scripts and evaluation tools are version-controlled using Git. Important experimental milestones are tagged with semantic versions (e.g. v1.0.0) and accompanied by changelog entries. Each CSV file contains a timestamp and is linked via metadata to the specific Git commit used to generate it. For dataset versioning, stable releases of the results package—including CSV files, plots, model checkpoints, metadata.json and README—will be deposited in the institutional research data repository as immutable, versioned snapshots, each assigned a DOI. Each release will reference the corresponding software version and provide a summary of differences from prior releases. Older releases will remain publicly available to ensure reproducibility. The metadata.json file records the dataset version and the associated code tag to guarantee traceability. Large artefacts, such as complete model checkpoints and extended log files, are stored outside of the Git repository (e.g. in the research data repository) and referenced via DOI or a persistent URL in the README and metadata, ensuring long-term accessibility without overloading the version control system. The dataset will include both human-readable and machine-readable metadata to ensure that it can be easily interpreted, located, and reused by others. The human-readable documentation will be provided in a README file and will contain a description of the dataset’s purpose, the experimental context, and the methodology used. It will also explain the folder structure and all included file types, provide a complete table defining all CSV columns with their units and meaning, and include notes on model configurations, the method used for simulating noise, the evaluation metrics applied, and any relevant limitations. In addition, the README will clearly state the licenses for both source data and derived results, and provide citation and reuse guidance. Complementing this, machine-readable metadata will be supplied in a metadata.json file. This file will include the project title, information on authors (including ORCID where available) and institutional affiliation, the dataset version and corresponding code commit identifier, details of the datasets used including names, sources, URLs, and licenses, as well as information on the output metrics, file naming conventions, and noise levels applied. It will also specify the software and scripts that were used to generate the data, include license information for both the produced data and associated code, and refer to the intended reuse context and any external repositories or DOIs. At file level, each CSV will include human-readable header lines specifying key metadata such as the experiment timestamp, the code commit hash used during data generation, and the dataset license. Each entry will contain an ISO 8601–compliant timestamp column, and metric names will follow standard COCO evaluation conventions to maximise interoperability and consistency. The repository will also include a CITATION.cff file to provide structured citation guidance for referencing both the codebase and result dataset. When the dataset is deposited in the research data repository, additional metadata such as keywords, abstract, licensing details, authorship, funding information, and versioning will be added, and a DOI will be assigned to support persistent access and citation. This will help others to identify, discover and reuse our data. Additionally, we will provide common metadata such as title, description, or keywords when publishing data in open access repositories. In such a case, we will follow the default template provided by the repository, such as Data Cite Metadata or Dublin Core. As far as possible, we will use controlled vocabularies for our data to allow inter-disciplinary interoperability and machine-actionability. To ensure that others can validate the data analysis and reuse the dataset responsibly, documentation will be provided in several ways. A detailed README file will explain the purpose of the experiment, the dataset structure, the evaluation methods, the definitions of performance metrics, model configurations, and how to reproduce the results using the provided scripts. Each results file contains metadata such as timestamp, code version (Git commit), dataset reference, noise level, and model parameters to make the analysis traceable. In addition, all scripts used to generate the results will be included in the repository together with a requirements file listing the software dependencies. A machine-readable metadata.json file will describe the dataset, metrics, data sources, licenses, and links to external archived versions (via DOI). Known limitations and usage recommendations will also be documented, including the note that some robustness results are synthetic simulations, and that models such as MSDA and TransVOD were not evaluated due to time constraints. Together, this documentation provides everything needed to validate the analysis, reproduce the results, and enable correct and informed reuse of the data. 9 III.2 Data quality control Data quality control: The following data quality checks will be done: standardised data capture and repeated samples or measurements. IV Data Storage, Sharing, and Long-Term Preservation IV.1 Data storage and backup during the research process Storage and backup facilities: For the duration of the project, storage and backup of data will be ensured by Moritz Zideck (acting as the person responsible for data management and DMP) in cooperation with the system operator. The data will not be stored on the servers of TU Wien. P1 (Temporal Context in Computer Vision Detection Model - Results) will be stored on Other. External storage will be used because it is required from the project designers R1 (Visdrone-VID) will be stored on Other. External storage will be used because it is required from the project designers R2 (XS-VID) will be stored on Other. External storage will be used because it is required from the project designers Data security and protection of sensitive data: We pay strict attention to compliance with the relevant institutional and national data protection policies. At this stage, it is not foreseen to process any sensitive data in the project. If this changes, advice will be sought from the data protection specialist at TU Wien, and the DMP will be updated. Access to data during research: dataset ID selected project members all other project members the public P1 writing reading only reading only R1 reading only reading only reading only R2 reading only reading only reading only IV.2 Data sharing and long-term preservation Data publication and access conditions: As far as possible, obtained datasets will be published in repositories. Details on access conditions, reuse licenses, reasons for restrictions, etc. are collected in the table below.