Full text
Posted on 15 Jun 2025 — CC-BY-NC-SA 4 — https://doi.org/10.36227/techrxiv.175000742.26542722/v1 — e-Prints posted on TechRxiv are preliminary reports that are not peer reviewed. They should not b... Andreas Valvis1, Vasiliki Balaska1, Ioannis Kansizoglou1, Loukas Bampis1, and Antonios Gasteratos1 1Affiliation not available June 15, 2025 1
SynthPark: Leveraging Synthetic Data for Occupancy Prediction in Autonomous Logistics Andreas Valvis1, Vasiliki Balaska1, Ioannis Kansizoglou1, Loukas Bampis2and Antonios Gasteratos1 Abstract— Towards a more digitalized and automated future, recent years have witnessed a significant increase in the deployment of Unmanned Aerial Vehicles (UAVs) across various domains, particularly in surveillance and logistics. However, object detection in obscure environments such as parking lots, particularly for trucks, remains challenging due to the scarcity of real-world datasets. Previous datasets such as, Common Objects in Context (COCO), PASCALVisual Object Classes (VOC), and PKLot are valuable for general object and parking occupancy detection but lack aerial-view truck images, limiting their applicability for real-time monitoring of truck parking availability using UAV imagery. To address this limitation, this paper proposes a synthetic dataset generated in an open-source simulation environment. In this digital twin environment, a realistic parking lot scene was designed and filled with trucks under dynamic lighting, camera angles, and spatial configurations. To prove its applicability, the synthetic dataset was used to train an object detection model. Early stage experiments indicate promising accuracy results which are even translated to real-world conditions. This study demonstrates the efficacy of our synthetic data in addressing the limitations of existing baseline datasets. Additionally, it overcomes legal and ethical constraints associated with aerial image capture, providing a scalable framework for enhancing truck parking detection under diverse conditions. I. INTRODUCTION Over the past years, Unmanned Aerial Vehicles (UAVs) have become integral to a wide range of applications, from surveillance to logistics. Therefore, modern industries and logistics structures need to tackle a new set of real-world problems. One of the challenges in these applications is robust object detection, particularly in dynamic and diverse environments such as parking lots. Vehicles like trucks and parking spots introduce substantial spatial and visual variability, complicating this procedure. As a result, training detection models effectively is not feasible due to the scarcity of realworld datasets. During the past decades, the MNIST [1], 1A. Valvis, V. Balaska, I. Kansizoglou, and A. Gasteratos are with the Department of Production and Management Engineering, Democritus University of Thrace, Xanthi, Greece. [email protected], [email protected], [email protected], [email protected] 2L. Bampis is with the Department of Electrical and Computer Engineering, Democritus University of Thrace, Xanthi, Greece. [email protected] COCO [2], PASCAL-VOC [3], and PKLot [4] datasets have been widely used in computer vision and machine learning, each serving distinct purposes based on their content and structure. However, the collection of such data often raises ethical and legal concerns, particularly regarding privacy in the use of aerial imagery. Moreover, the collection of those data are expensive, time-intensive, and limited in their ability to capture a wide range of scenarios. Our work aligns with the SPATRA.EU project1, which aims to access UAVs for real-time monitoring of parking availability in designated lots. However, one of the fundamental issues faced refers to lack of a dataset that captures trucks in such environments from the perspective of an UAV. To address this problem, synthetic data generation can offer a promising solution, allowing the creation of new, diverse datasets in controlled virtual settings. Such a strategy is already proven, with similar challenges addressed by generating synthetic data and training a machine learning model with their synthetic dataset, achieving results that underline the potential of simulation to overcome data scarcity [5]. In addition, other works have proposed simulation environments for applications in UAVs and Industry 5.0, demonstrating their applicability and impact on other domains, as well [6] [7]. Building upon this, in this paper a novel dataset2 is proposed to train an object detection model for UAV-based truck parking occupancy prediction using synthetic data. A realistic parking lot environment was developed in Unity3, populated with trucks and designed to replicate diverse real-world conditions, including varying lighting, UAV camera angles, and spatial arrangements. High-fidelity synthetic images were rendered to create a comprehensive dataset for training and validating the model. The findings demonstrate that this synthetic data pipeline effectively addresses the lack of real-world UAV data, delivering sufficient detection 1More details about the EU-funded SPATRA project can be found at https://spatra-project.eu/. 2Our complete dataset can be found in https://github.com/eeeAndrew/SynthPark 3More information about the Unity graphics engine can be found at https://unity.com/.
performance. This work advances the application of synthetic data in deep learning, offering a scalable and adaptable framework for enhancing perception in autonomous UAV navigation and robotics. The following sections are constructed as follows. Section II analyzes the related literature. Section III explains the design of our methodology, while Section IV defines our experimental protocol and presents the obtained results for benchmarking the proposed dataset. Finally, Section V concludes the study and outlines directions for future research. II. RELATED WORK The collection of reliable, accurate, and high volume datasets for training machine learning architectures is a never-ending process that supports some of the most significant advancements in recent technology. COCO [2], or Common Objects in Context, is a large-scale dataset with over 330,000 images of everyday scenes, featuring 80 object categories and rich annotations for object detection, segmentation, and captioning, making it ideal for complex tasks requiring contextual understanding. In addition, PASCAL-VOC [3] comprises around 20,000 images across 20 object classes, such as vehicles, animals, and household items, with annotations for object detection, segmentation, and classification, widely used for benchmarking object recognition algorithms. PKLot [4] focuses on parking lot scenarios, containing approximately 12,400 images of parking spaces captured under varying weather and lighting conditions, annotated to classify spaces as occupied or vacant. This dataset is valuable for classifying parking spaces as occupied or vacant, making it suitable for car parking detection systems, particularly in ground-level conditions. The work in [8] focuses on advancing visual odometry and SLAM in complex environments by introducing a dataset that leverages the strengths of event cameras. Recognizing the limitations of existing datasets—often monocular, low-resolution, or lacking precise ground truth—they developed a comprehensive data collection pipeline. Their dataset includes high-resolution, synchronized stereo event and standard camera data, enriched with accurate ground truth from motion capture systems and GNSS/INS. Yoshinobu et al. [9] explored visual navigation in unstructured environments, emphasizing the importance of identifying traversable regions. Some approaches bypass this by using deep learning to directly infer movement direction from images, though they face challenges in acquiring large, high-quality datasets. To address this, the study has proposed data acquisition systems and dataset generation methods, releasing imagepath pairs and sensor data for broader use. Last, synthetic dataset generation method has been introduced in [10] using CityEngine for procedural 3D modeling and Unreal Engine 4 for high-fidelity rendering. This pipeline produces multi-modal data, including spectral images, semantic segmentation maps, depth information, and surface normals. The synthetic environments replicate realistic urban conditions and UAV trajectories, enhancing model performance in urban scene understanding tasks. Another approach, presents the FakePS [11] which is an extensible pipeline for generating synthetic data in simulated parking scenarios. It incorporates pixel-level domain adaptation to increase the realism of synthetic images using unlabelled realworld data. This method significantly improves the training of parking spot detection models, while reducing the dependence on manually annotated datasets. However, the above parking-focused dataset like primarily addresses cars and objects under ground-level view without any images from trucks. The lack of suitable datasets motivates our approach, which involves generating a synthetic dataset and training a state-of-the-art machine learning model to address the parking occupancy problem in truck parking lots. III. FRAMEWORK DESIGN In this section, we present the environment’s architecture and all the necessary tools and techniques to achieve the generation of a robust dataset. A. Environment Construction Fig. 1: Google Earth Parking Lot Multiple assets for various tasks have been developed by the open-source Unity 3D graphics platform. Objects included in these assets are commonly found in real life, such as buildings, cars, roads, and trees. A digital environment was developed using 3D models to virtually recreate the real area of a parking lot located in Serbia (Lat: 45.0466667, Long: 19.2255555). The next step involved overlaying an actual image of the area, which allowed the lines of the parking lot to be accurately drawn. A satellite image of the target location was obtained from Google Earth to extract accurate geospatial measurements as illustrated in Fig. 1. The selected area spans 250 meters in length and 90 meters in width, with a total perimeter of 720 meters and an approximate area of 22.000 square meters. This also helped to preserve the colors of the objects, as the image served as a visual
reference. To create a more realistic environment, surFig. 2: Top-down view image on the simulated terrain rounding terrain was added to the scene, as illustrated in Fig. 2. Realism is considered crucial both for developing a robust dataset that accurately represents a real parking lot and for effectively training any object detection machine learning model. To achieve this, a dynamic lighting system was developed in Unity to replicate realworld conditions. Using Unity’s High Definition Render Pipeline (HDRP), a physically based sky model was configured with a directional light source to simulate the movement of the sun [5], [10], [11]. This setup allowed the capture of various scenarios at different times of the day, from dusk to dawn. Realistic shadows and ambient occlusion were incorporated into the system, adjusted to match diverse lighting conditions encountered by UAVs in real-world operations. For rendering, Unity’s shaders were used to achieve high-quality textures for trucks, buildings, streets, and surrounding objects, preserving fine details such as surface wear and reflections. Next, the objects are placed in the simulation environment, while maintaining the absolute lengths based on the reference image introduced earlier. The core entities included in the environment are roads for vehicle movement, the parking lot with lines to illustrate the available parking spots, and vehicles. To further advance the machine learning occupancy detector with occluded objects, trees, street lights, fences, and a gas station were included in the environment. These elements were added to simulate real-world occlusions that challenge parking occupancy detection, such as leaflets partially obscuring vehicles or trees casting shadows. Furthermore, environmental clutter such as signage, adjacent buildings, and randomly parked vehicles was integrated. These additions increased the visual complexity of the scene and further testing the capabilities of object detection models to differentiate trucks from background noise and overlapping elements. Figure 3 and Fig. 4 provide representative examples, demonstrating how shadows and viewing angles create occlusions and enhance the robustness of the dataset for truck parking occupancy detection in real-world UAV deployments. Fig. 3: Simulated Parking Lot with shadows - Front View Fig. 4: Simulated Parking Lot without shadows - Top View B. Dataset Generation In real-world applications, the creation of an annotated dataset requires an agent capable of aerial navigation and image capturing from various viewing angles within the study area. In this work, Unity’s camera Prefab tool was used to simulate a drone’s camera in order to rotate and change position. Hence, our images were recorded from varying altitudes and locations in the area, using manual inputs. To ensure robustness of the dataset collected, diverse occupancy scenarios were simulated, such as fully occupied, semi-occupied, or empty parking lots. To that end, multiple objects were repositioned at different locations before capturing frame, aiming to improve the model’s ability to generalize the predictions in unseen footage. Finally, the rendering process utilized Unity’s virtual camera system to capture high-resolution images, simulating aerial viewpoints at altitudes ranging from 10 to 50 meters, resulting in a unique perspective, orientation, and scene composition for each image, emulating the variability encountered in real UAV surveillance missions. Object instances from each recorded image were annotated with two different labels following a multi-class :‘truck‘, representing all visible trucks, and ‘parking spot‘, indicating areas that could potentially be occupied by a vehicle. To address the limitations of dataset size and promote the model’s ability to generalize across unseen conditions, extensive data augmentation techniques were employed, resulting in a substantially enriched dataset suitable for training detection models.
C. Dataset Augmentation To further enhance the model’s ability to generalize across a wide range of real-world conditions, a comprehensive data augmentation pipeline was employed. The motivation behind each selected technique was to simulate realistic variations in UAV surveillance data, thus improving model robustness and reducing overfitting. Geometric transformations were applied to introduce spatial variability. In particular, horizontal and vertical flips (50% probability) to help the training process in generalizing across different object arrangements, or variations in the UAV’s flight path and the camera’s pitch, especially relevant in symmetrical environments like parking lots. Random shifts (up to 5%), scaling (up to 10%), and rotations (up to 20%), were incorporated to simulate positional and angular differences caused by drone movement or environmental constraints. In addition, photometric adjustments were introduced to account for changes in lighting and sensor behavior. Brightness and contrast alterations (50%) simulated different times of day and shadowing effects. Gamma correction (30%) addressed non-linear exposure variations from different capture devices. Hue-saturationvalue shifts (40%) and RGB channel shifts (30%) were applied to introduce color variability due to atmospheric conditions or device-specific color calibration, encouraging the model to rely less on color and more on object shape and structure. Moreover, to replicate sensor noise and motion artifacts, several blurring techniques were integrated. Standard blur, Gaussian blur, and motion blur (each with 30% probability) mimicked common distortions in drone-captured footage, such as camera vibration, low quality image or drone movement. Gaussian noise (40%) was added to simulate electronic noise, a frequent issue in low-cost UAV systems. Perspective distortions (30%) were used to emulate different viewpoints and tilts that naturally occur in aerial navigation, training the model to recognize objects from various angles. More advanced weather-based effects, such as artificial rain and fog (each with 30% probability), were added to expose the model to challenging environmental conditions often encountered during outdoor UAV missions. Finally, Contrast Limited Adaptive Histogram Equalization (CLAHE) was applied with a 30% probability to enhance image contrast in low-light or uneven illumination scenarios, ensuring visibility of features across the scene [12]. All the above data augmentation techniques generate 250 different images along with their labels, improving the model’s ability to generalize and predict unseen data. Figure 5 demonstrates some of the produced dataset results. (a) Augmenter sample 1: Right-shifted perspective and brightness increase. (b) Augmenter sample 2: Noise emulating low-quality camera sensor. (c) Augmenter sample 3: Weather effect simulating rain. (d) Augmenter sample 4: Brightness and contrast adjustment replicating different lighting conditions. Fig. 5: Representative examples of augmented dataset instances IV. EVALUATION AND RESULTS A. Experimental Setup To validate the quality and applicability of the created dataset, it is essential to train a neural network and evaluate the performance metrics both during training and upon completion. Specifically, version 114of YOLO (You Only Look Once), the latest release available from Ultralytics, was utilized. This version was selected because it is considered state-of-the-art and, according to the literature, outperforms all previous versions in terms of accuracy and speed. For training YOLOv11 on the dataset, pretrained weights obtained from Roboflow were used, specifically those trained on the PKLot dataset, which contains 12,416 images. These were split into the following: •Training set: 8.691 images •Validation set: 2.483 images •Testing set: 1.242 images (corresponding to a 70/20/10 split). The use of pretrained weights is crucial in tasks involving relatively small datasets to maintain the already learned low-level features, such as edges, shapes, and basic object structures from large-scale datasets. Additionally, starting from pretrained weights significantly reduces the number of epochs required to achieve good performance. Beyond the weights themselves, the model architecture is also transferred, which is already optimized for object detec4The YOLOv11 implementation used for this paper can be found in: https://docs.ultralytics.com/models/yolo11/
tion tasks, such as vehicle and parking spot identification. The model was subsequently fine-tuned using our generated simulated dataset, in order to better adapt to the spatial characteristics of trucks and parking areas. The whole process was conducted on machine equipped with an NVIDIA RTX 3060 GPU (8GB VRAM). The batch size was 16 and the learning rate was automatically selected by the Ultralytics API within the range [1e−5,1e−2]. Last, training was performed for up to 500 epochs, with early stopping applied to prevent overfitting. The training process was halted when the validation loss remained stable for several consecutive epochs, indicating no significant improvement. B. Model Evaluation Metrics The model demonstrated satisfactory performance on several evaluation metrics for inferring the regions and labels of the trucks and parking spots. The mean Average Precision (mAP) at Intersection over Union (IoU) thresholds 0.5 to 0.95 (mAP50-95) reached 0.761 during testing, indicating a solid detection capability. The mAP at IoU threshold 0.5 (mAP50) was 0.942, which highlights the model’s accuracy in correctly identifying objects with a relatively high degree of overlap. Furthermore, the model exhibited a precision of 0.935 and a recall of 0.877, showing a good balance between correctly identifying trucks and minimizing missed detections. These values reflect the model’s effectiveness in retrieving relevant objects while maintaining a low rate of false positives.Throughout the training process, both training and validation losses consistently decreased. The training box loss and classification loss were 0.7755 and 0.4837, respectively, while the validation box loss and classification loss were 0.6992 and 0.4572, respectively. This suggests that the model was effectively learning to minimize errors during training while maintaining solid performance during validation, demonstrating its ability to generalize across unseen data. In terms of efficiency, the model requires approximately 28.6 GFLOPs and contains 11.1 million parameters, offering a good trade-off between computational complexity and detection performance. Furthermore, inference speed was measured at around 3.74 ms per image. Overall, these evaluation results confirm that the model can predict parking spots and trucks while maintaining computational efficiency. C. Evaluation Metrics of the Parking Occupancy detector To determine whether a parking spot is occupied, Non-Maximum Suppression (NMS) is applied to the model’s output. NMS is a post-processing technique used to remove redundant overlapping bounding boxes that the model may predict for the same object. In the case of parking occupancy detection, the model might detect multiple boxes around the same pot due to a variety of factors, such as slight variations in the position or scale of the trucks. NMS helps to filter out these duplicate detections by retaining only the most confident (highest scoring) bounding box for each parking spot. This ensures that only one detection per occupied parking spot is considered, thereby improving the accuracy of occupancy predictions and avoiding multiple false detections for a single spot. To use this method, the ground truth data for the set of testing images were additionally labeled into two classes: •Class 1: Defines the occupied spots with a label of ‘1‘, where the bounding box of trucks and parking spots overlap. •Class 0: Defines the unoccupied spots with a label of ‘0‘, where the bounding boxes are independent. The predictions are then compared against ground truth annotations in the test images to evaluate performance using several metrics, including True Positive (TP), True Negative (TN), False Positive (FP), and False Negative (FN). Those metrics reflect the proportion of total predictions (both occupied and unoccupied) that the model correctly predicted. However, the performance and effectiveness of the presented method are further evaluated using additional comprehensive metrics: overall accuracy, precision, recall, and F1 score. Accordingly, the performance of the evaluation metrics is summarized in Table I. As shown, the proposed method demonstrates satisfactory results, achieving 74% accuracy, 68.2% precision, 67% recall, and an F1 score of 69%. Furthermore, TABLE I: Model Evaluation Metrics Model’s Evaluation Metrics Performance Accuracy 74% Overall Accuracy 64,44% Precision 68,2% Recall 67% F1 Score 69% the Precision-Recall (PR) curve, shown in Fig. 6, was generated using 35 unseen testing images and provides a comprehensive visual representation of the trade-off between precision and recall across various classification thresholds. The curve exhibits a steady decline with noticeable fluctuations, dropping to a Precision of approximately 0.6 at a Recall of 0.5, and further to around 0.3 at a Recall of 0.8, before reaching a Precision of about 0.1 at a Recall of 1.0. These fluctuations suggest variability in the model’s ranking of positive predictions, potentially due to the distribution of confidence scores or the presence of hard-to-classify samples. The mAP was calculated as the area under this PR curve at 67.9%
Fig. 6: Testing - Precision recall curve V. CONCLUSIONS This paper presented a new dataset for parking occupancy prediction, leveraging synthetic data generated in Unity. By creating a controlled, synthetic dataset, our study addresses the challenge of limited real-world aerial imagery, particularly for truck parking scenarios. A YOLOv11-based object detection model was developed and trained on this synthetic dataset, demonstrating its effectiveness in predicting parking occupancy. The approach ensures satisfactory performance while avoiding cost and legal concerns associated with real-world UAV imagery. Finally, it highlights the significant potential of synthetic data to enhance model robustness and practical applicability in robotics. Future work will focus on deploying this model on a physical drone to monitor parking lot occupancy in real-world scenarios, such as large-scale truck parking lots, ensuring its practical applicability in autonomous logistics systems. ACKNOWLEDGMENT This research received funding from the European Union’s Horizon Europe research and innovation program under grant No. 101129658. In addition, it was partially supported by the project TAEDR-0535864 “Network of Excellence for the Development, Dissemination and Application of Digital Transformation Technologies in the Greek Manufacturing Industry” carried out within the framework of the National Recovery and Resilience Plan Greece 2.0, funded by the European Union - NextGenerationEU. REFERENCES [1] L. Deng, “The mnist database of handwritten digit images for machine learning research,” IEEE Signal Processing Magazine, vol. 29, no. 6, pp. 141–142, 2012. [2] T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll´ ar, and C. L. Zitnick, “Microsoft COCO: common objects in context,” in Computer Vision - ECCV 2014 - 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V, ser. Lecture Notes in Computer Science, D. J. Fleet, T. Pajdla, B. Schiele, and T. Tuytelaars, Eds., vol. 8693. Springer, 2014, pp. 740–755. [Online]. Available: https://doi.org/10.1007/978-3-319-10602-1 48 [3] M. Everingham, S. M. A. Eslami, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes challenge: A retrospective,” International Journal of Computer Vision, vol. 111, no. 1, pp. 98–136, Jan. 2015. [4] M. Everingham, S. M. A. Eslami, L. V. Gool, C. K. I. Williams, J. M. Winn, and A. Zisserman, “The pascal visual object classes challenge: A retrospective,” Int. J. Comput. Vis., vol. 111, no. 1, pp. 98–136, 2015. [Online]. Available: https://doi.org/10.1007/s11263-014-0733-5 [5] E. Bonetto and A. Ahmad, “Synthetic data-based detection of zebras in drone imagery,” in European Conference on Mobile Robots, ECMR 2023, Coimbra, Portugal, September 4-7, 2023, L. Marques and I. Markovic, Eds. IEEE, 2023, pp. 1–8. [Online]. Available: https://doi.org/10.1109/ECMR59166.2023. 10256293 [6] M. Kosmidis, S. Antoniou, V. Balaska, I. Kansizoglou, and A. Gasteratos, “Towards warehouse 5.0: A framework of human-centered technologies,” in IEEE International Conference on Imaging Systems and Techniques, IST 2024, Tokyo, Japan, October 14-16, 2024. IEEE, 2024, pp. 1–6. [Online]. Available: https://doi.org/10.1109/IST63414.2024.10759201 [7] T. Mitroudas, V. Balaska, A. Psomoulis, and A. Gasteratos, “Light-weight approach for safe landing in populated areas,” in 2024 IEEE International Conference on Robotics and Automation (ICRA), 2024, pp. 10 027–10 032. [8] A. Hadviger, V.-J. ˇ Stironja, I. Cviˇ si´ c, I. Markovi´ c, S. Vraˇ zi´ c, and I. Petrovi´ c, “Stereo visual localization dataset featuring event cameras,” in 2023 European Conference on Mobile Robots (ECMR), 2023, pp. 1–6. [9] Y. Uzawa, S. Matsuzaki, H. Masuzawa, and J. Miura, “Dataset generation for deep visual navigation in unstructured environments,” in 2023 European Conference on Mobile Robots (ECMR), 2023, pp. 1–6. [10] Q. Gao, X. Shen, and W. Niu, “Large-scale synthetic urban dataset for aerial scene understanding,” IEEE Access, vol. 8, pp. 42 131–42 140, 2020. [11] J. Chen, L. Zhang, Y. Shen, Y. Ma, S. Zhao, and Y. Zhou, “A study of parking-slot detection with the aid of pixel-level domain adaptation,” in 2020 IEEE International Conference on Multimedia and Expo (ICME), 2020, pp. 1–6. [12] L.-h. He, Y.-z. Zhou, L. Liu, W. Cao, and J.-h. Ma, “Research on object detection and recognition in remote sensing images based on yolov11,” Scientific Reports, vol. 15, no. 1, p. 14032, 4 2025. [Online]. Available: https://doi.org/10.1038/ s41598-025-96314-x