Full text
Leveraging Synthetic Data for Deep-Learning-Based Road Crack Segmentation from UAV Imagery Andriani Panagi[0009−0000−9212−2504] and Christos Kyrkou[0000−0002−7926−7642] KIOS Research and Innovation Center of Excellence University of Cyprus, Cyprus, Nicosia, Cyprus {panagi.andriani, kyrkou.christos}@ucy.ac.cy https://www.kios.ucy.ac.cy/ Abstract. One of the critical tasks in monitoring of road infrastructure is the identification of road cracks. Recent efforts have been made in utilising Unmanned Aerial Vehicles (UAVs) to automate this task without interfering with the road network traffic and infrastructure. However, high-quality annotated datasets that allow the development of reliable deep learning models for this purpose, are scarce. Synthetic data generation offers a promising alternative to mitigate this issue by reducing annotation costs and enhancing the dataset’s variability. This paper represents a comparative study of two state-of-the-art deep learning models - UNet with EfficientNet and UNet with MobileNetfor road crack segmentation trained with three loss functions (Dice Loss, Focal Loss, and Weighted Binary Cross-Entropy Loss (WBCEL)) by using synthetic datasets generated with three different ways. Performance evaluation indicates that the best results were achieved using UNet with a MobileNet encoder, trained with WBCEL and synthetic images, yielding an mIoU of 63.52% tested on real crack imagery. A more granular analysis underlines how both synthetic data realism and the choice of the loss function impact segmentation accuracy. Our preliminary study concludes that well-designed synthetic data and appropriate loss functions have the potential to allow better generalization of the model to real-world scenarios. Keywords: Synthetic Image Generation ·Semantic Segmentation ·UAV Imagery ·Road Crack Segmentation ·Convolutional Neural Networks 1 Introduction Monitoring and maintaining road networks is critical for ensuring continuous transport, minimising disruptions, and fostering and supporting economic productivity. Conversely, decayed infrastructure can have disastrous consequences, including increased accident risks, higher vehicle maintenance costs, and transportation delays. Road defects, such as cracks, not only contribute to inefficiencies in commerce and transportation but also threaten human lives. The World Health Organization (WHO) reports that 1.19 million traffic accidents occur
2 A. Panagi and C. Kyrkou annually, injuring 20 to 50 million people [1], with road damage often playing a vital role. Uneven surfaces can lead to loss of vehicle control, reduced traction between the tires and the road, and increased driver stress and attention, ultimately leading to increased chances of accidents. Traditional methods for detecting road damage typically involve human visual inspection or vehicle-mounted cameras for infrastructure monitoring [7], [28]. Human inspectors are able to estimate not only the visible damage but also contextual circumstances, such as surroundings or structural issues that may occur. However, human in the loop solutions often lead to delays and limited throughput. Meanwhile, vehicle-mounted cameras can cover a wide range of road networks relatively fast and acquire high-resolution images that can be analysed automatically or manually inspected when needed. Nevertheless, these methods have several limitations: manual inspections are highly labour-intensive, time-consuming, and costly, particularly when applied to large sections of road networks. Additionally, they put workers’ safety at risk, especially when they are on busy highways or in severe weather conditions. Vehicle-mounted systems, while automated, can generate a vast amount of data which makes the processing of it extremely resource-intensive. Traditional methods are therefore less effective when used in proactive maintenance and repair strategies as they may be inconsistent and occasionally fail to detect vital yet minor road defects, which ultimately minimise their effectiveness. Identifying road damages using UAVs provides an alternative and timeeffective method for infrastructure monitoring [25] without disrupting road network operations. One of the advantages of UAVs is their ability to cover large, hard-to-reach areas such as highways and rural roads in less time, without interfering with traffic flow compared to traditional methods. Equipped with highresolution cameras and sensors, UAVs can capture detailed imagery of the road surface from multiple angles, therefore improving the likelihood of detecting subtle defects such as cracks. Moreover, their deployment is more cost-effective and adaptable to various monitoring scenarios. However, using UAVs for automated road inspection presents several challenges. Cracks typically occupy only a small percentage of image pixels, which requires deep learning-based segmentation models to be highly precise in detecting even the smallest defects for enabling quick responses. Additionally, acquiring training data and manually annotating the images is often a complex and timeconsuming task that requires significant human and infrastructure resources. Other factors, such as weather conditions, safety concerns and regulatory restrictions, further complicate the data acquisition process. A more efficient alternative is the use of synthetic image data. Synthetic image data refers to artificially generated images through computer algorithms or simulations, rather than those captured by traditional imaging devices. These images closely replicate real-world visuals and are useful for training machine learning models in object detection, segmentation, and recognition tasks, providing a scalable solution for model development. Moreover, the generation of
Synthetic Data for Crack Segmentation from UAV Imagery 3 synthetic data automatically creates pixel-accurate annotations, reduces annotation costs, and accelerates the experimentation cycles. While synthetic data generation has been shown to improve detection performance [29], its effectiveness compared to real data, the choice of model architecture, and the impact of various synthetic data generation techniques remain open questions especially in the context of road crack segmentation from UAV imagery. This paper aims to address these gaps by conducting a comparative study of segmentation models (UNet with EfficientNet and MobileNet encoders), synthetic data generation methods, and loss functions (Dice Loss, Focal Loss, and Weighted Binary Cross-Entropy) for road crack segmentation using UAV imagery. The generalisation ability of models trained on synthetic data is evaluated on a real dataset. The key contributions of this paper are summarised as follows: •We propose a framework for generating synthetic road crack images by first segmenting road regions from UAV-captured images with a top down view and then superimposing crack patterns using three different techniques to ensure realistic integration with the road surface (Figure 1). •We evaluate and compare the performance of different deep learning segmentation model configurations trained on the synthetic datasets, evaluating them with various metrics to assess the impact of three different loss functions on segmentation accuracy. •We derive the best combination of data generation, backbone network and loss function for road crack segmentation for UAV imagery. Fig. 1. Overlaying crack images on road infrastructure background images captured from a UAV with a top-down view. This involves ensuring that the crack images are appropriately sized and positioned to blend seamlessly into the road surface. The end result is the road image with its binary mask. 2 Literature Review Over the past years, extensive research has been conducted on general crack detection, segmentation, and classification using various approaches. These include
4 A. Panagi and C. Kyrkou image processing techniques [17], [10], [26], model-based methods [8], [14], [16] and synthetic data generation [12], [19]. Initially, using conventional image processing, A. Mohan and S. Poobal present a thorough analysis of 50 studies focused on the automatic identification of concrete surface cracks and their depth using image processing methods [17]. In addition to describing the various image processing methods according to the kind of image utilized, the survey also does a thorough study, paying particular attention to the accuracy and error rates of each approach. However, such methods have been replaced by machine learning approaches that offer improved accuracy and efficiency. Specifically deep learning techniques dominate in crack detection. This can be seen in both [8] and [14], where Convolutional Neural Networks (CNNs) were employed for pixel-level semantic segmentation of road crack damages. These studies concentrate on the development of deep neural networks for crack detection on road surface images, which is a widely adopted approach among researchers. Additionally, a comparative study presented in [24] evaluates state-of-the-art deep learning models for semantic crack detection in construction materials. The authors evaluate the performance of various existing algorithms in capturing fine details for crack segmentation and suggest that generating realistic synthetic image data could be a potential solution to mitigate data scarcity issues during model training. 2.1 Relevant Synthetic Data Generation Methods Kanaeva and Ivanova in [12] generated an image dataset by imposing road crack images on top of backgrounds taken from publicly available datasets such as KITTI and Cityscapes. They used this dataset to train UNet [21] and Mask R-CNN [9] models. When tested on real images of cracked roads, these models provided an mIoU score of 47%. However, the mismatch of background and overlaid image perspective adopted in their methodology is a limitation to the realism of the synthesised dataset. In addition, Rill-García et. al. [19] proposed a new tool, named "Syncrack", for creating images with accurate labels of cracks to improve the accuracy of pixel-level prediction. For generating images, the synthesised crack shapes were overlaid onto custom-made background images. To produce a fully synthetic dataset, noisy annotations were generated using the pixel-accurate crack masks. The segmentation models, such as UNet, that were trained on this dataset, performed similarly to those that were trained on real world datasets, providing greater precision, particularly in detecting the crack width. Aditionally synthetic data augmentation for dam crack detection is proposed in [29]. The paper presented and introduced a concept similar to the one developed for this research but this time for crack detection on the surface of dams. The method used the generation of synthetic crack images and training a deep neural network, yielding good results, which further demonstrates how synthetic data increases the ability to train models when real data may not be accessible or sufficient.
Synthetic Data for Crack Segmentation from UAV Imagery 5 Other methods, such as Generative Adversarial Networks (GANs), have been popular recently, since the success of these models has motivated many researchers. A semantically-driven generative adversarial network for generating crack images is called CrackGauGAN proposed in recent work [3]. This network can generate realistic-looking crack images that are high in fidelity with diverse data distributions. This not only outperforms the latest semantic layout-based GANs but also shows greater improvement in generating crack images. Although the aforementioned methods propose state-of-the-art improvements in the synthesis of images, they do not adequately address key factors such as background noise and the unique challenges posed by top-down UAV imagery, which is crucial for road surface analysis. Moreover, while individual studies have demonstrated improvements in segmentation using synthetic data, there is limited comparative analysis on how different synthetic data generation techniques impact model performance in real-world scenarios. Our study proposes a comprehensive comparative research, evaluating multiple segmentation models and synthetic data configurations to determine the most effective approach for robust road crack segmentation with UAV imagery. By directly comparing model performance across different datasets and loss functions, we aim to provide insights into how synthetic-to-real transfer can be optimized for UAV-based infrastructure monitoring. 2.2 Existing Crack Segmentation Datasets Crack segmentation datasets have significantly evolved and expanded over time. Most of them are employed as benchmarks for tasks involving segmentation, detection, and classification. Nevertheless, there are some limitations to the existing crack segmentation datasets. One drawback is that models are unable to learn much about how the crack might relate to the surrounding environment because most images of the already existing datasets lack a contextual background. Scale inconsistencies, limited variability, and differences in viewpoints—along with most images being captured by vehicle-mounted cameras instead of UAVs—are common issues in these datasets, reducing their generalization power in real-world applications. In the following, we provide an overview of some existing datasets which can be leveraged to create synthetic images from a UAV perspective (also summarized in Table 1). -CRACK500 Dataset: The CRACK500 Dataset [30] consists of more than 500 high-resolution pavement images collected by using a smartphone as the data sensor. Each image was annotated by multiple annotators. The dataset contains pavement cracks of variable widths. -Crack Segmentation Dataset by Roboflow: There are 4029 images of concrete damage in the dataset [20]. The images in the dataset have undergone augmentations such as rotation, saturation, and brightness adjustments. Although the dataset consists solely of concrete crack damage, the
6 A. Panagi and C. Kyrkou Table 1. Summary of crack segmentation datasets. Dataset View Captured With Number of Images Image Size CRACK500 [30] Top view Smartphone camera 500 3264x2448 Crack Segmentation Dataset by Roboflow [20] Top view - 4029 416x416 CrackTree200 [31] Top view - 206 800x600 CFD [23] Top view Smartphone Camera 118 480x320 GAPs [6] Top view Vehicle Mounted Camera 1969 1920x1080 DeepCrack [15] Top view - 537 544x384 CrackSeg9k [13] Multiple views Multiple Methods 9255 400x400 OmniCrack30k [2] Multiple views Multiple Methods 30K Multiple image sizes cracks can still be applied to pavement scenarios, as they have been processed in a way that allows them to effectively mimic pavement cracks after overlaying. -CrackTree200 Dataset: The CrackTree200 Dataset consists of 206 pavement images, each with dimensions of 800×600 pixels, featuring a variety of crack types annotated at pixel level. The dataset captures images in complex asphalt environments, characterized by challenges such as shadows, occlusions, low contrast, and noise [31]. -CFD Dataset: The dataset includes 118 annotated images of road cracks, each of 480×320 pixels [23]. These images were captured on urban roads in Beijing and exhibit a significant presence of noisy pixels, such as oil spots and water stains. Additionally, some images were taken under poor lighting conditions, further complicating the annotation process. -GAPs Dataset: A total of 1969 grayscale images with a resolution of 1920×1080 pixels are included in [6]. The images have been annotated manually. The surface material depicted in the images includes pavement from three different German federal roads. -DeepCrack Dataset: This dataset [15] entails 537 images with manual annotation maps. Each image is made available to a pixel-wise segmentation map, which presents to be a mask exactly covering the crack regions. All of the images are of a fixed size of 544×384 pixels. -CrackSeg9k Dataset: With 9255 images that aggragate other smaller open-source datasets, the dataset [13] is one of the biggest, most varied,
Synthetic Data for Crack Segmentation from UAV Imagery 7 and reliable crack segmentation collection. It consists of 10 sub-datasets namely Crack500 [30], DeepCrack [15], SDNET2018 [5], CrackTree200 [31], GAPs [6], Volker, Rissbilder [18], CFD [23], Masonry [4], and Ceramic [22], with images being preprocessed and resized to 400×400. -OmniCrack30K: It is the largest benchmark for crack segmentation containing nearly 30K images from over 20 datasets [2]. The dataset features cracks found in various materials, including asphalt, ceramic, concrete, masonry, and steel. Although UAV images are included in this dataset, they do not depict road cracks. The existing datasets primarily focus on segmenting and detecting cracks in structures or roads using close-up images captured by ground-based or vehicle cameras, often neglecting contextual information. In contrast, our dataset provides images with annotations based on aerial imagery (captured with a UAV), incorporating contextual information from the surrounding areas near the roads. This broader perspective enhances its applicability for real-world crack detection tasks. 3 Methodology 3.1 Synthetic Image Data Generation Creating high-quality road crack segmentation datasets requires pixel-level annotations, which is a highly time-consuming process as each pixel corresponding to a crack must be accurately labeled in the image. Alternatively, generating synthetic image data using auxiliary datasets can substantially reduce the time and effort required for manual annotation, making it a highly efficient and scalable solution. Obtaining relevant background images that will be used as the basis for superimposing the crack damage is crucial for the generation of the synthetic images (Figure 2). In our case, top-down realistic road infrastructure images were captured by a UAV (with an altitude of 120m) to provide a comprehensive view of the surface. Along with these, images of cracks and their corresponding binary masks were sourced from existing annotated datasets, though these originated from different contexts. The binary masks were used to isolate the cracks, allowing their integration onto our own road images. To ensure diversity in crack damages we combined multiple crack images sourced from different publicly available datasets such as Crack500 [30], DeepCrack [15], CrackTree200 [31], GAPs [6], Volker, Rissbilder [18], CFD [23], and the Crack Segmentation Dataset by Roboflow [20].
8 A. Panagi and C. Kyrkou Fig. 2. Examples of crack image overlays: (a) The image displays the superimpose of cracks without any pre-processing performed to the crack overlay. (b) The crack images are scaled and then overlaid. (c) Both scaling and blending techniques are used during the deployment of cracks on the background images. Alongside each image, the corresponding binary mask is generated.The scaling and blending techniques used in (b) and (c) improve the visual coherence of the cracks with the background, making them appear more natural with the road. The process of overlaying cracks onto background images starts with first isolating the road segment. This ensures that cracks are applied only to the relevant road area, resulting in a more realistic outcome. A segmented mask of the road from the background images is used to guide the synthetic crack deployment. Valid positions (e.g., all (x,y) pixel points of the image that belong to the road area) for crack placement are determined to avoid unrealistic positioning, such as cracks appearing on non-road areas. The crack image is adjusted to match the color and scale of the road, ensuring a natural integration into the background. In detail, an inverse mask is created from the original binary mask of the crack. This inverse mask is used to black out the area where the crack will be applied, preserving the rest of the region in the background image. The inverse mask is applied to the region of interest (ROI) 1in the background, ensuring that only the background area outside the crack region remains visible, while the area where the crack will be placed is blacked out. The crack image is further isolated by applying the binary mask to it, keeping only the crack region visible and blacking out the rest. The crack region is then resized to match the 1The region of interest refers to the specific portion of the background image where the crack will be placed. It is defined by the coordinates (x, y), which represent the top-left corner of the area, and the dimensions h (height) and w (width) of the crack image. By slicing the background image from y to y+h (vertical range) and from x to x+w (horizontal range), the function extracts the area of the background where the crack will be overlaid, ensuring that the crack fits within this region.
Synthetic Data for Crack Segmentation from UAV Imagery 9 size of the background ROI, ensuring consistent dimensions between the crack and the background areas. Finally, the crack region is added to the blacked-out ROI from the background, effectively overlaying the crack on the background. This combined region is then placed back into the original background image at the specified position, resulting in a seamless integration of the crack into the background. During this process, certain images containing vehicles on the road were identified. Since overlaying cracks on the cars would result in unrealistic damage of the road, these images were removed from the dataset so that the accuracy is maintained. Additionally, cracks were superimposed on road images where the roads had no natural cracks. The entire process of generating synthetic images with superimposed cracks is fully automated in a form of a pipeline. Human intervention is only required to manually remove images where cracks are mistakenly applied to cars, ensuring realistic results. Synthetic images are generated in different ways of varying complexity. In the first case (referred to as the Raw Cracks Dataset), crack damage was superimposed directly onto the background images without any further pre-processing. In the second case, the crack overlays (referred to as Cracks with Scaling Dataset) were subsequently scaled to pixel values between the range 100-150 in order to better match the road color statistics. In the third case, referred to as Cracks with Scaling and Blending Dataset, in addition to scaling, a blending technique is used by doing a weighted sum of the initial pixel value with the crack image value, to ensure that the crack damages mix seamlessly with the road surface background. 3.2 UNet as the Baseline Network The UNet architecture [21] comprises a down-sampling and an up-sampling path, where the encoder layers in the down-sampling direction reduce the spatial resolution of the input while capturing contextual information. In the up-sampling direction, the decoder layers reconstruct the segmentation map by decoding the encoded data and incorporating information from the down-sampling path through skip connections. For crack segmentation, UNet processes a full-resolution input image (e.g., 512 ×512 pixels). During the down-sampling phase, the encoder progressively reduces the spatial dimensions of the image while increasing the number of channels by repeatedly applying 3 ×3 convolutional layers and max-pooling operations. This enables the encoder to capture hierarchical feature representations of the input. The up-sampling path takes the feature map from the bottleneck and reconstructs it to match the original input size using 2 ×2 up-convolution layers, followed by two 3 ×3 convolutional layers. These operations increase the spatial resolution while decreasing the number of channels. Skip connections from the down-sampling path are integrated into the up-sampling path to help the decoder accurately localize and refine features. Ultimately, the output is a binary segmentation map in which each pixel is classified as either a crack or part of the background.
16 A. Panagi and C. Kyrkou Fig. 5. An illustration of crack segmentation results on real test data, using models trained on the "Cracks with Scaling and Blending" synthetic dataset. From top to bottom: original image, ground truth, detected cracks using UNet-EfficientNet and UNet-MobileNet with the three different loss functions. 5.5 Further Fine-Tuning of Models on Real Crack Images The best seven models in terms of accuracy and mIoU presented in Table 2 are selected for further fine-tuning using a dataset with human-annotated road crack damages. The fine-tuning process involved freezing the initial layers of the model’s encoder and training the model for additional 10 epochs on the finetuning set which entails 43 images containing real road crack damages annotated by human experts.
Synthetic Data for Crack Segmentation from UAV Imagery 17 The outcomes of the metrics for the fine-tuned models are displayed in Table 4. While some of these models, before fine-tuning, such as the UNet-MobileNet Focal, Scaling, had low crack pixel accuracy, 0.2389, upon fine-tuning this model was able to achieve the best crack pixel accuracy of 0.6358, significantly improving it. With the best F1-Score of 0.7439 and the highest mIoU value of 0.6564, the UNet-EfficientNet Focal, Scaling and Blending model demonstrated an optimum precision-recall balance. Generally, the models fine-tuned with scaling and blending had generated better F1-score and mIoU values compared to those trained with raw inputs or scaling alone. The comparison of loss functions shows that while Focal loss models, including UNet-EfficientNet Focal Scaling and Blending, achieved a better overall performance balance, WBCEL-trained models, like UNet-MobileNet WBCEL, Scaling, excelled in pixel crack accuracy of 0.6358 and in recall of 0.8053 (Figure 6). Our experiments emphasize the importance of combining synthetic data with human annotations as a more effective strategy for training machine learning models in low-data regime applications. Fig. 6. Results of best fine-tuned model (UNet-MobilenetNet / WBCEL / Scaling) tested on real crack images depicting crack damages on the road. From left to right: original image, ground truth binary mask, predicted binary mask.
18 A. Panagi and C. Kyrkou Table 4. Results of fine-tuned segmentation models on a real crack test set. Model / Loss / Dataset Pixel Accuracy (Crack) mIoU Precision Recall F1-Score UNet-MobilenetNet / WBCEL / Scaling 0.6358 0.6168 0.6607 0.8053 0.7016 UNet-MobileNet / Focal / Scaling 0.6033 0.6301 0.6803 0.7907 0.7157 UNet-EfficientNet / Focal / Scaling and Blending 0.5846 0.6564 0.7204 0.7918 0.7439 UNet-MobilenetNet / Focal / Raw 0.5759 0.6043 0.6524 0.7716 0.6866 UNet-EfficientNet / WBCEL / Scaling and Blending 0.5572 0.6496 0.7173 0.7781 0.7363 UNet-MobileNet / WBCEL / Scaling and Blending 0.5056 0.6551 0.7451 0.7521 0.7413 UNet-MobileNet / Focal / Scaling and Blending 0.4925 0.6452 0.7288 0.7503 0.7304 6 Conclusion and Future Work This research explores the feasibility of training state-of-the-art deep learning models with synthetic data for road crack segmentation from UAV images and assessing their performance in real-world scenarios. In order to provide synthetic data that can serve as the foundation for training a segmentation model, we employed image-based techniques to incorporate externally sourced crack damages onto road surfaces. We then trained deep learning models based on the UNet architecture, and assessed the impact of synthetic datasets on segmentation performance. When tested on images with real crack damage, the best model achieved an mIoU of 63.52% and crack pixel accuracy of 33.26%. Further fine-tuning with a real crack dataset resulted in an increase in crack pixel accuracy to 63.58%, while the mIoU remained stable at approximately 61.68%. Future work may include generating more realistic crack patterns using GANs or diffusion models, expanding hybrid datasets to improve generalizability, and fine-tuning with larger and more varied datasets. Further, deploying these models on UAVs for real-time crack detection could be used to expand automated road maintenance and infrastructure inspection. Acknowledgments Funded by the European Union’s Horizon Europe Research and Innovation Actions under grant agreement No. 101147850 - EvoRoads. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union. Neither the European Union nor the granting authority can be held responsible for them. References 1. Ashleigh, C. (n.d.). Road traffic injuries. Health, WHO. Retrieved September 26, 2024, from https://www.who.int/health-topics/road-safety#tab=tab_1
Synthetic Data for Crack Segmentation from UAV Imagery 19 2. Benz, C., Rodehorst, V., (2024). OmniCrack30k: A Benchmark for Crack Segmentation and the Reasonable Effectiveness of Transfer Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 3876-3886. Retrieved from: https://openaccess.thecvf.com/content/CVPR2024W/VAND/papers/Benz_ OmniCrack30k_A_Benchmark_for_Crack_Segmentation_and_the_Reasonable_ Effectiveness_CVPRW_2024_paper.pdf 3. Chu, H., Chen, W., & Deng, L. (2024). CrackGauGAN: Semantic layoutbased crack image synthesis for automated crack segmentation. The International Association for Automation and Robotics in Construction (IAARC). https://doi.org/10.22260/isarc2024/0002 4. Dais, D., Bal, İ. E., Smyrou, E., & Sarhosis, V. (2021). Automatic crack classification and segmentation on masonry surfaces using convolutional neural networks and transfer learning. Automation in Construction, 125, 103606. https://doi.org/10.1016/j.autcon.2021.103606 5. Dorafshan, S., Thomas, R. J., & Maguire, M. (2018). SDNET2018: An annotated image dataset for non-contact concrete crack detection using deep convolutional neural networks. Data in Brief, 21, 1664–1668. https://doi.org/10.1016/j.dib.2018.11.015 6. Eisenbach, M., et al. (2017). How to get pavement distress detection ready for deep learning? A systematic approach. In International Joint Conference on Neural Networks (IJCNN), 2039–2047. https://doi.org/10.1109/IJCNN.2017.7966101. 7. Gavilán, M., et al. (2011). Adaptive Road Crack Detection System by pavement Classification. Sensors, 11(10), 9628–9657. https://doi.org/10.3390/s111009628 8. Han, C., Ma, T., Huyan, J., Huang, X., & Zhang, Y. (2021). CrackW-Net: A novel pavement crack image segmentation convolutional neural network. IEEE Transactions on Intelligent Transportation Systems, 23(11), 22135–22144. https://doi.org/10.1109/tits.2021.3095507. 9. He, K., Gkioxari, G., Dollar, P., & Girshick, R. (2018). Mask R-CNN. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(2), 386–397. https://doi.org//10.1109/TPAMI.2018.2844175 10. Hoang, N. (2018). Detection of Surface Crack in Building Structures Using Image Processing Technique with an Improved Otsu Method for Image Thresholding. Advances in Civil Engineering,. https://doi.org/10.1155/2018/3924120 11. Howard, A., et al. (2019). Searching for MobileNetV3. arXiv.org. Retrieved from: https://arxiv.org/abs/1905.02244 12. Kanaeva, I. A., & Ivanova, J. A. (2021). Road pavement crack detection using deep learning with synthetic data. IOP Conference Series: Materials Science and Engineering, 1019(1), 012036. https://doi.org/10.1088/1757-899X/1019/1/012036. 13. Kulkarni, S., Singh, S., Balakrishnan, D., Sharma, S., Devunuri, S., & Korlapati, S. C. R. (2023). CrackSeg9K: a collection and benchmark for crack segmentation datasets and frameworks. In Lecture notes in computer science, 179–195. https://doi.org/10.1007/978-3-031-25082-8_12. 14. Lau, S. L. H., Chong, E. K. P., Yang, X., & Wang, X. (2020). Automated pavement crack segmentation using U-Net-based Convolutional Neural Network. IEEE Access, 8, 114892–114899. https://doi.org/10.1109/access.2020.3003638. 15. Liu, Y., Yao, J., Lu, X., Xie, R., & Li, L. (2019). DeepCrack: A deep hierarchical feature learning architecture for crack segmentation. Neurocomputing, 338, 139–153. https://doi.org/10.1016/j.neucom.2019.01.036.
20 A. Panagi and C. Kyrkou 16. Mandal, V., Uong, L., & Adu-Gyamfi, Y. (2018). Automated Road Crack Detection Using Deep Convolutional Neural Networks. 2018 IEEE International Conference on Big Data, 5212-5215. https://doi.org/10.1109/bigdata.2018.8622327 17. Mohan, A., & Poobal, S. (2018). Crack detection using image processing: A critical review and analysis. Alexandria Engineering Journal, 57(2), 787–798. https://doi.org/10.1016/j.aej.2017.01.020. 18. Pak, M., & Kim, S. (2021). Crack detection using fully convolutional network in Wall-Climbing robot. In Lecture notes in electrical engineering, 267–272. https://doi.org/10.1007/978-981-15-9343-7_36 19. Rill-García, R., Dokladalova, E., & Dokládal, P. (2022). Syncrack: Improving pavement and concrete crack detection through synthetic data generation. Retrieved from https://hal.science/hal-03451685/. 20. Roboflow Team. (2024). Crack Instance Segmentation Dataset and Pretrained Model. Retrieved September 27, 2024 from: https://universe.roboflow.com/ university-bswxt/crack-bphdr?ref=ultralytics 21. Ronneberger, O., Fischer, P., & Brox, T. (2015). U-Net: Convolutional Networks for Biomedical Image Segmentation. In Lecture notes in computer science, 234–241. https://doi.org/10.1007/978-3-319-24574-4_28. 22. Santos, G., Junior, Ferreira, J., Millán-Arias, C., Daniel, R., Casado, A., Junior, & Fernandes, B. J. T. (2021). Ceramic Cracks Segmentation with Deep Learning. Applied Sciences, 11(13), 6017. https://doi.org/10.3390/app11136017 23. Shi, Y., Cui, L., Qi, Z., Meng, F., & Chen, Z. (2016). Automatic road crack detection using random structured forests. IEEE Transactions on Intelligent Transportation Systems, 17(12), 3434–3445. https://doi.org//10.1109/tits.2016.2552248. 24. Shi, Z., Jin, N., Chen, D., & Ai, D. (2024). A comparison study of semantic segmentation networks for crack detection in construction materials. Construction and Building Materials, 134950. https://doi.org/10.1016/j.conbuildmat.2024.134950 25. Silva, L. A., Leithardt, V. R. Q., Batista, V. F. L., Villarrubia González, G., & De Paz Santana, J. F. (2023b). Automated road damage detection using UAV images and deep learning techniques. IEEE Access, 11, 62918–62931. https://doi.org/10.1109/access.2023.3287770. 26. Talab, A. M. A., Huang, Z., Xi, F., & HaiMing, L. (2015). Detection crack in image using Otsu method and multiple filtering in image processing techniques. Optik, 127(3), 1030–1033. https://doi.org/10.1016/j.ijleo.2015.09.147 27. Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. arXiv.org. Retrieved from : https://arxiv.org/abs/1905. 11946. 28. Varadharajan, S., Jose, S., Sharma, K., Wander, L., & Mertz, C. (2014). Vision for road inspection. IEEE Winter Conference on Applications of Computer Vision, 115–122. https://doi.org/10.1109/wacv.2014.6836111 29. Xu, J., Yuan, C., Gu, J., Liu, J., An, J., & Kong, Q. (2022). Innovative synthetic data augmentation for dam crack detection, segmentation, and quantification. Structural Health Monitoring, 22(4), 2402–2426. https://doi.org/10.1177/14759217221122318. 30. Zhang, L., Yang, F., Zhang, Y. D., & Zhu, Y. J. (2016). Road crack detection using deep convolutional neural network. In IEEE International Conference on Image Processing (ICIP), 3708–3712. https://doi.org/10.1109/icip.2016.7533052. 31. Zou, Q., Cao, Y., Li, Q., Mao, Q., & Wang, S. (2011). CrackTree: Automatic crack detection from pavement images. Pattern Recognition Letters, 33(3), 227–238. https://doi.org/10.1016/j.patrec.2011.11.004.