scieee AI-readable full text Open interactive document viewer

Multichannel Convolutional Networks for Classification and Localization of Construction and Demolition Waste

Zbíral, Tomáš; Nežerka, Václav

Abstract

Effective management of construction and demolition waste (CDW) remains a critical concern for both environmental sustainability and economic efficiency. Inadequate sorting mechanisms result in suboptimal resource use and diminished recycling opportunities. In response, this paper presents a sophisticated machine learning approach that utilizes multichannel Convolutional Neural Networks (CNNs) to enhance the sorting and recycling of CDW. Our methodology incorporates a two-phase deep learning model, beginning with a U-Net architecture for detailed, pixel-level segmentation of waste materials. This initial phase is essential for accurately identifying diverse waste types within complex debris captured via RGB imaging. Following this, a Residual Network (ResNet) based CNN is applied to achieve reliable multi-class classification of the segmented debris. Our comparative analysis shows that the U-Net framework surpasses traditional background subtraction techniques in Python, providing more efficient and accurate segmentation. Early findings indicate notable advancements in both classification and localization precision, facilitating improved CDW sorting and recycling processes. Additionally, this research aids in the progression towards real-time automated waste management systems, thus promoting practices that enhance environmental sustainability. To encourage ongoing research and practical implementations, we have made available the codes, training and testing datasets, and additional resources related to this study.

Full text

Multichannel Convolutional Networks for Classification and Localization of Construction and Demolition Waste Tom´ aˇ s Zb´ ıral[0000−0002−6977−7796], V´ aclav Neˇzerka[0000−0001−7360−0507] Abstract Effective management of construction and demolition waste (CDW) remains a critical concern for both environmental sustainability and economic efficiency. Inadequate sorting mechanisms result in suboptimal resource use and diminished recycling opportunities. In response, this paper presents a sophisticated machine learning approach that utilizes multichannel Convolutional Neural Networks (CNNs) to enhance the sorting and recycling of CDW. Our methodology incorporates a two-phase deep learning model, beginning with a U-Net architecture for detailed, pixel-level segmentation of waste materials. This initial phase is essential for accurately identifying diverse waste types within complex debris captured via RGB imaging. Following this, a Residual Network (ResNet) based CNN is applied to achieve reliable multi-class classification of the segmented debris. Our comparative analysis shows that the U-Net framework surpasses traditional background subtraction techniques in Python, providing more efficient and accurate segmentation. Early findings indicate notable advancements in both classification and localization precision, facilitating improved CDW sorting and recycling processes. Additionally, this research aids in the progression towards real-time automated waste management systems, thus promoting practices that enhance environmental sustainability. To encourage ongoing research and practical implementations, we have made available the codes, training and testing datasets, and additional resources related to this study. Tom´ aˇ s Zb´ ıral CTU Prague, Faculty of Civil Engineering, Th´ akurova 2077/7, 166 29 Praha 6 e-mail: zbiratom@ cvut.cz V´ aclav Neˇzerka CTU Prague, Faculty of Civil Engineering, Th´ akurova 2077/7, 166 29 Praha 6 e-mail: vaclav. [email protected] 1 2 Tom´ aˇ s Zb´ ıral, V´ aclav Neˇzerka 1 Introduction Pursuing sustainable development necessitates prudent and cost-effective waste management practices that align with the principles of the circular economy [1, 2]. To this end, the European Parliament and the Commission enacted Directive No 98/2008, mandating EU member states to achieve a minimum of 70% recycling by weight by 2020. While the average recycling rate for construction and demolition waste (CDW) in the EU hovers around 90%1, the majority of these materials are subjected to downcycling. On a global scale, developing nations such as China, generating approximately 2 billion tons of CDW annually, exceed the combined output of all EU states [3]. Among CDW, the most commonly recycled materials are soils, concrete, and ceramics, typically repurposed for use in embankments, backfills, and foundation underlays. Occasionally, these materials are recycled as aggregates for new concrete mixes or as micro-fillers [4–7]. The primary obstacle to valorizing crushed CDW in higher-value applications like concrete production is inadequate sorting [8]. Studies, such as the multi-agent evolutionary game analysis by Su [9], have underscored the significant potential of CDW classification research in enhancing recycling and reuse. Furthermore,Davis et al. [10] emphasized that automating the classification of CDW materials could drastically reduce sorting costs. Recent advancements have seen the development of various sophisticated sensing technologies (image, spectroscopic, spectral, UV sensitive, etc.) for waste sorting [11, 12]. However, at an industrial level, sorting is predominantly manual, hampered by the physical similarities of the waste fragments. There is thus a growing impetus to replace manual sorting with machine learning-assisted robotic vision technologies, such as RGB cameras, hyperspectral imaging, and X-ray sensors. Initially applied in municipal waste separation, these technologies have achieved sorting accuracies exceeding 90% [13]. CDW sorting has also begun to benefit from these technological advancements [14, 15]. Yet, challenges in achieving high accuracy and precise material boundary detection remain. Addressing this issue,Dong et al. [16] developed a boundary-aware model capable of identifying and segmenting distinct materials within structural debris. CNNs, highly adept at recognizing hierarchical visual patterns, are pivotal for enhancing classification accuracy in complex scenarios like CDW sorting. For example, Xiao et al. [17] successfully employed CNNs to classify various CDW materials—such as wood, brick, and concrete—with accuracies surpassing 80%. Similarly, Ku et al. [18] demonstrated the effectiveness of hyperspectral and 3D imaging in automated CDW sorting, achieving accuracy rates around 90%. Additionally, Lin et al. [19] utilized machine learning to differentiate visually distinct CDW fragments, reaching accuracy levels between 75 and 80%. Notably, Hoong et al. [8] used neural networks for classifying recycled aggregates, achieving up to 97% accuracy with a substantial image library. 1https://ec.europa.eu/eurostat/databrowser/view/cei_wm040/default/table Multichannel CNNs for Classification and Localization of CDW 3 Our research differentiates itself by focusing on advanced feature extraction using standard RGB cameras—a method not thoroughly investigated in previous studies or lacking the required accuracy. This paper outlines a novel methodology for the rapid segmentation and precise classification of CDW using U-Net and ResNet models. We provide extensive datasets, source codes, and pre-trained models to ensure that our approach is transparent, reproducible, and extendable by both the research community and industry practitioners. 2 Data Preparation The data preparation phase is crucial in the development of deep neural networks as it greatly influences the models’ performance, efficiency, and generalization ability. Effective data preparation encompasses several critical processes, including data cleaning, augmentation, and splitting, each enhancing the model’s robustness and quality. 2.1 Data Cleaning The initial dataset, as discussed in our previous study [20], was unsuitable for automatic segmentation due to various issues: images captured from mixed material piles and the inclusion of irrelevant objects like structures and human elements. Through rigorous data cleaning, we ensured the reliable generation of ground truth masks, crucial for accurate segmentation modeling. From the original compilation of 2,664 images, a refined set of 2,133 images remained, segmented as follows: 538 images of AAC, 315 images of asphalt, 622 images of ceramics, and 658 images of concrete (Fig. 1). 2.2 Ground Truth Masks For effective supervised learning in segmentation tasks, it is essential to create accurate ground truth masks delineating the regions of interest (Fig. 2). We utilized the rembg library, a robust tool for background removal across diverse images. The remove function within rembg leverages multiple models such as U-2-Net [22], IsNet [23], and Sam [24], ensuring reliable and computationally intensive background extraction. However, it must be noted that this approach can be used only for preparing the dataset as the segmentation is computationally extremely demanding, as described next. Each generated mask was visually inspected. 4 Tom´ aˇ s Zb´ ıral, V´ aclav Neˇzerka AACAsphaltCeramicsConcrete Fig. 1 Examples of image datasets for the examined CDW [21]. AAC Asphalt Ceramics Concrete Fig. 2 Paired images and corresponding masks created by the rembg’s remove function for each material type. 2.3 Data Augmentation Data augmentation plays a pivotal role in enhancing a model’s generalization, crucial when dataset sizes are limited. This technique involves generating new samples through various transformations applied to existing data, such as rotation, flipping, scaling, translation, and noise addition (e.g., blur, brightness changes). The primary benefits include improved model generalization to prevent overfitting, increased dataset diversity, and enhanced robustness—key for applications like image recognition under diverse and unpredictable conditions. In our study, we employed the albumentations library [25], a powerful and flexible tool designed for efficient data augmentation in deep learning contexts. This library provides an extensive array of augmentation techniques tailored for image data, including basic transformations (flips, rotations, scaling), color adjustments, Multichannel CNNs for Classification and Localization of CDW 5 and advanced manipulations (cutout, grid distortions, elastic transformations). It seamlessly integrates with deep learning frameworks like PyTorch [26], used extensively in our project. Below is a code snippet illustrating our data augmentation pipeline: 1def augment_image(image, mask): 2# Define the augmentation pipeline for both image and mask 3transform =A.Compose([ 4A.Rotate(limit=25, p=1), 5A.HorizontalFlip(p=0.5), 6A.VerticalFlip(p=0.5), 7A.RandomBrightnessContrast(p=0.3), 8A.OneOf([ 9A.Blur(blur_limit=5, p=0.2), 10 A.GaussianBlur(blur_limit=5, p=0.2), 11 ], p=0.2), 12 ToTensorV2(), 13 ]) The finalized dataset, post-cleaning and augmentation, comprised 4,266 images, half of which were augmented, to train the U-Net and ResNet models. 2.4 Data Splitting Next, we partitioned the dataset into three subsets: training, validation, and test. The training set is fundamental for developing the neural network’s core capabilities, while the validation set assists in fine-tuning hyperparameters and selecting the most effective model architecture. Iterative adjustments based on the feedback from the validation set after each epoch were instrumental in enhancing the model’s generalization capabilities and in mitigating overfitting. The test set was utilized to provide an unbiased assessment of the model’s performance under conditions that simulate real-world scenarios. We utilized a 70/15/15 split ratio for the segmentation and classification tasks to balance the extensive need for training data with sufficient validation and testing, ensuring the model’s robustness and reliability in practical applications. 3 Segmentation Due to high demands of the rembg model, it was essential to train own dedicated segmentation model based on U-Net architecture [27] to meet the demands of realtime segmentation. Runtime comparisons conducted on both models with a batch of 100 images revealed that the rembg model, while comprehensive, fell short in real-time applicability with a total runtime of 840 seconds, averaging 8.4 seconds 6 Tom´ aˇ s Zb´ ıral, V´ aclav Neˇzerka per image. In contrast, our custom U-Net model significantly outstripped the rembg, clocking in at merely 0.37 seconds per image—a 22.7×faster. This marked increase in speed is pivotal for applications requiring real-time processing. 3.1 U-Net Architecture In our adaptation of the U-net architecture, we incorporated batch normalization, an enhancement not utilized in the original U-Net framework. The architecture of our U-Net model is detailed in the following code snippet: 1class DoubleConv(nn.Module): 2def __init__(self, in_channels, out_channels): 3super(DoubleConv, self).__init__() 4self.conv =nn.Sequential( 5nn.Conv2d(in_channels, out_channels, kernel_size=3, stride=1, padding=1, bias=False),↩→ 6nn.BatchNorm2d(out_channels), 7nn.ReLU(inplace=True), 8nn.Conv2d(out_channels, out_channels, kernel_size=3, stride=1, padding=1, bias=False),↩→ 9nn.BatchNorm2d(out_channels), 10 nn.ReLU(inplace=True), 11 ) 12 13 def forward(self, x): 14 return self.conv(x) 15 16 class Unet(nn.Module): 17 def __init__(self, in_channels=3, out_channels=1, features=[64, 128,256,512]):↩→ 18 super(Unet, self).__init__() 19 self.downs =nn.ModuleList() 20 self.ups =nn.ModuleList() 21 self.pool =nn.MaxPool2d(kernel_size=2, stride=2) 22 23 # Down part of Unet 24 for feature in features: 25 self.downs.append(DoubleConv(in_channels, feature)) 26 in_channels =feature 27 28 # Up part of Unet 29 for feature in reversed(features): 30 self.ups.append( 31 nn.ConvTranspose2d( 32 feature*2, feature, kernel_size=2, stride=2, padding=0↩→ 33 ) 34 ) 35 self.ups.append(DoubleConv(feature*2, feature)) Multichannel CNNs for Classification and Localization of CDW 7 36 37 self.bottleneck =DoubleConv(features[-1], features[-1]*2) 38 self.final_conv =nn.Conv2d(features[0], out_channels, kernel_size=1)↩→ 39 40 def forward(self, x): 41 skip_connections =[] 42 for down in self.downs: 43 x=down(x) 44 skip_connections.append(x) 45 x=self.pool(x) 46 47 x=self.bottleneck(x) 48 # Now we need to go up 49 skip_connections =skip_connections[::-1] 50 51 for idx in range(0,len(self.ups), 2): 52 x=self.ups[idx](x) 53 skip_connection =skip_connections[idx//2] 54 55 if x.shape != skip_connection.shape: 56 x=TF.resize(x, size=skip_connection.shape[2:]) 57 58 concat_skip =torch.cat((skip_connection, x), dim=1) 59 x=self.ups[idx+1](concat_skip) 60 61 return self.final_conv(x) 3.2 Segmentation Accuracy Metrics Segmentation accuracy is pivotal for evaluating the efficacy of image segmentation models. This chapter delves into two widely employed metrics for gauging segmentation performance: the F1 score (𝐹1) and Intersection over Union (𝐼𝑜𝑈). These metrics offer a balanced and comprehensive assessment of a model’s segmentation accuracy. 𝐹1 is the harmonic mean of precision and recall, defined as 𝐹1=2𝑃 𝑅 𝑃+𝑅,(1) where 𝑃is precision, represented as 𝑃=𝑃true 𝑃true+𝑃false , denoting true positives and false positives respectively, and 𝑅is recall, defined as 𝑅=𝑃true 𝑃true+𝑁false , where 𝑁false denotes false negatives. 𝐼𝑜𝑈 metric, also known as the Jaccard index, quantifies the overlap between the predicted and the actual segmentation areas, defined as: 𝐼𝑜𝑈 = |𝐴∩𝐵| |𝐴∪𝐵|,(2) 8 Tom´ aˇ s Zb´ ıral, V´ aclav Neˇzerka where 𝐴and 𝐵represent the sets of pixels in the predicted and ground truth regions, respectively. The numerator |𝐴∩𝐵|indicates the intersection area, while the denominator |𝐴∪𝐵|denotes the union of both regions, highlighting the extent of overlap and exclusion. 3.3 Optimization and Training of U-Net Model The U-Net model was optimized using the sigmoid activation function, binary crossentropy with logits for the loss function, and the Adam optimizer with a learning rate of 0.001. The training was conducted over 35 epochs. The progression of training and validation loss, alongside the accuracy metrics across epochs, is detailed in Figs. 3, 4, and 5. The validation loss shows a plateau around epoch 15, indicating the beginning of model convergence. Optimal performance was observed at epoch 32, with the validation 𝐼𝑜𝑈 peaking at 0.902 and 𝐹1 score at 0.928. The highest performance metrics on the test dataset were also noted during this epoch, with an 𝐼𝑜𝑈 of 0.863 and an 𝐹1 of 0.941. The variability in 𝐼𝑜𝑈 across different materials is illustrated in Table 1, reflecting differences in capture conditions and segmentation challenges across material types. Despite some materials showing lower 𝐼𝑜𝑈 values, such as the concrete class at 0.767, the overall performance remains high, demonstrating the efficacy of the U-Net model in segmenting complex CDW materials. 0 5 10 15 20 25 30 35 Epoch 0.05 0.10 0.15 0.20 Loss Train Loss Validation Loss Fig. 3 Training and validation loss as a function of epoch during the U-Net model training. These modifications and the extended explanations aim to enhance clarity and provide a deeper insight into the scientific processes and results of the project. Multichannel CNNs for Classification and Localization of CDW 9 0 5 10 15 20 25 30 35 Epoch 0.6 0.7 0.8 0.9 IoU Train IoU Validation IoU Fig. 4 Training and validation 𝐼𝑜𝑈 as a function of epoch during the U-Net model training. 0 5 10 15 20 25 30 35 Epoch 0.70 0.75 0.80 0.85 0.90 0.95 F1 Train F1 Validation F1 Fig. 5 Training and validation 𝐹1 scores as a function of epoch during the U-Net model training. Material Class Mean 𝐼 𝑜𝑈 AAC 0.943 Asphalt 0.865 Ceramics 0.935 Concrete 0.767 Table 1 Mean 𝐼𝑜𝑈 values achieved by the U-Net model on the training dataset for different materials. 16 Tom´ aˇ s Zb´ ıral, V´ aclav Neˇzerka AAC Asphalt Ceramics Concrete Predicted label AACAsphaltCeramicsConcrete True label 90.3% (159) 0.0% (0) 0.0% (0) 9.7% (17) 0.0% (0) 98.1% (106) 0.0% (0) 1.9% (2) 0.5% (1) 0.0% (0) 99.5% (205) 0.0% (0) 5.6% (11) 0.5% (1) 0.0% (0) 93.9% (185) Accuracy: 95.34% Fig. 11 Confusion matrix from epoch 61, detailing specific classification challenges and successes of the ResNet model. of the segmented images, considerably improves the system’s performance. The optimized U-Net model achieved real-time processing speeds, thereby enhancing operational efficiency, while the ResNet model demonstrated superior accuracy in material classification, surpassing previous benchmarks set by conventional models. However, the study is not without limitations. The precision of segmentation and classification still varies with the material type, indicating an area for future improvement. Moreover, the implementation of these models in real-world conditions might present unforeseen challenges, necessitating further on-site validations. Future research should focus on refining these models to handle a broader array of materials and conditions, potentially integrating additional sensory data to enhance the models’ applicability and resilience. Further exploration into transfer learning could also reduce the need for extensive training datasets, speeding up the adaptation of these models to new waste types and environments. The authors have no conflicts of interest to declare that are relevant to the content of this chapter. This work was funded by the European Union under the project ROBOPROX (no. CZ.02.01.01/00/22 008/0004590), by the European Union’s Horizon Europe Framework Programme (call HORIZON-CL4-2021-TWIN-TRANSITION-01-11) under grant agreement No. 101058580, project RECONMATIC (Automated solutions for sustainable and circular construction and demolition waste management), and the Czech Technical University in Multichannel CNNs for Classification and Localization of CDW 17 Prague, grant agreement No. SGS24/003/OHK1/1T/11 (Application of state-of-the-art technologies for enhancing sustainability of the construction sector). 18 Tom´ aˇ s Zb´ ıral, V´ aclav Neˇzerka References [1] T. Joensuu, H. Edelman, A. Saari, Circular economy practices in the built environment, Journal of Cleaner Production 276 (2020) 124215. doi:10. 1016/j.jclepro.2020.124215. [2] B. I. Oluleye, D. W. Chan, A. B. Saka, T. O. Olawumi, Circular economy research on building construction and demolition waste: A review of current trends and future research directions, Journal of Cleaner Production 357 (2022) 131927. doi:10.1016/j.jclepro.2022.131927. [3] L. Zheng, H. Wu, H. Zhang, H. Duan, J. Wang, W. Jiang, B. Dong, G. Liu, J. Zuo, Q. Song, Characterizing the generation and flows of construction and demolition waste in China, Construction and Building Materials 136 (2017) 405–413. doi:10.1016/j.conbuildmat.2017.01.055. [4] R. Hl˚uˇzek, J. Trejbal, V. Neˇzerka, P. Demo, Z. Proˇ sek, P. Tes´ arek, Improvement of bonding between synthetic fibers and a cementitious matrix using recycled concrete powder and plasma treatment: from a single fiber to FRC, European Journal of Environmental and Civil Engineering 26 (2020) 3880–3897. doi: 10.1080/19648189.2020.1824821. [5] Z. Proˇ sek, J. Trejbal, V. Neˇzerka, V. Goli´ aˇ s, M. Faltus, P. Tes´ arek, Recovery of residual anhydrous clinker in finely ground recycled concrete, Resources, Conservation and Recycling 155 (2020) 104640. doi:10.1016/j.resconrec. 2019.104640. [6] J. Valentin, J. Trejbal, V. Neˇzerka, T. Valentov´ a, M. Faltus, Characterization of quarry dusts and industrial by-products as potential substitutes for traditional fillers and their impact on water susceptibility of asphalt concrete, Construction and Building Materials 301 (2021) 124294. doi:10.1016/j.conbuildmat. 2021.124294. [7] V. Neˇzerka, Z. Proˇ sek, J. Trejbal, J. Peˇ sta, J. Ferriz-Papi, P. Tes´ arek, Recycling of fines from waste concrete: Development of lightweight masonry blocks and assessment of their environmental benefits, Journal of Cleaner Production 385 (2023) 135711. doi:10.1016/j.jclepro.2022.135711. [8] J. D. L. H. Hoong, J. Lux, P.-Y. Mahieux, P. Turcry, A. A¨ ıt-Mokhtar, Determination of the composition of recycled aggregates using a deep learningbased image analysis, Automation in Construction 116 (2020) 103204. doi: 10.1016/j.autcon.2020.103204. [9] Y. Su, Multi-agent evolutionary game in the recycling utilization of construction waste, Science of The Total Environment 738 (2020) 139826. doi:10.1016/ j.scitotenv.2020.139826. [10] P. Davis, F. Aziz, M. T. Newaz, W. Sher, L. Simon, The classification of construction waste material using a deep convolutional neural network, Automation in Construction 122 (2021) 103481. doi:10.1016/j.autcon.2020. 103481. [11] S. P. Gundupalli, S. Hait, A. Thakur, A review on automated sorting of sourceseparated municipal solid waste for recycling, Waste Management 60 (2017) 56–74. doi:10.1016/j.wasman.2016.09.015. Multichannel CNNs for Classification and Localization of CDW 19 [12] W. Lu, J. Chen, Computer vision for solid waste sorting: A critical review of academic research, Waste Management 142 (2022) 29–43. doi:10.1016/j. wasman.2022.02.009. [13] J. Yang, Z. Zeng, K. Wang, H. Zou, L. Xie, GarbageNet: A unified learning framework for robust garbage classification, IEEE Transactions on Artificial Intelligence 2 (2021) 372–380. doi:10.1109/tai.2021.3081055. [14] Z. Wang, H. Li, X. Zhang, Construction waste recycling robot for nails and screws: Computer vision technology and neural network approach, Automation in Construction 97 (2019) 220–228. doi:10.1016/j.autcon.2018.11. 009. [15] Z. Wang, H. Li, X. Yang, Vision-based robotic system for on-site construction and demolition waste sorting and recycling, Journal of Building Engineering 32 (2020) 101769. doi:10.1016/j.jobe.2020.101769. [16] Z. Dong, J. Chen, W. Lu, Computer vision to recognize construction waste compositions: A novel boundary-aware transformer (BAT) model, Journal of Environmental Management 305 (2022) 114405. doi:10.1016/j.jenvman. 2021.114405. [17] W. Xiao, J. Yang, H. Fang, J. Zhuang, Y. Ku, Classifying construction and demolition waste by combining spatial and spectral features, Proceedings of the Institution of Civil Engineers—Waste and Resource Management 173 (2020) 79–90. doi:10.1680/jwarm.20.00008. [18] Y. Ku, J. Yang, H. Fang, W. Xiao, J. Zhuang, Deep learning of grasping detection for a robot used in sorting construction and demolition waste, Journal of Material Cycles and Waste Management 23 (2020) 84–95. doi:10.1007/ s10163-020-01098-z. [19] K. Lin, T. Zhou, X. Gao, Z. Li, H. Duan, H. Wu, G. Lu, Y. Zhao, Deep convolutional neural networks for construction and demolition waste classification: VGGNet structures, cyclical learning rate, and knowledge transfer, Journal of Environmental Management 318 (2022) 115501. doi:10.1016/j.jenvman. 2022.115501. [20] V. Neˇzerka, J. Trejbal, T. Zb´ ıral, Dataset of construction and demolition waste images: aerated autoclaved concrete (AAC), asphalt, ceramics, and concrete (Feb. 2023). doi:10.5281/zenodo.7618350. [21] V. Neˇzerka, T. Zb´ ıral, J. Trejbal, Machine-learning-assisted classification of construction and demolition waste fragments using computer vision: Convolution versus extraction of selected features, Expert Systems with Applications 238 (2024) 121568. doi:10.1016/j.eswa.2023.121568. [22] X. Qin, Z. Zhang, C. Huang, M. Dehghan, O. Zaiane, M. Jagersand, U2net: Going deeper with nested u-structure for salient object detection, Pattern Recognition 106 (2020) 107404. [23] X. Qin, H. Dai, X. Hu, D.-P. Fan, L. Shao, L. V. Gool, Highly accurate dichotomous image segmentation, in: Lecture Notes in Computer Science, 2022. doi:10.1007/978-3-031-19797-0_3. [24] A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Doll´ ar, R. Girshick, Segment anything, 20 Tom´ aˇ s Zb´ ıral, V´ aclav Neˇzerka in: 2023 IEEE/CVF International Conference on Computer Vision (ICCV), 2023. doi:10.1109/iccv51070.2023.00371. [25] A. Buslaev, V. I. Iglovikov, E. Khvedchenya, A. Parinov, M. Druzhinin, A. A. Kalinin, Albumentations: Fast and flexible image augmentations, Information 11 (2020). doi:10.3390/info11020125. [26] A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al., Pytorch: An imperative style, highperformance deep learning library, Advances in neural information processing systems 32 (2019). [27] O. Ronneberger, P. Fischer, T. Brox, U-net: Convolutional networks for biomedical image segmentation, in: Medical image computing and computerassisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18, 2015, pp. 234–241. doi:10.1007/978-3-319-24574-4_28. [28] K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778. doi:10.1109/cvpr.2016.90.