scieee AI-readable full text Open interactive document viewer

Grapevine reconstruction and pruning points identification based on Deep Learning

Wang, Yiyi

Abstract

Despite the growing prevalence of robotics in agriculture, there is still a limited amount of research specifically targeting the automation of grapevine management. Due to the complexity, the pruning task during the dormant season demands skilled laborers, who are increasingly scarce during the winter months. For the potential prospect, CANOPIES, a H2020 European Project, seeks to pioneer a new collaborative approach between humans and robots in precision agriculture, specifically targeting permanent crops such as table-grape vineyards. Its goal is to enable farm workers to seamlessly collaborate with robot teams to carry out harvesting and pruning more efficiently and effectively. Aiming to fulfill the need for winter pruning task for the CANOPIES project, this master’s thesis presents a novel design of perception in the pruning procedure. The algorithm is mainly developed based on deep learning, providing selective solutions for different application scenarios for the tasks such as plant reconstruction, vineyards organ segmentation and potential pruning point identification. The developed system highly satisfies the need of CANOPIES project for winter pruning. In plant reconstruction section, several models with satisfactory performance are obtained whose best mIoU is 88.72%, and shortest computational cost is 0.1124 seconds. While in the part of vine organ segmentation, the best network gets mPA50 of 60.43% also with a high speed. Finally, a node graph is generated to show the topological structure of the relationship between different organs along with pruning point localization.

Full text

DRAFT Master’s Thesis Master Degree in Automatic Control and Robotics Grapevine Reconstruction and Pruning Points Identification Based on Deep Learning April 30, 2024 Autor: Yiyi Wang Director: Antoni Grau Escola Tècnica Superior d’Enginyeria Industrial de Barcelona page. 1 Abstract Despite the growing prevalence of robotics in agriculture, there is still a limited amount of research specifically targeting the automation of grapevine management. Due to the complexity, the pruning task during the dormant season demands skilled laborers, who are increasingly scarce during the winter months. For the potential prospect, CANOPIES, a H2020 European Project, seeks to pioneer a new collaborative approach between humans and robots in precision agriculture, specifically targeting permanent crops such as table-grape vineyards. Its goal is to enable farm workers to seamlessly collaborate with robot teams to carry out harvesting and pruning more efficiently and effectively. Aiming to fulfill the need for winter pruning task for the CANOPIES project, this master’s thesis presents a novel design of perception in the pruning procedure. The algorithm is mainly developed based on deep learning, providing selective solutions for different application scenarios for the tasks such as plant reconstruction, vineyards organ segmentation and potential pruning point identification. The developed system highly satisfies the need of CANOPIES project for winter pruning. In plant reconstruction section, several models with satisfactory performance are obtained whose best mIoU is 88.72%, and shortest computational cost is 0.1124 seconds. While in the part of vine organ segmentation, the best network gets mPA50 of 60.43% also with a high speed. Finally, a node graph is generated to show the topological structure of the relationship between different organs along with pruning point localization. Keywords: Robotic pruning, vineyard automation, deep learning page. 2 page. 3 Contents 1 Introduction ............................................................. 7 1.1 Motivation............................................................. 7 1.2 Objectives............................................................. 8 1.3 ThesisLayout.......................................................... 8 2 LiteratureReview......................................................... 11 2.1 Robotic Systems Applied in Agriculture Field . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 2.2 StudyFocusingonVineRruning........................................... 11 3 FieldEnvironmentandMaterial............................................ 15 3.1 FieldEnvironment...................................................... 15 3.2 Robotandsensors....................................................... 15 3.3 DataCollectionandStructure ............................................. 16 4 SystemOverview......................................................... 19 5 PlantReconstruction...................................................... 23 5.1 AlignmentofdepthandRGBimages........................................ 23 5.2 PlantSegmentation...................................................... 25 5.2.1 Method1: Lightness Segmentation . . . . . . . . . . . . . . . . . . . . . . . 25 5.2.2 Method2: Semantic segmentation of deep learning . . . . . . . . . . . . . 27 5.3 PointCloudGeneration.................................................. 33 6 VineOrganSegmentation................................................. 37 6.1 ModelSelection ........................................................ 37 6.2 DataAnnotation........................................................ 38 6.3 DataAugmentation ..................................................... 39 6.4 ExperimentandResults.................................................. 40 7 PruningPointIdentification ............................................... 45 7.1 NodeGraph ........................................................... 45 7.2 PruningPointLocalization................................................ 45 page. 4 8 Conclusion............................................................... 49 9 FutureWork ............................................................. 51 Acknowledgements....................................................... 53 page. 5 List of Figures 1 LogoofProjectCANOPIES ............................... 7 2 Vineyardfield ....................................... 15 3 Robotandsensor ..................................... 16 4 flatshots .......................................... 17 5 upwardshots ....................................... 17 6 RGBimagesfromdatasets................................ 17 7 Pipelineofthesystem .................................. 19 8 upward shots by lightness segmentation . . . . . . . . . . . . . . . . . . . . . . . . 26 9 flat shots by lightness segmentation . . . . . . . . . . . . . . . . . . . . . . . . . . 26 10 result of lightness segmentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26 11 Plant segmentation annotation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 12 Fast-SCNNstructure ................................... 28 13 UNetstructure....................................... 29 14 KNetkernelupdatehead................................. 30 15 rawRGBimage ...................................... 32 16 Fast-SCNNsegmentation ................................ 32 17 U-Netsegmentation.................................... 32 18 K-Netsegmentation.................................... 32 19 Segmentationresults ................................... 32 20 Deployment ........................................ 32 21 rawRGBimage ...................................... 34 22 Segmentationresults ................................... 34 23 Filtered3Dmodel..................................... 34 24 Pointcloudgeneration .................................. 34 25 Structure of Yolov8 multi-task . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38 26 annotationforYOLOv8.................................. 39 27 annotation for YOLOv8 multi-task . . . . . . . . . . . . . . . . . . . . . . . . . . . 39 28 dataaugmentationofYOLO............................... 40 29 criteria of YOLOv8 during training . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 30 outputofYOLOv8 .................................... 41 31 output of YOLOv8 multi-task . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42 32 organsegmentation.................................... 47 33 node graph and pruning point of flat shots . . . . . . . . . . . . . . . . . . . . . . 47 34 organsegmentation.................................... 48 35 node graph and pruning point of upward shots . . . . . . . . . . . . . . . . . . . . 48 page. 7 1 Introduction 1.1 Motivation This profound interest in human-robot collaboration has given rise to a dedicated field of study, with numerous universities and corporations now actively exploring this area. One notable initiative is the European Project Canopies 2020[38], which zeroes in on human-robot collaboration within the strenuous realm of agriculture. This project marks a pioneering effort to bring collaborative robotics into table grape vineyard operations, specifically addressing harvesting and pruning challenges. The challenges in these agricultural settings are manifold. The dynamic nature of outdoor lighting and field conditions complicates perception. Manipulation tasks are complicated by the fragility of the fruits and plants. Additionally, the complexity and density of the plant structure highly effect both navigation and perception. Figure 1: Logo of Project CANOPIES This thesis endeavors to take incremental steps towards addressing these challenges in a scenario that demands attention. Generally, it provides novel perception system of winter grapevines which helps pruning procedure to be highly automatic. Before diving into the deep, a general introduction of the field environment and materials is addressed. Then, the thesis are divided into three main parts: plant reconstruction, vine organ segmentation and pruning point identification. In order to prune effectively with robots, a precise 3D model of the vine becomes essential for identifying pruning locations and planning the motion of the robot. However, dense and sprawling canes and complicated background present challenges in accurate modeling. With several high precision deep neural network models, the plant is able to be extracted from color images. Then, the alignment of color and depth images is obtained by a specific optical algorithm. With the output masks from NN models and the aligned depth images, 3D plant models page. 15 3 Field Environment and Material 3.1 Field Environment The vineyard used for this study was located at the Corsira Agricultural Cooperative Society (Aprilia, Italy). Aprilia’s vineyard features a unique double-roof structure, with European grapes (Vitis Vinifera) positioned nearly 2 meters above ground and rows spaced 2.5 meters apart, each hosting evenly spaced vines every 2 meters shown in Fig.2. This specific architecture, used exclusively for table grape vines, allows the plants to grow in a controlled 360-degree direction from the main trunk, with branches extending radially up to 1.5 meters. This growth pattern creates a canopy of interlocking branches spanning across rows. The structure’s stability is enhanced by a network of intersecting cables that provide support for the vine branches, which also crisscross from various directions. The presence of weeds, both within and between the rows, and the tangled canopy overhead suggest insufficient vineyard management. Past pruning mistakes have resulted in vines that lack uniform shape, each appearing distinct from the next. Additionally, the soil is notably muddy and uneven during this season, adding to the navigational challenges of the environment. Figure 2: Vineyard field 3.2 Robot and sensors We have used a robot that is an evolution of the Tiago++ robot manufactured by PAL Robotics. This robot was specifically designed for the CANOPIES European project. It is mounted on the Alitrack platform, a caterpillar platform, and it has two robot arms. Each one has 7 degrees of freedom and in each one of them, it is located a Realsense D435i stereo camera and in the head of the Tiago++ there is another Realsense D435i shown in Fig.3(a)). page. 16 The Realsense D435i camera sensor was configured with a resolution of 1280 ×720 pixels, and frames were captured at a rate of 30 frames per second. The RGB camera has a field of view (FOV) of 69 degrees horizontally and 42 degrees vertically. In terms of depth settings, the camera used stereoscopic depth technology, providing a depth field of view of 87 degrees horizontally and 58 degrees vertically. At maximum resolution, the minimum depth distance (Min-Z) is of 28 cm. The depth output resolution was configured to be 1280 ×720 pixels, the same as the RGB camera, and the depth images were captured at a rate of 30 frames per second. The Realsense D435i mounted on the robot wrist can be seen in Fig.3(b). ((a)) CANOPIES Tiago++ bi-manipulator ((b)) gripper with the Realsense D435i camera on the wrist Figure 3: Robot and sensor 3.3 Data Collection and Structure There are two main sets of datasets used in this project. One was recorded in January 2022, and another was recorded in April 2023. The former one is for flat shots(see Fig.4) and the later one is for upward shots(see Fig.5). Both of them were collected at different hours to cover illuminations and weather conditions. In addition, images were obtained at different distances and from different angles. These two sets of datasets has the purpose of designing classifiers for vineyard grape cluster detection and tree branches detection. These data was captured in real time and real scenarios and it used the following sensors: RGB-D images, LiDAR depth data, GPS RTK. And, data collected was in the form of ROS bag data. In ROS bag, the topics are listed below: •/camera/accel/imu_info: Information of IMU •/camera/color/camera_info: Intrinsic camera info, including K,R and P matrix •/camera/color/camera_raw: Raw RGB image page. 17 •/camera/depth/color/points: 3D point cloud •/camera/depth/color/img_rect_raw: Raw depth image (not aligned to RGB images) •/camera/extrinsics/depth_to_color: Extrinsic camera info which implies the relationship between RGB and depth images Figure 4: flat shots Figure 5: upward shots Figure 6: RGB images from datasets page. 19 4 System Overview Pruning is a complex task, involving a comprehensive assessment of each tree to determine the optimal branches for removal, thus guiding the tree’s growth for the upcoming year. Typically, this process involves a holistic view of the tree, identifying branches for pruning based on factors like their age or condition. Expertise plays a crucial role, with experienced pruners relying on established guidelines tailored to individual plant conditions. This meticulous process requires precise identification of branch types, understanding of plant structure, and accurate detection of nodes and their positions along branches. Automating vine pruning presents significant challenges, both for robots and humans, due to the complexity involved in replicating these intricate decision-making processes. To tackle the problem, a complete system is developed, whose general process flow can be seen in Fig.7. There are three main tasks: the first is reconstruct the plant into a 3D model. Next, it is necessary to have a segmentation map at a pixel level of each plant organ. Finally, an algorithm to identify the pruning points is developed. Figure 7: Pipeline of the system The main blocks of the system are the following: •Stereo camera: The stereo camera releases RGB and depth images separately. •Depth image alignment: Align the depth images with RGB ones. •Semantic segmentation: it is the one of cores of the system, which uses RGB images as page. 20 input, and processes them to segment plant and leaves. •3D model: Combine output masks from semantic segmentation with RGB and depth images to generate 3D point cloud to show the plant structure. •Instance segmentation: It is another critical part of the pipeline, which detects the buds, joint and segments different type of canes at a pixel level. •Node graph: With the segmented plant organ masks, an algorithm is developed to generate the node graph, which shows the topology of each selected cane. •Potential pruning point identification: Together with node graph, the localization ((u, v) coordinates) of pruning points are provided according to the relationship between each organ. •Pruning position: Localization in (u, v, Z)coordinates of points will be transformed into camera frame (X, Y, Z). The methods of the main blocks will be explained deeply in the next sections. page. 21 page. 23 5 Plant Reconstruction Ensuring the robot accurately locates the plant during pruning is paramount. A robust 3D reconstruction serves two crucial purposes: firstly, recognizing the plant as an obstacle to prevent the manipulator from colliding with the vine, and secondly, precisely identifying the cutting points. To reach the goal, this section proposes several solutions to have a precise reconstruction to fit different working scenarios. 5.1 Alignment of depth and RGB images Before starting to extracting the plant, an alignment of depth and RGB images is necessary. As I mentioned in the third chapter, depth and RGB are recorded in different frames, which will affect overlapping the segmentation results and point cloud registration later. Since there are point cloud generated along with raw RGB image while collecting data, transforming point cloud to aligned RGB and depth images will be executable. Given the camera intrinsics matrix (X, Y, Z)in the camera coordinate system, the projection of the point onto the image plane results in a 2D point (u, v)in pixel coordinates. This projection can be calculated using the following formula: u=fx·X Z+cx v=fy·Y Z+cy where fx, fyare the focal lengths along the x and y axes, respectively. And cx, cyare the principal point coordinates. After running the algorithm1, alignment is obtained. However, the numerical values of the depth map obtained by this method are not real, but the relative ratio between depth values. So, this works fine for reconstructing the model, but it’s not enough for pruning point positioning. So a further step is targeting on align depth to RGB images, and keep the real depth value at the same time. Another algorithm2 is used to perfectly get the goal which uses the features of camera intrinsics and extrinsics. First, the depth image is loaded along with camera intrinsics and extrinsics. The algorithm constructs a 3D point cloud from the depth image using pixel-topoint conversion. Then, the points are transformed from the depth camera coordinate system to the RGB camera coordinate system using rotation Rand translation T.Finally, the points are projected back to pixel coordinates in the RGB image using point-to-pixel conversion. page. 30 Figure 14: KNet kernel update head Training process Training process was done using MMSegmentation[5], an open source semantic segmentation toolbox based on PyTorch. Its unified benchmark toolbox for semantic segmentation offers a modular design, allowing users to construct customized frameworks by combining different components effortlessly. Moreover, the toolbox prioritizes high efficiency, boasting training speeds that are faster than or on par with other codebases. This comprehensive solution empowers users to experiment and deploy state-of-the-art semantic segmentation methods efficiently and effectively. Before starting training, a specific class is defines which involves 3 classes: background, tree and leaves. During training, background is not removed because the these pixels also contain a lot of information, such as human body, trellis wire, gripper of robot and other objects. Fast-SCNN and U-Net were both trained for 40000 iterations, and K-net was done for 20000 iterations. All of them using Adam as optimizer. Results and Analysis The evaluation criteria and costing time of each model is shown in table.1. 1. Fast-SCNN: •Performance Metrics: Fast-SCNN achieves an mIoU of 81.06%, and the mean accuracy of 89.29% ,indicating its ability to accurately segment objects in the images. •Time Efficiency: Fast-SCNN demonstrates remarkable efficiency with a low processing time.This suggests that Fast-SCNN can process images quickly while maintaining good segmentation performance. 2. U-Net: •Performance Metrics: U-Net achieves an mIoU of 79.57%, indicating slightly lower performance compared to Fast-SCNN in terms of segmentation accuracy. However, the other metrics also show competitive performance. •Time Efficiency: U-Net exhibits higher processing time compared to Fast-SCNN. 3. K-Net: page. 31 •Performance Metrics: K-Net demonstrates the highest IoU among the models, with a value of 88.72%. This indicates superior segmentation accuracy compared to both FastSCNN and U-Net. The other metrics also show excellent performance, highlighting KNet’s effectiveness in segmentation tasks. •Time Efficiency: Although K-Net delivers exceptional segmentation accuracy, its computational time is almost 10 times longer than the one of Fast-SCNN. Overall, Fast-SCNN offers a good balance between performance and efficiency, making it suitable for real-time applications. U-Net provides competitive performance with moderate computational cost, while K-Net delivers superior segmentation accuracy at the expense of slightly higher processing time. Table 1: Comparison of Model Performance and Time Model Metrics Time (s) mIoU mAcc mDice mF-score mPrecision mRecall Fast-SCNN 81.06 89.29 89.20 89.20 89.11 89.29 0.1124 U-Net 79.57 87.15 88.19 88.19 89.33 87.15 0.6033 K-Net 88.72 95.03 93.78 93.78 92.61 95.03 0.9293 Although the above evaluation criteria is quite good, images for validation are only four. It is necessary to take a look on other raw images which are not included in this segmentation task. The result is shown in Fig.19. For the first upward shot, only K-Net perfectly extract the plant without involving in human body. And in the next image set which background is more complex, Fast-SCNN misidentified the computer as a tree. And the plants extracted are not as detailed as the other two models. In the third set, U-Net doesn’t segment the leaves as expectation and only thick canes are detected. In summary, K-Net is the most accurate model among all of them, but its memory and computational cost is too big to be used in real-time tasks. So, segmentation tasks for database maximizes the advantages of this model. While for the performance of Fast-SCNN, maybe some small barbs will be missing and some misidentification will happen. Due to its low computational cost and time efficiency, this model can be utilized in the field for real-time tasks. Model Deployment In order to perform tasks in real-world applications where resources and hardware are not defined, model deployment is followed after training models. Here, the pytorch models are converted to ONNX (Open Neural Network Exchange) format, which provides an interoperable framework for exchanging models between different deep learning framework. In addition, ONNX models can be deployed across different platforms and devices, including mobile devices, edge devices, and cloud servers. By converting a PyTorch model to ONNX, there is more flexibility in real world applications. By using MMDeploy[4], an open-source deep learning model deployment toolset, fine-tuned Fast-SCNN and U-Net are converted. After that, the weight size shrinked to half original size and segmentation accuracy still keeps a high level as shown in Fig.20 Finally, ONNX models are be optimized for inference performance using tools ONNX Runtime[6], which provides efficient execution of ONNX models across different hardware platforms. page. 32 Figure 15: raw RGB image Figure 16: Fast-SCNN segmentation Figure 17: U-Net segmentation Figure 18: K-Net segmentation Figure 19: Segmentation results ((a)) raw image ((b)) output of pytorch ((c)) output of onnx Figure 20: Deployment page. 33 5.3 Point Cloud Generation After segmentation, a filtered 3D point cloud is generated following these process: Overlap segmented mask separately with RGB and depth image. Then morphology process is applied to refine the segmented regions and improve segmentation accuracy. Then, convert to 3D model using Open3D point cloud generator with RGB colors assigned from the RGB image and depth values retained from the processed depth image. Finally, statistical outlier removal is applied to the generated point cloud to filter out noise and outliers, resulting in a clean and accurate 3D point cloud representation. Details are shown in Algorithm.3 Algorithm 3 Point Cloud Generation Input: RGB image, aligned depth image, segmented mask Output: filtered 3D point cloud Step 1: Overlap segmented mask on RGB image Create a blank image same size as RGB image as overlapped RGB image. For each pixel (x, y)in the segmented mask: If mask value is 1: Set corresponding pixel in overlapped RGB image to [0, 200, 0] If mask value is 2: Set corresponding pixel in overlapped RGB image to original color End If End For Step 2: Overlap segmented mask on depth image Create a blank image same size as RGB image as overlapped RGB image. For each pixel (x, y)in the segmented mask: If mask value not 0: Set corresponding pixel in overlapped depth image to depth value End If End For Step 3: Apply dilation and erosion to overlapped depth image Apply dilation and erosion operations with a (3,3) kernel to the overlapped depth image Step 4: Convert overlapped RGB-D images to 3D point cloud For each pixel (x, y)in the overlapped RGB and depth images: If depth value is valid: pcd = o3d.geometry.PointCloud.create_from_rgbd_image(overlapped RGB, overlapped depth) End If End For Step 5: Apply statistical outlier removal to point cloud cl, ind =pcd.remove_statistical_outlier(nb_neighbors = 20, std_ratio = 2.0) filteredpcd =pcd.select_by_index(ind) Output the filtered 3D point cloud The results are shown in Fig.24. The upward shot (left column) whose structure is relatively simple and the distance is small provides a highly accurate 3D model. While the flat shot (right column), some canes are disconnected and even not shown. Moreover, the color of point cloud is not exactly the same as the RGB one. Since the segmentation is quite perfect, the reason for these flawless is the depth image. During data collection, depth and RGB sensors may not page. 34 always be perfectly synchronized, leading to occasional mismatches in the number of captured frames. Obviously, the longer the duration of filming, the more mismatches will happen. Figure 21: raw RGB image Figure 22: Segmentation results Figure 23: Filtered 3D model Figure 24: Point cloud generation page. 35 page. 37 6 Vine Organ Segmentation Vine organ segmentation enables precise identification and delineation of different parts of the vine, whose data can inform decisions about resource allocation during pruning, such as determining the number of buds to retain for the upcoming growing season. By optimizing bud density and distribution, growers can ensure efficient resource utilization. As Fig.2 shows, vine plants mainly consists of trunks, branches and canes. To segment the canes, and tell the age of them is the most critical thing for all kinds of pruning rules. In addition, the buds on canes also play an important role when localizing the pruning points. So in this section, two state of the art neural networks are introduced and fine-tuned to segment 1st-year canes and 2nd-year canes. Meanwhile, detect the joints of 1st-year canes and all the buds. 6.1 Model Selection •Yolov8: YOLOv8[16], short for "You Only Look Once version 8," is a state-of-the-art algorithm that belongs to the YOLO family of models. YOLOv8 builds upon previous versions of YOLO, covering a full range of vision AI tasks. It leverages advanced backbone and neck architectures, enhancing feature extraction and object detection performance. By adopting an anchor-free split Ultralytics head, it achieves superior accuracy and efficiency compared to anchor-based methods. With a deliberate emphasis on optimizing the accuracy-speed tradeoff, YOLOv8 strikes a balance suitable for real-time object detection across various application domains. Although YOLOv8 is a powerful network, only one kind of task can be done in one time by using the existing models. In order to detect the buds and joints, the detection task is converted to segmentation task as well. For comparison, another multi-task model is introduced below. •Yolov8 multi-task: YOLOv8 Multi-Task[36] extends the capabilities of YOLOv8 by introducing multi-task learning, where the model simultaneously performs multiple related tasks during training. In addition to object detection, YOLOv8 Multi-Task can handle tasks such as detection, instance segmentation, and depth estimation within a single unified framework. The network is configured into 3 tasks: segmenting 1st-year cane, segmenting 2nd-year cane and detect the buds and joints. The structure is shown in Fig.25. The network has a backbone that processes the input image, followed by separate "necks" for segmentation and detection tasks. This suggests that the architecture works for simultaneous object detection and pixel-wise segmentation at multiple scales, which is fulfuill the need of this task that requires both context and detail.In addition, by sharing the backbone and using efficient upsampling methods, the network can perform both tasks without redundant computation. What’s more, multi-scale processing often leads to higher accuracy, as the network can detect and segment objects of various sizes and complexities. page. 38 Figure 25: Structure of Yolov8 multi-task 6.2 Data Annotation The same as the previous section, data annotation is done by using software X-AnyLabeling. There are total 44 images in the dataset which are labeled in 4 classes: 1st-year cane, 2nd-year cane, bud and joint. And 40 images are for training and another 5 for validation. The 1st-years canes are relatively thinner and in lighter color, while the 2nd-year ones are thicker and connected to branches. Besides, the joints, where the 1st-years canes are connected with other parts of plant are also labeled. The buds are structures present on the canes where new shoots may grow. Dividing the grapevine into these four main categories allows us to generate potential pruning points. However, due to the different needs of two networks, the same dataset must be labeled in two ways. •YOLOv8: Since it’s a totally instance segmentation task, all the instances are labeled as polygons in one json file to make format transformation easier. An example of these annotation concepts can be seen in Fig.26, where red parts stands for 1st-year canes, and green ones are 2nd-year canes. Buds and joints are respectively labeled in blue and yellow blocks. Then, these json files are all exported as YOLO txt using the embedded function in annotator software. page. 39 Figure 26: annotation for YOLOv8 •YOLOv8 multi-task: To launch three tasks simultaneously, the images should be annotated targeting the objects which belong to each task as shown in Fig.27 ((a)) buds and joints detection ((b)) 1st-year cane seg ((c)) 2nd-year cane segn Figure 27: annotation for YOLOv8 multi-task 6.3 Data Augmentation There are several data augmentation techniques embedded in ultralytics system, such as Mosaic Augmentation which blends four training images into one, enhancing object detection models’ ability to handle diverse object scales and translations. And Random Affine Transformations introduce variability through random rotation, scaling, translation, and shearing. MixUp Augmentation creates composite images by linearly combining two images and their labels. Albumentations, a robust library, offers various augmentation techniques. HSV Augmentation introduces randomness by altering the Hue, Saturation, and Value of images, while Random Horizontal Flip randomly flips images horizontally. All the techniques mentioned are integrated into different augment strategies which will be applied in different stage of training. For example, Mosaic Augmentation is only utilized during first 90% epochs of training as Fig.28 shows. page. 46 Algorithm 4 Node Graph Generation Input: Prediction results and raw RGB image Output: Node graph for each result in prediction results do Extract bounding boxes and segmentation masks Initialize graph G for each polygon in 1st-year cane masks do Convert polygon points to mask Initialize sub node graph for current loop and node lists for buds and joints for each box in joint boxes do Calculate joint region and check intersection with cane mask if intersection exists then Add joint center(cXjoint, cYjoint)to joint nodes and sub node graph (int(cXjoint), int(cYjoint)) ←jointnodes sub_node_graph[int(cYjoint), int(cXjoint)] ←255 end if end for for each box in bud boxes do Check if bud center (cXbud, cYbud)is inside polygon if bud is inside then Add bud center to bud nodes and sub node graph (int(cXbud), int(cYbud)) ←budnodes sub_node_graph[int(cYbud), int(cXbud)] ← 255 end if end for for (Y, X)insub_node_graph do if sub_node_graph(Y, X)! = 0 then G.add_node((X, Y )) end if end for Sort bud nodes by distance to nearest joint node Connect nodes in sorted bud nodes list Connect nearest bud node to joint node end for Draw nodes and edges in complete node graph G on original image end for page. 47 ((a)) raw RGB image ((b)) organ segmentation Figure 32: organ segmentation ((a)) node graph ((b)) potential pruning point Figure 33: node graph and pruning point of flat shots page. 48 As we can see in Fig.35, the potential pruning points are drawn in green. To prevent to cut at an invalid location, some points are hidden. In conclusion, this node graph enables pruning identification easier and more efficient. But the result is highly depends on the vine organ segmentation, which needs to be improved in the future. ((a)) raw RGB image ((b)) organ segmentation Figure 34: organ segmentation ((a)) node graph ((b)) potential pruning point Figure 35: node graph and pruning point of upward shots page. 49 8 Conclusion In summary, this thesis introduces a new system for autonomous winter vine pruning, which generally satisfy the need of CANOPIES project. Before starting the project, the field environment was analyzed in order to dress a more suitable solution to the real world application. And all the data are collected in real field under different occasions, promising the reliability of the results. With 2 kinds of dataset shot from different angles, the software becomes more robust and stable. However, the images are not aligned in the same frame, affecting system performance at the beginning. Thanks to the different structure of 2 datasets, camera intrinsics and extrinsics are both utilized to develop an optical solution to align RGB and depth images. Although this process is easy to be ignored, but it indeed reduces errors for all the following perception process. Moreover, this system provides several fine-tuned network to extract and reconstruct the plants fitting different working environment. These three segmentation models are classical in the deep learning field, and each one represents the segmentation level in a specific period. High speed model FastSCNN, is perfect for real-time task in vineyards. Even for devices with low memory GPU or solo CPU, this model still gives a fast and robust segmentation. While the most accurate one, K-Net, provides the possibility to generate a database for research in the lab. No only in autonomous pruning, it will be useful also for the researches about human robot interaction. Finally, a precise 3D model is obtained. This outcome may have possibility to be extended to VR environment modeling in the future. Due to various factors such as limited data on grapevines of different ages and varying capture conditions, the accuracy of detected grapevine items used for generating potential pruning points can be compromised. To address these challenges, it is necessary to implement data augmentation to enrich the dataset and to capture new images across diverse grapevine conditions. By observing and exploring the auto-data augmentation embedded in YOLO during training process, my training strategies are also modified and improved. Because of the potential inaccuracies in segmentation, two new released networks are trained and evaluated. Both of them are ’SOTA’ model, which have a strong competitiveness compared to others. Although the accuracy is not as good as expectation, the speed and computational cost are quite satisfying. Then, a vine organ segmentation map is created based on instance segmentation of grapevines. Following the organ segmentation, which encapsulates topographical and geometrical information of various grapevine organs. A methodology successfully generates a substantial set of potential pruning points. Analyzing the the relationship between different kinds of organs, I proposed an algorithm to build the node graph, to show the topological structure of each selected cane. From this set, the actual pruning points can ultimately be chosen. Additionally, the algorithm eliminate the candidates which might lead to erroneous location because of curvature. Nonetheless, this preliminary solution enables the robot to perform autonomous pruning of grapevines. page. 51 9 Future Work Although the system fulfills the need and performs quite well in some sections, there is still a lot of flawless when it is really operated. For example, the integration of models, high computational cost for accuracy and the error of depth image itself. To perfect the system, future work should be done in following aspects: 1. Enhanced dataset: Compared with different models, actually it is the dataset that affects the performance the most. In this project, only 40-50 images are labeled, which is a really small scale. Despite the help of data augmentation, the features are still limited. So, including depth information into annotation will enhance the diversity and geometric information. Besides, including data across different seasons and growth stages will also enhance the robustness and accuracy of our segmentation models. 2. Advanced Segmentation Algorithms: In recent years, model updates have become faster and faster. Multi-modal and multi-task networks will become the mainstream. At the same time, the integration of large models also enhances the model performance. For segment the plant, trying to find a speed and accuracy balanced model will be the possible solution. 3. 3D Modeling skeletonlization: Observing the flat shots in this project, the canes are more or less curtained by leaves. But for the density of plant and low resolution of images, to skeletonlize the model is really challenging. However, its advantage is also worthy. With skeletonlized 3D model, the amount of information contained in pictures has risen to another dimension, which enables pruning point identification easier. And it helps connect the missing points to generate a cane-solo point cloud. 4. 2 stages model for instance segmentation: For vine organ segmentation part, two networks both belong to YOLO family, which is 1 stage model. They detect and segment at the same time, winning the time, but sacrificing the accuracy. Since the performance is not so satisfying, testing 2 stages models such as Mask-RCNN or RTMDet and making comparison will be very interesting. 5. Autonomous Robot Navigation: Another avenue of future work will be the enhancement of autonomous navigation algorithms to better handle the complex terrain of vineyards. This includes improving the robot’s ability to navigate between rows and around obstacles, thereby ensuring efficient and safe pruning operations. 6. Application in real field: For now, the system is only a state of art design. When it is actually put into use, more practical problems will emerge, such as hardware mismatch and resource shortage. Some environmental factors that are different from those in the laboratory will greatly disturb the performance of the model. page. 53 Acknowledgements Firstly, I want to thank professor Edmundo Guerra, Antoni Grau and Alberto Sanfeliu for their helping during the project. Then I’m really appreciate being access to CANOPIES project and IRI. I also want to say thanks to all my friends in Barcelona. With your accompany, I never feel lonely in this city. Finally I’d like to thank my family and my partner for supporting me unconditionally. I love you all! page. 54 page. 55 Bibliografia [1] T. Botterill et al. “A Robot System for Pruning Grape Vines”. In: Journal of Field Robotics 34.6 (2017), pp. 1100–1122. url:https://doi.org/10.20870/oeno-one.2019.53.2.2416. [2] Attila Budai et al. “Robust Vessel Segmentation in Fundus Images”. In: International Journal of Biomedical Imaging 9 (2013). [3] Z. Cai and N. Vasconcelos. “Cascade R-CNN: High quality object detection and instance segmentation”. In: IEEE Trans. Pattern Anal. Mach. Intell 43 (2021), pp. 1483–1498. [4] MMDeploy Contributors. OpenMMLab’s Model Deployment Toolbox. https : / / github . com/open-mmlab/mmdeploy. 2021. [5] MMSegmentation Contributors. MMSegmentation: OpenMMLab Semantic Segmentation Toolbox and Benchmark.https://github.com/open-mmlab/mmsegmentation. 2020. [6] ONNX Runtime developers. ONNX Runtime.https://onnxruntime.ai/. Version: x.y.z. 2021. [7] M.A. Ebrahimi et al. “Vision-based pest detection based on SVM classification method”. In: Computers and Electronics in Agriculture 137 (2017), pp. 52–58. [8] Miguel Fernandes et al. “Grapevine Winter Pruning Automation: On Potential Pruning Points Detection through 2D Plant Modeling using Grapevine Segmentation”. In: 2021 IEEE 11th Annual International Conference on CYBER Technology in Automation, Control, and Intelligent Systems (CYBER) (2021). url:https://doi.org/10.48550/arXiv.2106. 04208. [9] Marco Giacchetti. “Perception Tools for Collaborative and Autonomous Pruning for BiManipulator Robot in Table Grape Vineyards”. In: (2023). [10] A. Gollakota and M. Srinivas. “Agribot — A multipurpose agricultural robot”. In: Annual IEEE India Conference (2011), pp. 1–4. [11] S. Gongal A.and Amatya, Q. Karkee M.and Zhang, and K. Lewis. “Sensors and Systems for Fruit Detection and Localization: A Review”. In: Computers and Electronics in Agriculture 116 (2015), pp. 8–19. [12] J. Grimm et al. “An Adaptive Approach for Automated Grapevine Phenotyping using VGG-based Convolutional Neural Networks”. In: (2018). url:http://arxiv.org/abs/ 1811.09561. [13] P. Guadagna, M. Fernandes, and F. et al. Chen. “Using deep learning for pruning region detection and plant organ segmentation in dormant spur-pruned grapevines”. In: Precision Agric 24 (2023), pp. 1547–1569. url:https://doi.org/10.1007/s1111902310006-y. [14] K. He et al. “Mask R-CNN”. In: IEEE Trans. Pattern Anal. Mach. Intell 14.4298 (2017), pp. 386–397. [15] Changho Hwang et al. Tutel: Adaptive Mixture-of-Experts at Scale. 2022. arXiv: 2206.03382. [16] Glenn Jocher, Ayush Chaurasia, and Jing Qiu. Ultralytics YOLO. Version 8.0.0. Jan. 2023. url:https://github.com/ultralytics/ultralytics. [17] W. Kazmi et al. “Indoor and outdoor depth imaging of leaves with time-of-flight and stereo vision sensors: Analysis and comparison”. In: ISPRS journal of photogrammetry and remote sensing 99 (2014), pp. 128–146. url:http://arxiv.org/abs/1811.09561. [18] A Kicherer et al. “Automatic image-based determination of pruning mass as a determinant for yield potential in grapevine management and breeding”. In: Australian journal of grape and wine research 29(3) (1975), pp. 286–291. url:https://doi.org/10.3390/ rs14184495. [19] Alexander Kirillov et al. “Segment Anything”. In: arXiv:2304.02643 (2023).