Deep learning and computer vision for assessing the number of actual berries in commercial vineyards
Abstract
Fernando Palacios would like to acknowledge the research founding FPI grant 286/2017 by Universidad de La Rioja, Gobierno de La Rioja. This work is supported by National Funds by FCT – Portuguese Foundation for Science and Technology, under the project UIDB/04033/2020.
Full text
Research Paper Deep learning and computer vision for assessing the number of actual berries in commercial vineyards Fernando Palacios a,b , Pedro Melo-Pinto c,d , Maria P. Diago a,b , Javier Tardaguila a,b,* a Televitis Research Group, University of La Rioja, 26006, Logro~ no, Spain b Instituto de Ciencias de La Vid y Del Vino (University of La Rioja, CSIC, Gobierno de La Rioja) 26007, Logro~ no, Spain c CITABdCentre for the Research and Technology of Agro-Environmental and Biological Sciences, Universidade de Tr as-os-Montes e Alto Douro, 5000-801, Vila Real, Portugal d Departamento de Engenharias, Escola de Ci^ encias e Tecnologia, Universidade de Tr as-os-Montes e Alto Douro, 5000-801, Vila Real, Portugal article info Article history: Received 8 May 2021 Received in revised form 8 April 2022 Accepted 12 April 2022 Published online 26 April 2022 Keywords: Grapevine yield components Non-invasive sensing technologies Precision viticulture SegNet architecture The number of berries is one of the most relevant yield components that drives grape production in viticulture. The goal of this work was to estimate the number of actual berries per grapevine using computer vision and deep learning in commercial vineyards. Images from the visible range (RGB) were acquired from a set of 96 grapevines (Vitis vinifera L.) at pea-size berry stage using a red, green and blue camera (RGB). At harvest, the number of berries and per vine was manually assessed as the ground-truth values. The algorithm involved computer vision to detect berries in the images and to extract canopy features, in order to gain information about canopy occlusion. These were used by the machine learning regression models built to estimate the number of actual berries per vine. A SegNet architecture was used to segment individual berries and several canopy related features. Four datasets were created combining the number of estimated visible berries and different canopy features. Three different regression models were tested on the four datasets. The best results were achieved with support vector regression (SVR) on a dataset including six canopy features. This method yielded a root mean squared error (RMSE) of 205 berries, a normalised root mean squared error (NRMSE) of 24.99% and a coefficient of determination (R 2 ) of 0.83 between the number of estimated and the number of actual berries per vine. The results show that the number of actual berries in grapevines can be assessed with high accuracy up to 60 days prior to grape harvest, using the developed algorithm based on computer vision and deep learning. ©2022 The Author(s). Published by Elsevier Ltd on behalf of IAgrE. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). *Corresponding author. Televitis Research Group, University of La Rioja, 26006, Logro~ no, Spain. E-mail address: [email protected]s (J. Tardaguila). Available online at www.sciencedirect.com ScienceDirect journal homepage: www.elsevier.com/locate/issn/15375110 biosystems engineering 218 (2022) 175e188 https://doi.org/10.1016/j.biosystemseng.2022.04.015 1537-5110/©2022 The Author(s). Published by Elsevier Ltd on behalf of IAgrE. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
1. Introduction Grapevine yield components'assessment is one of the most relevant tasks in precision viticulture. An accurate and early assessment of these parameters could provide valuable information for grapegrowers to adjust cultural practices, improve fruit quality, and optimise grape harvest. Conventional methods for assessing yield components have classically consisted on a) counting and weighting a set of clusters manually harvested within a group of vines previously selected (Martin, Dunstone, Dunn, 2003), or b) counting the number of inflorescence primordia in dormant winter bud dissection and using historical cluster weights (Wohlfahrt, Collins, Stoll, 2019). These methods are destructive, time-consuming, and often do not represent the spatial variability of commercial vineyards, as typically a low number of vines are sampled. In fact, some authors have further studied different strategies of vineyard sampling to optimise yield components’ assessment in commercial vineyards (Oger, Laurent, Vismara, Tisseyre, 2021;Oger, Vismara, Tisseyre, 2021). In recent decades innovative and non-invasive technologies have provided efficient systems for collecting massive data and images for yield forecast in commercial vineyards. New sensing technologies and artificial intelligence have been recently applied in agriculture. Likewise, emerging technologies can be used for crop prediction (Klompenburg, Kassahun, Catal, 2020) and crop monitoring with several advantages versus conventional methods (Sankaran, Mishra, Ehsani, Davis, 2010). Computer vision is one of the most used technologies in agriculture (Li, Wang, Wang, 2010;Linker, Cohen, Naor, 2012;Rakun, Stajnko, Zazula, 2011). In viticulture it has been used to predict parameters of relevance, such as vine pruning weight (Kicherer et al., 2017), canopy features and coverage (Diago, Aquino, Millan, Palacios, Tardaguila, 2019;Zheng, Panicker, Stott, 2021), disease detection (Liu, Gold, Combs, CadleDavidson, Jiang, 2021;Pai &Thomas, 2021) and grapevine varieties identification (Pereira, Morais, Reis, 2019). For grapevine yield components’ appraisal, computer vision has been applied to address different grapevine organ detection problems (shoots, flowers, berries, clusters …) at specific phenological stages, from budburst (Liu, Cossell, Tang, Dunn, Whitty, 2017) to harvest (Herrero-Huerta, Gonz alez-Aguilera, Rodriguez-Gonzalvez, Hern andez-L opez, 2015;Liu &Whitty, 2015;Liu, Zeng, Whitty, 2020) or even combining multiple phenological stages (Aguiar et al., 2021), both under laboratory and field conditions. In other works, algorithms for flower counting from visible range inflorescence images (Liu et al., 2018;Rahim, Utsumi, Mineno, 2021;Tello, Herzog, Rist, This, Doligez, 2019) acquired using a red, green and blue camera (RGB) have been developed, while other authors focused model development on whole vine images (Grimm et al., 2019;Palacios et al., 2020). Specifically, for berry counting, the works can be divided into two approaches, those focused on cluster images, and those using whole vine images. Within the first group, Grosset^ eteetal.(2012),Aquino, Barrio, Diago, Millan, and Tardaguila (2018) and Coviello, Cristoforetti, Jurman, and Furlanello (2020) developed smartphone applications for berry counting in RGB cluster images using different computer vision methodologies. While the first two developed algorithms based on traditional image processing techniques and machine learning, the third one used a deep learning approach for detecting berries. Similarly, Buayai, Saikaew, and Mao (2021) presented also a deep learning approach to detect the grapevine cluster and its berries on RGB cluster images taken in the field. Then, the authors presented a regression model to estimate the number of actual berries of the cluster using a set of features extracted from the image. The works focused on whole vine images mainly involved the use of a mobile platform for image acquisition. Nuske et al. (2014) and Aquino, Millan, Diago and Tardaguila (2018) utilised a ground moving vehicle to acquire RGB vine images at several phenological stages, and peasize, respectively, that were further used to develop a supervised classifier for detecting berries with the help of berry descriptors extracted from the images. Grimm et al. (2019) and Zabawa et al. (2020) applied a deep learning semantic segmentation approach for detecting and quantifying such as ripe berries, in the first work, and three different phenological stages before harvest, in the second work, on vine images acquired both manually and using a mobile phenotyping platform. P erezZavala, Torres-Torriti, Cheein, and Troni (2018) followed an approachbasedonsupervisedlearningfordetectingbothvisible clusters on images from datasets in different illumination conditions. Instead of visible berry counting, other works have followed a grape pixel counting approach (Dunn &Martin, 2004; Marani, Milella, Petitti, Reina, 2021; I~ niguez et al., 2021). Silver and Monga (2019) tested several slightly different approaches based on deep learning to transform near harvest vine images andgrapepixelsintoayieldestimationbyaddingdenselayersto CNNs to transform the convolutional feature maps outputs. The number of visible berries has been used in some of the previous works as indicator of the final yield at harvest, but it is only a portion of the number of actual berries. In cluster images, the berries can be occluded only by other berries, while in vine images, the berries can be occluded by other elements, such as leaves, shoots, posts, wires and also other berries and clusters. Victorino, Braga, Santos-Victor, and Lopes (2020) evaluated the visibility of the yield components in natural conditions. Most of the previous works have been focused on the detection of the visible berries (Aquino, Barrio et al., 2018; Coviello et al., 2020;Grimm et al., 2019;Grosset^ ete et al., 2012; Zabawa et al., 2020). In Nuske et al. (2014) work, basal leaf removal was performed on the imaged side of the canopy, therefore removing most of the berries’ occlusion by leaves. Aquino, Millan, et al. (2018) applied a full defoliation of the vines and calibrated a linear model for yield prediction using only the number of detected berries. These works have achieved a yield estimation or a number of actual berries estimation under ideal conditions where the number of visible berries in the images is proportional to the number of actual berries. From the previous works where an estimation of the number of actual berries was performed, Buayai et al. (2021) extracted a set of five features from the RGB cluster images, including the number of detected berries, and used them as input in six machine learning regression methods for the number of actual berries estimation. In Buayai et al. (2021) work, individual cluster images were used, avoiding leaf occlusions, and leaving only the berry occlusions in the clusters. As previously mentioned, the number of visible berries is only a portion of the number of actual berries of the vine and the biosystems engineering 218 (2022) 175e188176
percentage of exposed berries can vary depending on the canopy conditions in the fruiting zone. Leaf and berry occlusions are the main challenging issues to estimate the number of actual berries per vine on images acquired under field conditions from whole vines in commercial vineyards. The main goal of this work was to develop a new approach to estimate the number of actual berries in commercial vineyards using computer vision and deep learning, taking into account vineyard canopy features to overcome canopy occlusion impact on the number of actual berries forecast. 2. Materials and methods The experimental procedure comprised several steps (Fig. 1), which started from on-the-go image acquisition in commercial vineyards. The clusters were then segmented in the vine images, and the output was used for berries'and canopy elements segmentation. The number of berries, obtained after a second segmentation step, and the canopy features, obtained after canopy elements’ segmentation, were then combined in different datasets, and used in several regression models to obtain an estimation of the number of actual berries per vine. 2.1. Experimental layout The trials were carried out during season 2018 in a commercial vineyard located in Vergalijo (lat. 4227046.000 N; long. 148013.100 W; Navarra, Spain). Six grapevine (Vitis vinifera L.) varieties were selected, three red (Cabernet Sauvignon, Syrah and Tempranillo) and three white (Malvasia, Muscat of Alexandria and Verdejo). For each variety, a set of 16 vines (96 vines in total) were labelled and vertically delimited using plastic tape. The vines were trained to a vertical shoot positioned (VSP) trellis system with 2 m row spacing and 1 m vine spacing. The vines were partially defoliated at pea-size stage (Fig. 2b). Three leaves per shoot were manually removed to increase fruit exposure and aeration around the fruiting zone, as it is routinely conducted in many winegrowing areas worldwide, such as Rioja (Spain). For validation, at harvest time all clusters of the 96 labelled vines were manually collected, separated per each vine, weighted in the field, and transported to the laboratory. Berries of the clusters were then destemmed and manually counted. The number of actual berries per each vine was then obtained. 2.2. Image acquisition RGB image acquisition was carried out at night-time, at pea size phenological stage (green berries of around 7 mm diameter, according to Coombe, 1995), on the 3rd July 2018 (66 days before harvest), using a mobile sensing platform developed at the Fig. 1 eFlow-chart of the full process for assessing the number of actual berries in commercial vineyards. Fig. 2 eMobile sensing platform used for on-the-go image acquisition at nigh time in a commercial vineyard: a) An all-terrain-vehicle (ATV) modified to incorporate an artificial illumination system and a camera for acquiring RGB images on-the-go, and b) example of canopy vine image acquired in a commercial vineyard using the mobile sensing platform. Leaf and berry occlusions can be clearly observed in the partially defoliated grapevine. biosystems engineering 218 (2022) 175e188 177
University of La Rioja (Fig. 2a). The platform was a customised all-terrain-vehicle ATV (Trail Boss 330, Polaris Industries, Minnesota, USA). It moved at a speed of 5 Km/h and comprised a custom-made aluminium structure for image acquisition. The structure included an artificial illumination system, to enable image capturing at night-time, as well as an automated triggering system for RGB image acquisition. An RGB Canon EOS 5D Mark IV (Canon Inc. Tokyo, Japan) camera, mounting a fullframe complementary metaleoxideesemiconductor (CMOS) sensor (35 mm and 30.4 MP) equipped with a Canon EF 35 mm F/ 2 IS USM lens was used to acquire images with dimensions of 6720 4480 pixels. This camera was positioned at a height of one m from the ground, and at a distance of 1.80 m (perpendicular) from the canopy. The camera settings were manually fixed at the beginning of the experiment. These camera parameters were the following: exposure time of 1/1600 s, International Organization for Standardization (ISO) speed of 2000 and f-number of f/5.6. The external lighting system and nighttime imaging were necessary for acquiring vine images with homogeneous illumination (allowing the camera parameters to remain constant and avoiding solar light variability on the vineyard canopy during daytime) and to differentiate the vine under evaluation from the vines of the parallel row. Therefore, acquiring the images at night-time reduced the visibility of the parallel row and eased the segmentation of the elements present on the images. The mobile platform was described and used in previous works (Diago et al., 2019;Palacios et al., 2020). 2.3. Segmentation of clusters, berries and canopy elements All image segmentation steps were performed using a semantic segmentation approach. This way we tried to recognise what was in an image at pixel level, assigning a class to each pixel of an input image, usually through the use of deep learning (Deng &Yu, 2014). The SegNet deep learning architecture (Badrinarayanan, Kendall, Cipolla, 2017) was used in this work to carry out the semantic segmentation tasks. SegNet is formed by an encoder and a decoder. The encoder extracts feature maps by performing down-sampling operations while the decoder reverses these operations and performs the final pixel-wise labelling. The encoder is formed by the layers of a pre-trained convolutional neural network (CNN), while the decoder mirrored the encoder. In this work the layers of the architecture developed by the Visual Geometry Group (VGG), denominated VGG16 (Simonyan &Zisserman, 2015), were used for the encoder. The training hyperparameters of the SegNet remained constant for each image segmentation step. 2.3.1. Clusters’ segmentation The segmentation of the pixels belonging to the clusters was a necessary step towards the proper segmentation of the individual berries (in order to avoid false positives at segmenting berries outside of the bounds of segmented clusters), and to extract canopy features related to the clusters’ exposure and berry occlusion. Training the SegNet architecture with full-resolution images (6720 4480) for segmenting clusters would greatly increase the training time, and downscaling them would affect the quality of the visualization of the clusters in the image, especially considering that they are generally the smallest elements in the image. Instead, a set of 1646 image patches with a size of 1120 1120 were randomly selected and extracted from 18 full-resolution images that were manually labelled. These images were acquired from different vines than the 96 ones labelled and harvested, but from the same varieties and vineyard, and they were only used for training SegNet model. The image patches were resized to 560 560 pixels in order to reduce the training time. From the whole set, half of the image patches contained clusters (823 image patches) while the remaining ones contained other elements (in general, canopy elements). In order to obtain better results on clusters of different size, image patches were extracted considering two scales for the full-resolution images: images at their original scale, and images scaled at 80% of their original scale. An example of the result of the clusters’ segmentation step for an is shown in Fig. 3. 2.3.2. Individual berries’ segmentation After clusters'segmentation, the individual berries segmentation step was performed. The same image patches'set used to train the SegNet model of the clusters’ segmentation step was employed to train a new SegNet model that segmented the berries. New classes were defined for labelling each pixel in the images: these were “background”,“center”and “contour”.The class “background”represented non-berry pixels, while “center”and “contour”corresponded to the pixels that contained the center, and the contour of the berries, respectively. This approach was similar to the one followed by Palacios et al. (2020) for segmenting individual flowers and by Zabawa et al. (2020) for counting grapevine berries. Labelling the center and the contour of the berries independently eased the separation of berries too close together and enabled a counting of the segmented berries by counting only the groups of adjacent pixels from “center”class. The result of the individual berries segmentation step for an example image is presented in Fig. 4. 2.3.3. Canopy elements’ segmentation As in the previous step, the canopy was segmented following a semantic segmentation approach. Images were labelled into six main canopy elements'classes: “gap”,“leaf abaxial”(lower side of the leaf), “leaf adaxial”(upper side of the leaf), “shoot”, “trunk”and “cluster”. In order to label the images that were used to train the canopy segmentation SegNet, a semiautomatic approach was followed (except for the cluster class, where the manually labelling used for the clusters’ segmentation step was employed as well). A multinomial logistic regression (MLR) model was trained for each grapevine variety using colour features, and the images from each variety were segmented using the model trained for each individual variety (this colour segmentation approach was previously addressed in the work of Palacios, Diago, & Tardaguila, 2019). It was necessary to use a different MLR model for each variety due to the differences in the colour hue of some canopy elements, such as leaves and shoots, among the different varieties. Then, the output segmentation of all MLR models were grouped in a unique set and used as the training masks for a SegNet-VGG16 model, to obtain a model that was more robust to colour changes. Using a MLR colourbiosystems engineering 218 (2022) 175e188178
based segmentation as the initial approach enabled a fast training image labelling that would consume an enormous amount of time if it was done completely manually, as it would require a pixel-wise labelling of six canopy classes. A set of 300 pixels (50 from each class) was manually selected for each model, and for each pixel, six colour features were extracted: red, green and blue values, from RGB colour space, and L, a, b, from Commission Internationale de l'Eclairage (CIE) L*a*b colour space (also referred as CIELAB). The RGB colour space enables an adequate segmentation for most of the classes, while the CIELAB colour space has the lightness channel, L, separated from the colour channels, a and b, which is useful to segment dark pixels (e.g. pixels from the “gap”class). The use of a combination of multiple colour spaces has proved to be successful in several image processing works (Diago et al., 2019;Maktabdar Oghaz, Maarof, Zainal, Rohani, Yaghoubyan, 2015). The result of the canopy segmentation step for an example image is presented in Fig. 3c. 2.4. Canopy features and models After performing the segmentation steps, several features were considered to be included in the estimation models. These features were calculated from the segmented berries and canopy elements based on the idea that they might be useful to estimate the number of berries successfully overcoming the canopy occlusion. An overall set of 28 features were obtained considering two possible ROIs: a rectangle containing the fruiting zone (Figure 5a and b), as most of the clusters will be on the fruiting zone and this is normally where leaf removal operations are performed, and the clusters bounding boxes (rectangles with the minimum size to contain each cluster; Figure 5c and d), as it could be useful to analyse if the visible clusters are partially occluded or fully visible (clusters mostly surrounded by “gap”pixels could indicate the absence of leaves while being mostly surrounded by “leaf abaxial”pixels that the vine was defoliated on the imaging side). For the 28 features the estimated visible berries and all canopy elements segmented were used, along with the area of the fruiting zone and a binary feature indicating if the grapevine variety was red or white. The features were the following: F1: Number of estimated visible berries, calculated as the number of adjacent pixel groups from berry “center” class. F2: Ratio between the number of “gap”class pixels in the fruiting zone and the area of the fruiting zone. F3: Ratio between the number of “leaf abaxial”class pixels in the fruiting zone and the area of the fruiting zone. F4: Ratio between the number of “leaf adaxial”class pixels in the fruiting zone and the area of the fruiting zone. Fig. 3 eExample of the clusters'segmentation step: a) original image of a commercial grapevine canopy b) image with clusters segmented and c) image segmented into six canopy classes. Fig. 4 eExample of individual berries'segmentation: a) original image and b) image with berries segmented (including the center, in red, and the contour, in blue). (For interpretation of the references to color/colour in this figure legend, the reader is referred to the Web version of this article.) biosystems engineering 218 (2022) 175e188 179
F5: Ratio between the number of “shoot”pixels in the fruiting zone and the area of the fruiting zone. F6: Ratio between the number of “trunk”pixels in the fruiting zone and the area of the fruiting zone. F7: Ratio between the number of “cluster”pixels in the fruiting zone and the area of the fruiting zone. F8: Ratio between the number of “leaf”class pixels in the fruiting zone (defined as sum of “leaf abaxial” and “leaf adaxial”pixels) and the area of the fruiting zone. F9: Number of “gap”class pixels in the fruiting zone. F10: Number of “leaf abaxial”class pixels in the fruiting zone. F11: Number of “leaf adaxial”class pixels in the fruiting zone. F12: Number of “shoot”class pixels in the fruiting zone. F13: Number of “trunk”class pixels in the fruiting zone. F14: Number of “cluster”class pixels in the fruiting zone. F15: Number of “leaf”class pixels in the fruiting zone. F16: The average of all cluster bounding boxes ratios between the number of “gap”class pixels and the area of the corresponding bounding box. F17: The average of all cluster bounding boxes ratios between the number of “leaf abaxial”class pixels and the area of the corresponding bounding box. F18: The average of all cluster bounding boxes ratios between the number of “leaf adaxial”class pixels and the area of the corresponding bounding box. F19: The average of all cluster bounding boxes ratios between the number of “shoot”class pixels and the area of the corresponding bounding box. F20: The average of all cluster bounding boxes ratios between the number of “trunk”class pixels and the area of the corresponding bounding box. F21: The average of all cluster bounding boxes ratios between the number of “cluster”class pixels and the area of the corresponding bounding box. F22: The average of all cluster bounding boxes ratios between the number of “gap”class pixels and the number of “cluster”class pixels in the corresponding bounding box. F23: The average of all cluster bounding boxes ratios between the number of “leaf abaxial”class pixels and the number of “cluster”class pixels in the corresponding bounding box. F24: The average of all cluster bounding boxes ratios between the number of “leaf adaxial”class pixels and the number of “cluster”class pixels in the corresponding bounding box. F25: The average of all cluster bounding boxes ratios between the number of “shoot”class pixels and the Fig. 5 eFruiting zone and bounding box analysis: a) Fruiting zone delimited in blue dashed lines in the original image and b) in the segmented image - c) Clusters'bounding box delimited red rectangles in the original image and d) in the segmented image. (For interpretation of the references to color/colour in this figure legend, the reader is referred to the Web version of this article.) biosystems engineering 218 (2022) 175e188180
number of “cluster”class pixels in the corresponding bounding box. F26: The average of all cluster bounding boxes ratios between the number of “trunk”class pixels and the number of “cluster”class pixels in the corresponding bounding box. F27: The area of the fruiting zone (expressed in pixels). F28: A binary feature indicating whether the grapevine variety is a red or white variety, being “1”in the case of a red variety or “0”for white varieties. Each feature was calculated per vine, not per image. Although one image was selected per vine, the width of each vine in the image was determined by the width of the ROI for the fruiting zone, therefore any segmented object outside the width of that ROI was considered part of an adjacent vine and it was discarded. Some statistics for features from F1 to F27 and the output of the models (i.e. number of actual berries) are presented in Table 1. The features were combined to create four different datasets (Table 2). The first dataset, denoted as D1, contained a set of four features. These features were a combination of the number of estimated visible berries and three main canopy features calculated over the fruiting zone (Fig. 5a) and presented in the work of Diago et al. (2019) as they were stated in that work as the main canopy features. The second dataset, denoted as D2, included the whole set of 28 features presented above. Datasets D3 and D4 were reduced sets of D2. A sequential feature selection algorithm (Kohavi &John, 1997) was applied to D2 in order to obtain a smaller, relevant, set of features, and to remove features that could be redundant or irrelevant to the model. This algorithm is a wrapper method (an algorithm that implements a learning method to select features) that starts from an empty set and adds a feature to a given learning method in each iteration until no improvement is achieved. In each iteration, a leave-one-out cross validation is performed with the selected features, and the root mean squared error (RMSE) is calculated as the error metric for that subset of features. The subset of features with the lowest RMSE is selected as the final set. Therefore, the features of D3 dataset were selected to minimise RMSE using a support vector regression (SVR) model, obtaining a dataset with six features, while the features of D4 dataset were selected to minimise RMSE using a multiple linear regression model, obtaining a dataset with four features. The four datasets were tested with three regression methods: multiple linear regression (Freedman, 2009), support vector regression with a linear kernel (Bishop, 2006) and random forest (Breiman, 2001). 2.5. Evaluation metrics The metrics considered to evaluate the performance of the different algorithm steps involved in image processing were the following: Precision ¼TP TP þFP (eq. 1) Recall ¼TP TP þFN (eq. 2) F1¼2Precision Recall Precision þRecall (eq. 3) IoU ¼TP TP þFP þFN (eq. 4) Where for an arbitrary class “C", TP are the true positives (number of samples correctly classified as “C"), FP are the false positives (number of non “C" samples incorrectly classified as “C") and FN are the false negatives (number of “C" samples not classified as “C"). Of these metrics, precision (eq. (1)) represents the proportion of pixels correctly classified for each particular class while recall (eq. (2)) can be defined as the proportion of pixels identified for each class. Ideally, an algorithm performing perfectly would have a precision and recall of 1.0. In practice, a trade-off has to be achieved between Table 1 eAverage and standard deviation values calculated for canopy features from F1 to F27 and for the output of the models (number of actual berries). Features Statistics Average Standard deviation F1 403 293 F2 0,175 0,055 F3 0,136 0,038 F4 0,462 0,118 F5 0,091 0,066 F6 0,075 0,033 F7 0,052 0,036 F8 0,598 0,110 F9 1,452Eþ06 7,305Eþ05 F10 1,104Eþ06 4,331Eþ05 F11 3,763Eþ06 1,580Eþ06 F12 7,312Eþ05 5,474Eþ05 F13 5,938Eþ05 2,847Eþ05 F14 4,117Eþ05 3,108Eþ05 F15 4,868Eþ06 1,814Eþ06 F16 0,212 0,058 F17 0,202 0,106 F18 0,108 0,057 F19 0,075 0,042 F20 0,046 0,031 F21 0,346 0,061 F22 0,785 0,652 F23 0,362 0,192 F24 0,683 0,395 F25 0,275 0,236 F26 0,159 0,146 F27 8,138Eþ06 2,572Eþ06 Output 819 499 Table 2 eDataset and vineyard canopy features included. Dataset Vineyard canopy features D1 F1, F2, F7, F8 D2 All features (from F1 to F28) D3 F1, F4, F11, F12, F17, F27 D4 F1, F2, F11, F17 biosystems engineering 218 (2022) 175e188 181
false positives, which reduces precision, and false negatives, which lowers recall. The F1 score (eq. (3)) is defined as the harmonic mean of precision and recall (Chinchor, 1992) and reflects the accuracy of the classification, as it combines the recall and precision metrics. Intersection over Union (IoU, eq. (4)) is an evaluation metric that provides information about the overlap between the predicted and the ground truth masks, with values ranging from 0 to 1. An IoU equal to zero means that there is no overlap between the actual and predicted mask, while an IoU of 1 corresponds to the situation in which the union of the masks is the same as their overlap, therefore both are completely overlapping (Rezatofighi et al., 2019). For the canopy and the clusters'segmentation steps the four metrics were used, while for berries’ detection step only precision, recall and F1 score were computed (as it was considered a detection problem). For the estimation models, several regression metrics were selected. These were the coefficient of determination (R 2 ), the root mean squared error (RMSE) and the normalised root mean squared error (NRMSE), calculated as the ratio of RMSE over the average of the number of actual berries. In addition, the occlusion rate was calculated as an indicative value of the occlusion percentage of each vine using the following equation: ORið%Þ¼ABiEVBi ABi 100 (eq. 5) being ABiand EVBithe number of actual and estimated visible berries of the i-th vine, respectively. 2.6. Image processing and model validation To evaluate the performance of the image processing algorithm, two different set of images were labelled. For clusters' and canopy segmentation steps, a set of 411 image patches (560 560 pixels) were labelled using the multinomial logistic regression presented on the canopy elements’ segmentation section. For the individual berries detection step, the centre of the berries from a set of 60 full-resolution (6720 4480 pixels) images were manually checked and their number was used as the number of visible berries of the images. To evaluate the performance of the estimation models, two cross validation methods were considered for the set of 96 vines: the leave one out cross validation (LOOCV) and 8-fold cross validation. In LOOCV, the models were trained with 95 vines in each iteration and tested on the remaining vine, while in the 8-fold cross validation, the models were trained with 84 vines (14 per grapevine variety) in each iteration and tested with 12 (two per grapevine variety). 3. Results and discussion 3.1. Segmentation of clusters and other canopy features The results presented in Table 3 show the segmentation performance for each canopy class of the SegNet model. It can be observed that a high precision was achieved for all classes, being poorer for “leaf abaxial”and “shoot”classes. In terms of recall, all classes achieved a value above 0.80, which calls for a very satisfactory identification of pixels of the different canopy features. The balance between these two metrics is expressed as the F1-score. This index yielded values above 0.80 for all classes, with the exception of “leaf abaxial”and “shoot”, that also showed lower values of the IoU metric. However, IoU scores above 0.5 are considered as “good”. Overall, it can be concluded that the segmentation performance was very similar for most classes, with the exception of “leaf abaxial”and “shoot”, that were the most difficult classes to segment. This could be related to the semiautomatic labelling method used for training the SegNet model which relied only on colour features to segment each canopy class. At pea-size stage, the abaxial side of the leaves and the shoots displays a similar shade of green, which may have led to errors in the colour segmentation that were subsequently transferred to the final segmentation performed by SegNet. Although the segmentation performance of “leaf abaxial”and “shoot”classes could be improved by conducting manual segmentation, it would be time consuming. Since the segmentation metric values for “leaf abaxial”and “shoot” classes using the semiautomatic labelling method can be considered adequate and satisfactory, manual segmentation is therefore discarded. The works of Diago, Krasnow, Bubola, Millan, and Tardaguila (2016) and Diago et al. (2019) also presented a colour segmentation approach to identify canopy elements using a classifier based on the Mahalanobis distance. While, Diago et al. (2016) did not report any evaluation metrics for the image segmentation, Diago et al. (2019) did presented results of recall (named as specificity) and F1-Score. These results were similar for the “trunk”class in both works, while for the rest of the classes, Diago et al. (2019) achieved better results. This could be due to the fact that these authors used images acquired near harvest, where most of the canopy elements were easily distinguishable from each other by colour, in contrast to the pea-size stage images (used in the present work), where clusters, shoots and leaves had a similar shade of green. In another study, Palacios et al. (2019) presented a similar approach of colour segmentation using a multinomial logistic regression (MLR). The results obtained in that work also overcame the ones presented in Table 3 for the models trained on each variety independently. This could be explained by the fact that Palacios et al. (2019) used individual models, which were trained and applied on each variety individually, and to the fact that they used images acquired Table 3 eCanopy segmentation performance metrics using the SegNet model. Vineyard canopy class Precision Recall F1-Score Intersection over union (IoU) Gap 0.915 0.877 0.896 0.812 Leaf abaxial 0.664 0.833 0.739 0.586 Leaf adaxial 0.952 0.834 0.889 0.800 Shoot 0.686 0.808 0.742 0.590 Trunk 0.760 0.921 0.832 0.713 Cluster 0.799 0.974 0.878 0.782 Average 0.796 0.874 0.829 0.714 biosystems engineering 218 (2022) 175e188182
near harvest (showing substantial differences in colour between berries and leaves). For the models that combined all varieties in Palacios et al. (2019) work, similar results were obtained for all classes in both works, excepting for “trunk” and “grape”(“cluster”class in this work) classes, where poorer results were obtained in Palacios et al. (2019) work. P erezZavala et al. (2018) developed a method for visible cluster detection on several datasets of images acquired under different vineyard illumination and occlusion conditions employing texture descriptors and several classifiers. The method also employed a machine learning clustering algorithm for segmenting clusters'area, obtaining average results of 80.90% for precision (similar to the one obtained in this work for cluster class) and 70.82% for recall. In the work of Marani et al. (2021) several pretrained network architectures were tested to segment clusters and other elements related to the canopy (post, wood, leaves and background). While no results were shown for these elements, their best results at segmenting clusters were achieved by the VGG19 network with a mean segmentation accuracy of 80.58% and a IoU of 45.64%. Although their results in terms of IoU for cluster class were lower than the ones presented in this work, their image acquisition conditions were also more complex (day-light conditions and low-resolution images). 3.2. Individual berries’ detection The individual berries'segmentation yielded metrics of precision equal to 0.626, recall of 0.860 and F1-score of 0.722. These results reflect that the algorithm was able to detect the majority of the berries in the images (recall of 0.860) at a cost of having false positives (precision of 0.626), and that a good balance between precision and recall existed (F1-Score of 0.722). In addition, Fig. 6 shows that the number of estimated visible berries, defined as the number of berries detected by the algorithm was highly correlated to the number of visible berries, defined as the number of berries manually labelled in the images. This suggests that the algorithm's error was proportional to the number of visible berries in the images, existing a certain fixed deviation, since the slope was 0.60 instead of 1. In a previous study, Nuske et al. (2014) on their berries detection achieved recall values between 0.66 and 0.89, depending on the variety and the illumination conditions, the highest recall values were similar to the one obtained in this work for this step. Aquino, Millan, et al. (2018) reported a precision of 0.96 and a recall 0.88 for their berries detection algorithm. It can be observed that their performance, in terms of recall, was similar to the one obtained in this work, indicating that both algorithms detect approximately the same percentage of berries, but in terms of precision the results in the work of Aquino, Millan, et al. (2018) are better, potentially due to the full defoliation performed by these authors, that may have prevented false positive detection on leaves. In another recent work, Grimm et al. (2019) achieved a F1-score of 0.876 when detecting ripe berries. This result also overcomes the one obtained in the present work but ripe berries (after veraison) are easier to detect using computer vision as they are not so easily confused with leaves due to their colour, as it is the case with berries at pea-size, which are of similar green colour than leaves (Fig. 3a). Zabawa et al. (2020) achieved a precision of 0.85, a recall of 0.94 and a F1-Score of 0.89 using also a semantic segmentation approach with more recent architectures. These authors utilised a U-shaped encoderdecoder architecture that involved a MobileNetV2 (Sandler, Howard, Zhu, Zhmoginov, Chen, 2019) as the encoder, that produced an efficient feature extraction (as it was designed for mobile and embedded vision applications), and a DeepLabV3þ (Chen, Papandreou, Kokkinos, Murphy, Yuille, 2018) as the decoder, which refined the segmentation results focusing on object boundaries. In the work of Zabawa et al. (2020), images of three different grapevine varieties were acquired using a mobile phenotyping platform called Phenoliner (Kicherer et al., 2017). This platform was a modified harvester that included a background, an illumination system, and five cameras that captured images with a resolution of 2592 2048 at a distance of 75 cm from the canopy. The results obtained in this step prove that the segmentation method is useful to actually detect the berries in grapevine RGB images acquired on-the-go. 3.3. Estimation of the number of actual berries A good assessment of the number of visible berries in the images does not guarantee per se a satisfactory and accurate estimation of the number of actual berries per vine. This can be observed in Fig. 7, that shows a moderate correlation between the number of estimated visible berries in the images and the number of actual berries per vine (determined by manual counting at harvest). In addition, although the average occlusion rate per vine was 62%, the occlusion rate of the lowest occlusion vine was 11% and that of the highest occlusion vine was 90% (data not shown). This may be potentially due to differences in the canopy architecture and distribution of the various canopy elements, causing occlusion of berries, and to the shifts of grape phenology between varieties. As a consequence, it can be assumed that the number of actual berries per vine cannot be estimated using only the number of visible berries in the images. This is also supported by the information shown in Table 4 and Table 5, Fig. 6 eLinear correlation between the number of estimated visible berries and the number of visible berries per image for the validation image set (n ¼60 images). biosystems engineering 218 (2022) 175e188 183