scieee AI-readable full text Open interactive document viewer

Applications of artificial neural networks in three agro-environmental systems: microalgae production, nutritional characterization of soils and meteorological variables management

Franco Ortellado, Blas Manuel

Abstract

Departamento de Ingeniería Agrícola y Forestal

Full text

PROGRAMA DE DOCTORADO EN CIENCIA E INGENIERÍA AGROALIMENTARIA Y DE BIOSISTEMAS TESIS DOCTORAL Applications of artificial neural networks in three agro-environmental systems: microalgae production, nutritional characterization of soils and meteorological variables management Presentada por Blas Manuel Franco Ortellado para optar al grado de Doctor por la Universidad de Valladolid Dirigida por: Dr. Luis Manuel Navas Gracia Dr. Luis Hernández Callejo Passion is what gets you through the hardest times that might otherwise make strong men weak, or make you give up. Neil deGrasse Tyson v Acknowledgments This research and the thesis development were possible thanks to the funding from the Spanish Ministry of Education and Science via the predoctoral scholarships “Formación de Profesorado Universitario” (FPU) that was conceived to me [grant number FPU15/01707] and the project LIFE+ Integral Carbon, as part of the LIFE program and financial instrument of the European Union [grant number ENV/ES/001251]. This thesis is the result of a long journey and effort of many people on both sides of the world. Thanks to all and every single person from Paraguay (ñane reta) and Spain for this collective achievement, without the spiritual support during these four years this work would not be possible. Firstly, thanks to all my family, parents, Aurea and Blas, brothers, Paulo, José and Celeste, grandparents, uncles, cousins, and those families and friends that the blood does not relate us, but the souls do it in such as strong bound. Thanks for all you shared with me, happy moments (and sad moments, I’m sure there are, but I do not even remember them), knowledge, wisdom, love, and faith in God. Thanks to the University of Valladolid for being the house where the knowledge and creativity expressed in the past few years. Thanks to the directors of this work, Luis Manuel Navas and Luis Hernández for the guide and support to accomplish this thesis. Tanks to Matilde and Oscar for the support in the last months for the writing of the thesis. Thanks to all the friends who were the closest during the stay in this strange but familiar country of what is known as the old continent. Thanks to Anais, Mª. Ángeles, Charo, Ángela R., Alberto, Carmen, Bea, Laura, Marta S., Ángela B., Vicky, Felipe, Montse, and Maria; to the newest friends in the last part of this journey Ángel, José, Marta J., Shabnam, Miguel Angel, Alfredo, Jorge and Mario. A special acknowledgment for whose were not just good friends but also terrific teachers in these years: Jorge, Gonzalo, Víctor and Michel. Once again, it was shown that the humblest people are always the ones whose more good things have to share with you. Thanks for being on the train in which we shared fragments of the trip but had different arrival stations. vi For all the people in Spain, maybe I am wrong with this but, a song from The Beatles comes into my head when I think about all this ― the lyrics sing “… you and I have memories longer than the road that stretches out ahead…”. If we are lucky enough, life will give us other chances to laugh together again. Ipahápe, Ñandejárape ñande jajerovia ha mba’eichapa ñanembo’e Ñandejára Ñe’ẽ: Upéicharõ, peñemomirĩkena ha peñemoĩ Ñandejára ipuʼakapáva poguýpe, ikatu haguáicha haʼe penembotuicha oguahẽ vove pe tiémpo (1 Peter 5:6). Aguije Tupã. INDEX OF CONTENTS ix Index of Contents Acknowledgments ............................................................................................................. v Index of Contents ............................................................................................................. ix List of Tables ................................................................................................................... xiii List of Figures ................................................................................................................... xv List of Equations ............................................................................................................. xix List of Acronyms ............................................................................................................. xxi Abstract/Resumen ......................................................................................................... xxv 1. Introduction ................................................................................................................... 3 1.1 Research problems and hypothesis ................................................................................ 4 1.2 Objectives ....................................................................................................................... 4 1.2.1 General objective ..................................................................................................... 4 1.2.2 Specific objectives .................................................................................................... 5 1.3 Methodology overview .................................................................................................. 6 1.4 Motivation for the research ........................................................................................... 6 1.5 Innovative aspects of the research ................................................................................ 7 1.6 Thesis structure .............................................................................................................. 8 2. Theoretical Background ............................................................................................... 11 2.1 Challenges in agriculture and new technologies .......................................................... 11 2.2 Microalgae culture ....................................................................................................... 19 2.2.1 Microalgae overview .............................................................................................. 19 2.2.2 Specificity in microalgae production systems ........................................................ 21 2.3 Agricultural soil fertility................................................................................................ 24 2.3.1 Machine vision and color theory............................................................................ 24 2.3.1.1 Computer vision algorithms ............................................................................. 25 2.3.1.2 Color theory and systems ................................................................................. 27 2.3.2 Soil color and fertility properties ........................................................................... 30 2.4 Agricultural meteorology data ..................................................................................... 33 xvi Figure 18. Microalgae suspension for monoalgal and mixed algal cultures (a); and sample for light absorption measure in the colorimeter (b). .................................. 65 Figure 19. ANN architecture schema, with 31 neurons in the input layer, 45 in the hidden layer and 4 in the output layer. ................................................................... 67 Figure 20. Soil color analysis and characterization by ANN experiment scheme. ... 68 Figure 21. ECa raster map from the Veris data acquisition system at 36 cm (a) and 90 cm (b) depth; Segments for soil sampling generated through the processed raster data (c). .................................................................................................................... 70 Figure 22. Image acquisition system; light and camera setup (a) and shooter camera app for smartphone (b). ........................................................................................... 71 Figure 23. X-Rite ColorChecker Classic, color calibration card. ............................... 72 Figure 24. ANN architecture schema, with 3 neurons in the input layer, 8 in the hidden layer and 3 in the output layer. ................................................................... 73 Figure 25. VWS development methodology scheme. ............................................. 76 Figure 26. Meteorological station network of the InfoRiego program of the Agrarian Technology Institute of Castilla y León. ................................................................... 77 Figure 27. Light absorption spectrum of monoalgal cultures as obtained from the colorimeter (left), and relative absorbance conversion (right). .............................. 85 Figure 28. Correlation between the experimental composition of the samples and that predicted by the artificial neural network developed for the samples contained in Table 1. ................................................................................................................. 90 Figure 29. The user interface of DigiCIELAB CV software with the “Color Calibration” option (a) to generate the ANN model for the color analysis, and the “Colorimeter” option (b) for the color measurement of samples with the calibrated model. ....... 93 Figure 30. DigiCIELAB software workflow scheme. The part "a" of the scheme corresponds to the “Color Calibration” option, and the part "b" is regarding the “Colorimeter” option. .............................................................................................. 94 Figure 31. CV process to detect the object of study and ROI in the digital photograph: Original image (a), hue plane extraction (b), binary image (c), filled image (d), Petri plate detection (e), ROI for soil color extraction in the image center (f). ............... 95 Figure 32. Munsell color classification for soil samples dataset. ............................ 98 Figure 33. Soils samples with their respective lightness (L*) and OM (organic matter) contents (a); and samples with their respective red component (a* and R) and Fe (iron) contents (b). ................................................................................................. 102 Figure 34. Boxplots and mean comparison of meteorological observations grouped by the season of the year. ...................................................................................... 107 xvii Figure 35. X-Rite ColorChecker Classic color charts enumerated.......................... 173 xix List of Equations Eq. 1. Perceptron mathematical model. .................................................................. 38 Eq. 2. Hardlim activation function equation. ........................................................... 44 Eq. 3. Sigmoid activation function equation. ........................................................... 45 Eq. 4. Hyperbolic tangent activation function equation. ......................................... 46 Eq. 5. Softsign activation function equation. ........................................................... 47 Eq. 6. Softmax activation function equation. .......................................................... 48 Eq. 7. Rectified linear unit activation function equation. ........................................ 48 Eq. 8. Z-score normalization formula. ...................................................................... 51 Eq. 9. Min-Max normalization formula. ................................................................... 51 Eq. 10. Chain rule for the backpropagation algorithm. ........................................... 53 Eq. 11. Mean square error formula. ........................................................................ 57 Eq. 12. Root mean square of error formula. ............................................................ 57 Eq. 13. Mean absolute error formula. ..................................................................... 57 Eq. 14. Mean absolute percentage error formula. .................................................. 58 Eq. 15. Cross-entropy error formula. ....................................................................... 58 Eq. 16. Correlation coefficient formula. .................................................................. 58 Eq. 17. Determination coefficient formula. ............................................................. 59 Eq. 18. Relative frequencies conversion formula. ................................................... 66 Eq. 19. Minimum Euclidean distance match algorithm. .......................................... 74 Eq. 20. Inverse distance weighted interpolation method formula. ......................... 80 xxi List of Acronyms ACE Average cross-entropy Al Aluminum ANN Artificial neural networks Ca Calcium CE Cross-entropy error CH4 Methane CIE Commission Internationale de l'Eclairage (International Commission on Illumination) CIELAB Device independent color system; where L* represents lightness; a* the redness/greenness axis; and value b* the yellowness/blueness axis CMOS Complementary metal-oxide semiconductor CO2 Carbon dioxide CPU Central processing unit CSA Climate-smart agriculture Cu Copper CV Computer vision DNN Deep neural network EC Electrical conductivity ECa Apparent electrical conductivity EDA Exploratory data analysis ETo Reference evapotranspiration Fe Iron FTP File transfer protocol GHG Greenhouse gases GPU Graphics processing unit HSI Hue, saturation and intensity HSL Hue, saturation and luminance HSV Hue, saturation and value IDW Inverse distance weighted ISDW Inverse square distance weighted ISO International organization of standardization K Potassium L*a*b* Lightness, redness/greenness and yellowness/blueness axis in the CIELAB color system MAE Mean absolute error MAPE Mean absolute percentage error xxii Mg Magnesium ML Machine learning MLP Multilayer perceptron Mn Manganese MRL Multiple linear regression MSE Mean square error MV Machine vision N Nitrogen Na Sodium NH3 Ammonia NO2 Nitrous oxide NO-3 Nitrate OM Organic matter P Phosphorous pH Potential of hydrogen Precip Precipitation R Correlation coefficient R2 Determination coefficient RAM Random access memory relu Rectified linear unit RFR Random forest regression RGB Device dependent color system based on the red, green and blue components RH Relative humidity RMSE Root mean squared error RNN Recurrent neural network ROI Region of interest ROI Regions of interest SOC Soil organic carbon St. dev. Standard deviation tanh Hyperbolic tangent Temp Temperature TSI Total solar irradiation vis-NIR Visible–near infrared VWS Virtual weather stations w/v The weight volume ratio in g of soil per mL of water WS Wind speed XOR Exclusive OR logic problem ABSTRACT xxv Abstract Agriculture is an essential human activity, highly dependent on meteorological conditions and focus of research and innovation to afront several challenges. Climate change, global warming, and the degradation of agricultural ecosystems are just a few of the problems that humans are facing for continuing the essential food production. Seeking the innovation in the agricultural sector, three main research topics were considered for this thesis; such as microalgae production, soil color and fertility, and meteorological data acquisition. These subjects have increasing roles in agriculture, specifically under the uncertainty in the future of food production. Microalgae are a healthy alternative for crops fertilization and soil sustainability; while the soil fertility parameters need to be more studied to aim lower cost and faster analysis methods to help the management. Agriculture, as a highly weather-dependent activity, needs meteorological data to anticipate events, planning, and management crops in an efficient mode. These topics were selected with the purpose to improve the current state of the art, propose new alternatives based, mainly, in the application of artificial neural networks (ANNs) as a novel manner to solve the problems and generate knowledge of direct application in crop systems. ANNs are a useful tool to modeling and solve complex nonlinear problems; they are a mathematical model of the animal brain and their ability to deal with complicated issues drive the scientific community to used them to find solutions hardly found with other techniques. The main objective of this thesis was to generate ANN models capable of addressing agricultural related problems as an alternative to traditional and more expensive methods for management, analysis, and data acquisition in the crop systems. For the microalgae culture experiments, monoalgal and mixed algal culture spectral signatures from light absorption measurements were analyzed. Additionally, an ANN was used alongside the spectral signature in order to create a model capable of classifying the microalgae cultures and determine the species present in suspension. The results show that the ANN was capable of distinguishing between monoalgal and mixed algal cultures, identifying the microalgae species in the monoalgal cultures and providing the approximate composition of mixed algal cultures. Regarding the soil study, soil samples were classified according to the Munsell color notation, and the obtained hues were used to group soils and perform statistical analysis over their fertility attributes. In addition, RGB and L*a*b* colors 4 1.1 Research problems and hypothesis The research topics, as mentioned earlier, microalgae production, soil color and fertility, and meteorological data are a niche for innovation, as other fields in the agricultural sector. This innovation is a crucial concern due to the fact of the challenges for the agricultural sector in the present and future scenarios. Considering this, ANNs were applied in this research to solve issues in the territory of these topics; in the microalgae assessment, the issue was to develop a fast and reliable method to analyze microalgae cultures; in the soil color and fertility study, the issue was to relate color and fertility in order to find relationships between both. Lastly, in the meteorological data assessment, the matter was to generate a method for data acquisition in locations with no availability of weather stations. In this thesis, the hypothesis formulated for the research topics were: - Microalgae light absorbance analysis through ANNs can elucidate the biological culture composition and be used as an alternative for the assessment of microalgae cultures. - Soil color analysis can provide information about fertility parameters, and ANNs can be applied to map this relationship and conceive a model for rapid soil evaluation. - Meteorological data acquisition can be performed by using interpolation techniques and ANNs, alongside real data from weather stations to interpolate values with high accuracy. 1.2 Objectives 1.2.1 General objective The general objective of the research is to generate ANN models capable of addressing agricultural related problems as an alternative to traditional and more expensive methods for management, analysis, and data acquisition in the crop systems. Introduction 5 1.2.2 Specific objectives The specific objectives were separated for each of the main research topics of the thesis. The established objectives are presented below. In the microalgae culture study: - Study light absorption spectra of the different microalgae species. - Evaluate the ability of ANN to differentiate between monoalgal and mixed algal cultures. - Determine the feasibility of using ANN to estimate the biological composition of mixed microalgae cultures. Regarding the soil color and fertility study: - Analyze soil color in order to determinate its capability to group soils according to fertility levels. - Develop a computer tool for soil color measurement and classification. - Describe soils fertility parameters using color analysis through ANNs. Concerning the meteorological data acquisition study: - Evaluate and compare the accuracy of several interpolation algorithms ― including the traditionally used and ANNs. - Study the seasons and the extreme phenomena effects in the performance of interpolation methods. - Devise a set of algorithms to acquire real meteorological data, process them, and generate accurate estimations in distinct locations, economically and straightforwardly. 6 1.3 Methodology overview To conduct the experimental phase of the thesis, three experiments were carried out. For the microalgae experience, monoalgal and mixed algal culture spectral signatures from light absorption measurements were analyzed. Additionally, an ANN was used alongside the spectral signature in order to create a model capable of classifying the microalgae cultures and determine the species present in suspension. For the soil study, soil samples were classified according to the Munsell color notation, and the obtained hues were used to group soils and perform statistical analysis over the fertility parameters. In addition, RGB and L*a*b* colors were used as input in ANNs to create models to describe the soil fertility parameters. The RGB and the L*a*b* colors were obtained from digital color photographs and computer software programmed to perform quick and accurate measurements. Finally, for the meteorological data acquisition study, daily data from a weather station network were used to perform interpolations with several methods, including ANNs and traditional techniques. With these interpolations algorithms, the development of virtual weather stations (VWS) was proposed, scripts automatically acquire, and process meteorological data were coded, and afterward, perform estimations in different locations without weather stations availability. 1.4 Motivation for the research ANNs studies are nowadays a multidisciplinary science, a crucial factor for that is their application to solving problems in different areas, pursuing the modernization of techniques, better prediction of phenomena, classification, and estimation of everyday and important matters. The room for applications of ANN are infinitely potential, and the scope of this tool is changing the manner scientific process data and conduct experiments. The agriculture is a niche for infinite research since with new days, new problems and challenges appear and need solutions, and at the same time, better solutions for older problems are also needed. With the aim of solving agricultural related subjects, this research was motivated and focused on microalgae culture analysis, soil color and fertility studies and meteorological data acquisition through interpolation methods. Introduction 7 Microalgae are an interesting group of microorganisms, and their cultivation is a promising tool for carbon sequestration and biomass utilization, for example, as an organic soil amendment ― therefore their agricultural and environmental importance will arise in the next few years. Regarding soil fertility, this parameter is one of the most important in the production ecosystem, and its knowledge is vital for the correct and sustainable management of soils. Concerning the meteorological data, agriculture as an activity that is highly dependent on the weather requires the availability of these type of data can be useful for improving, for instance, the irrigation calculus and planning, crop phenology, pest studies and modeling, among other aspects. 1.5 Innovative aspects of the research This thesis was designed with the practical application of technologies and techniques in mind, to solve specific problems in agricultural disciplines. The innovative aspects and contributions of this research are methodologies and tools that can be used directly in culture systems, been microalgae or conventional crops. The main contributions are cited below: - An ANN to elucidate microalgae species in suspensions was made. The model provides a fast and powerful tool for microalgae culture management at the commercial scale, provides information regarding the biological composition of cultures, and approximates the relative proportions of species in the suspensions. The developed method means a cheaper and faster analysis method in comparison with the typically used. - A scientific paper with the title “Monoalgal and mixed algal cultures discrimination by using an artificial neural network” (Appendix A) was published in the “Algal Research” journal. A Q1 journal in Biotechnology and applied microbiology journal with an impact factor of 3.75. - The development of computer software, the DigiCIELAB, which is a powerful tool to perform the color measurement in a faster and more affordable manner in comparison with traditional colorimeters. The software was selected by the University of Valladolid in the “Prometeo” 2017 program and intellectually protected as a result of this selection (Appendix B). 8 - The development of the VWS algorithms to acquire and process meteorological data from weather station networks with the ultimate purpose of estimate data where no weather station is available. The VWS is a feasible alternative for the acquisition of meteorological data of importance for agricultural activities. 1.6 Thesis structure After this overview of the thesis research problems, hypothesis, objectives, and other introductory aspects treated in this chapter. The remaining of the manuscript is organized as follows: a theoretical background with the literature of the current estate of the art in agriculture challenges, ANNs, brief microalgae production concepts, soil fertility and color theory and meteorological data importance ― with the descriptions of techniques used in each one of these topics ― is presented in Chapter 2. The materials used and the methodology applied in the experimental phase for the microalgae, soil color, and meteorological data studies are detailed in Chapter 3. The results of these experiments alongside with discussion with the most relevant literature to contrast the obtained results are presented, separately by research topic, in Chapter 4. The conclusion, the general from the overall research experience and the specifics regarding each topic, are illustrated in Chapter 5. Finally, the cited bibliography is referenced (Chapter 7), and the additional information is shown in the Appendix section. THEORETICAL BACKGROUND 11 2. Theoretical Background In this chapter, the theoretical background that supports the thesis development will be presented, including the basic concepts for the methodology applied to seek the research objectives. This section will describe a brief agriculture panorama, future challenges and possible manners to address the issues with technologies and alternative agricultural practices (section 2.1); the microalgae production and its agriculture importance (section 2.2), soil fertility and color analysis (section 2.3) and meteorological data for agriculture (section 2.4) are treated as well. Finally, an overview of ANNs, concepts, modeling and training process, and applications of this technique (section 2.5). 2.1 Challenges in agriculture and new technologies Achieving maximum crop yield at minimum cost is one of the goals of agricultural production from an economic point of view. However, since 1950, a climatological transformation is taking place as a consequence of several phenomena like deforestation, the emissions of greenhouse gases (GHG) increment, ozone loss, the increment of global temperature, change in precipitations regimes and other climate change effects, with a higher acceleration from the 90s (Smith et al., 2014; Steffen et al., 2015). These transformations, specifically regarding the anthropogenic climate change, are considered as one of the most significant environmental, social and economic threats to the future world (Ghosh et al., 2019) and agriculture is one of the most exposed to climatic impacts (Martins et al., 2019; Tran et al., 2019). Global food demand is expected to increase considerably in the near future as a result of the growing population. Agricultural research needs to step up to meet a more sustainable development for food production, human nutrition, climate change and environmental protection in a world with 9.7 billion people by 2050 (Thornton et al., 2018). Considering that, with the modernization of agriculture, high inputs of fertilizers, pesticides, and mechanical energy are demanded for labors (Erb et al., 2008; Wu et al., 2017), GHG emissions such as carbon dioxide (CO2) (Li et al., 2019), nitrous oxide (NO2) which is one of the most harmful GHG (Liu et al., 2019; Rowlings et al., 2013; Wolff et al., 2017) and others 12 such as methane (CH4), nitrate (NO3-), ammonia (NH3) will increase (Sanz-Cobena et al., 2017a). It is estimated that agriculture contributes 11% to the anthropogenic GHG emissions, approximately 5.3 Gt CO2 equivalents in 2010 (Figure 1), the trend indicates that it will increase 9% in 2030 and 18% in 2050 with respect to 2010, or up to 37% more considering the 90s as base reference (Tubiello et al., 2014). Figure 1. CO2 emissions estimations for the 1990s, 2000s and 2010 and projections for 2030 and 2050 for agricultural related activities (elaborate from data of Tubiello et al., 2014). Climate change is a danger for crops and food production (Fitzgerald et al., 2019; Manners and van Etten, 2018) and risk for smallholder farmers and livestock sector, particularly in dryland regions (Hansen et al., 2019; Herrero et al., 2015). Globally, annual climate variability accounts for roughly a third (32–39%) of the observed crop yield variability (Ray et al., 2015); increasing the yield losses in warmer years; it is estimated that for each degree Celsius (⁰C) of mean global temperature, there is a reduction in global wheat grain production of about 6% (Asseng et al., 2015). Minimizing and preventing environmental degradation is a critical sustainability challenge of the present and coming decades (Bais-Moleman et al., 4.6 5.0 5.3 5.8 6.3 3.0 3.5 4.0 4.5 5.0 5.5 6.0 6.5 7.0 1990s 2000s 2010 2030 2050 (in Gt CO2) Emissions of CO2 equivalent from agricultural related activities Theoretical Background 13 2019). Climate-smart agriculture (CSA) is widely promoted as an approach for reorienting agricultural development under the realities of climate change (Figure 2), seeks to meet three challenges: improve the adaptation capacity of farming systems to climate change, reduce the greenhouse gas emissions of these systems, and enhance agricultural productivity (Acosta-Alba et al., 2019; Jagustović et al., 2019; Thornton et al., 2018). Figure 2. Climate-smart agriculture foundations for the development of sustainable agriculture. A key reason for the emergence of the CSA concept is the recognition that agriculture, and related food security issues, require a synthesized approach which may not be achieved by tackling climate mitigation and adaptation objectives separately (Long et al., 2016). CSA is the combined strategies to respond to the challenges of making food security, providing public goods and ecosystem services to society (Bais-Moleman et al., 2019). Climate-Smart Agriculture (CSA) B. Promote adaptation A. Enhance agricultural productivity C. Reduce GHG emissions N2O CO2 CH4 20 From all these thousands of species, only a few are being produced commercially today, primarily for high-value products (Rickman et al., 2013). Microalgae have been proposed for a wide range of applications, from the production of foods and animal feed, cosmetics, biofuels and wastewater treatment processes (Borowitzka, 2013; Chisti, 2007; Enzing et al., 2014; Olguín, 2012). Microalgae are cultivated at an industrial scale in two widespread systems, the open-culture systems and the closed-culture systems (Figure 3). Open-culture systems, for example, open ponds and raceways, are the simplest and less expensive in comparison to closed-culture systems. Open-culture systems are almost always located outdoors and rely on natural light for illumination. Unfortunately, the microalgae culture can be easily contaminated, and little control of the operating conditions can be made on such systems (De Andrade et al., 2016). Closed-culture systems, such as tubular photobioreactors, allow a certain control level of operating conditions and to avoid contamination, being possible to obtain high-value algal products. Closed photobioreactors may be located indoors or outdoors, but the outdoor location is more common because it can make use of free sunlight (Molina Grima et al., 2003). Figure 3. Microalgae production system in an open-culture system (a) and a closedculture system (b). Despite the large variety of applications proposed, only a few are presently performed at the commercial scale using a limited number of algal strains. Examples of this are the production of carotenoids, beta-carotene and astaxanthin from Dunaliella salina and Haematococcus pluvialis (Forján et al., 2014), biomass for foods from Chlorella vulgaris and Spirulina platensis (Fradique et al., 2010), and a b Theoretical Background 21 biomass for aquaculture from Nannochloropsis gaditana, Tetraselmis suecica and Isochrysis galbana T-ISO (Shields and Lupatsch, 2012). Microalgae have other applications such as carbon sequestration by transferring the environmental CO2 into the microalgae mass (Acién et al., 2012; Sobczuk et al., 2002), wastewater treatment (Acién et al., 2016; Ledda et al., 2015), and microalgae biomass fertilizers (Mulbry et al., 2005; Raposo and Morais, 2011; Shaaban, 2001). For fertilization, microalgae are a source of one of the scarcest nutrient, the phosphorus, which is abundant in the algal mass (Egle et al., 2016; Melia et al., 2017; Mukherjee et al., 2015). Despite being effective, these alternative uses are not implemented at commercial scale yet. However, recent advances allow the scientific community to be optimistic in the development of production technologies which are more economically viable, and within the next 10 to 15 years these technologies will allow applications that are not viable at present (Acién et al., 2014). New technologies are in development to decrease the cost of microalgae production such as species selection, supplies optimization, and better condition controls (Collet et al., 2011; Madkour et al., 2012). These advances make the microalgae biomass one of the most attractive alternatives for fertilizers in agriculture and CO2 sequestration to achieve the CSA objectives of mitigating and reduce the GHG emissions and improve agricultural sustainability. 2.2.2 Specificity in microalgae production systems For microalgae cultures and the harvested product, it is often required the maintenance of monoalgal cultures when a specific compound is desired. Whereas when focusing on biofuel production or wastewater treatment, the utilization of mixed cultures is usually acceptable. In the case of mixed algal cultures, the relative composition is gradually modified according to changes in environmental or operational conditions (Godos et al., 2009). Monitoring the biological composition of microalgae cultures is a necessary task, generally performed through routine microscopic examination. By means of light microscopy, an expert can distinguish the presence of a “contaminating” microorganism and whether the current algal strain is close to the expected value. However, only highly skilled taxonomists are capable of correctly recognizing algal strains (their species and genera) using light microscopy 22 observation solely based on morphology – this is because most of the strains are small round cells with similar features, only a few have easily recognizable morphology. Some microalgae can be classified by a computer using high-resolution images for the morphological characterization (Walker and Kumagai, 2000). This classification is performed automatically only with simple forms such as circle, ellipse or cell sizes; complex forms or similar species require the intervention of an operator, but the process is under human error that can lead to confusion (Mirto et al., 2015). Alternative methods, based on omics allow to accurately identify the microalgae strains in cultures (Godhe et al., 2002), but these are expensive and require much time (reducing time can be useful for making operation process decisions). As a standard method, biochemical analyses, such as the chlorophyll to carotenoid ratio and the fatty-acid profile have also been used as tools for verifying the biological composition of microalgae cultures; however, their precision is limited. These methods are unable to identify the presence of low-level contamination; likewise, they take time, although their cost is much lower than that for the omics methods (Serive et al., 2017; Sydney et al., 2011). Microalgae species identification up to the phylum or class levels has been achieved based on the fluorescence properties of photosynthetic pigments using flow cytometry (Cellamare et al., 2010). Employing light-emitting diode induced fluorescence analysis is possible to differentiate between Anabaena sp. and Cylindrospermum sp. cells by comparing the fluorescence spectra (Ng et al., 2017). Microalgae have photosynthetic pigments providing different spectral signatures for different species; thus, it is possible to build classes based on the presence of pigments. As a result, the relative content of chlorophyll, carotenoids and other pigments can be used to differentiate between diatoms, red/green/brown microalgae, and cyanobacteria groups through their light absorption spectra (Serive et al., 2017). Microalgae species can be distinguished by their spectral signature. As an example, Rivularia M-216 exhibits a different absorbance signature to that of Anabaena variabilis ― the heterocyst absorbance from Rivularia is more than double than that from A. variabilis at wavelengths between 540 and 620 nm; this variation is a result of the different phycocyanin and chlorophyll contents (Nozue et al., 2017). The spectral signatures of Botryococcus braunii, Chlorella sp. and Chlorococcum littorale allow identifying them by comparing their absorption indexes (Lee et al., 2013). Moreover, the absorption spectrum in the 400 to 700 nm Theoretical Background 23 range is used to determine the extinction coefficient of the biomass, a specific microalgae strain and culture conditions indicator (Rubio Camacho et al., 2003). Based on this, absorption properties could be a possible approach to distinguish between microalgae species (Coltelli et al., 2017). ANNs are a powerful tool for finding relationships between experimental data and the phenomena behind these data. ANNs have been used to predict harmful algal blooms in lakes (Recknagel, 1997; Tian et al., 2017), as well as microalgae growth and biomass concentration under laboratory conditions and outdoor environments (García-Camacho et al., 2016; Sharon Mano Pappu et al., 2013). In microalgae identification using ANN, extracted features such as the perimeter, shape, area and Fourier Transform of microalgae micrographics were used to train a model capable of identifying the genera Navicula, Scenedesmus, Microcystis, Oscillatoria and Chroococcus (Mosleh et al., 2012). Microalgae micrographic image processing and color analysis were used with ANN to achieve taxonomic accuracy of up to 99% by first detecting a cell in the image and subsequently extracting the detected cell color (Coltelli et al., 2017). Considering the capacity of ANNs and the light spectral signature properties of microalgae, it is possible to consider a fast and economical method derived from these two elements for the biological monitoring of cultures, especially those that need a specific species composition, pure or in proportions. 24 2.3 Agricultural soil fertility Fertility is the soil ability to supply essential plant nutrients in adequate amounts, proportions, and moments for the growth and reproduction of plants (Bünemann et al., 2018). In the last few decades, intensive agricultural management practices in European agriculture have resulted in soil degradation (Freibauer et al., 2004; Virto et al., 2015). OM and N content in soils are now 30% to 60% lower than their undisturbed (virgin) equivalents, about 44% of southern Europe lands exhibits low OM content as a result of intensive cultivation (Rusco et al., 2001). In the framework of sustainable agriculture, better soil management must be addressed to preserve and increase soil fertility in agricultural lands. Among the environmentally friendly labors such as non-tillage, pesticides reduction and administration of fertilizers, these can be used more effectively to minimize supplies spending and contamination. A fast method for estimating soil nutritional content can help in the process of decision-making. Soil color contents some information about fertility (Barrios and Trejo, 2003; Fleming et al., 2004; Gray and Morant, 2003). The analysis of this property through machine vision can help to characterize soils. In consideration, the description of machine vision, color theory, and soil color properties will be reviewed in the section below. 2.3.1 Machine vision and color theory Computer vision (CV) is the science of the design and operation of the software for the analysis and processing of image, whereas Machine Vision (MV), means a more global concept: the study of the hardware environment, the image acquisition techniques and the software, CV, needed for the development of applications (Davies, 2018). Following its origin in the 1960s, MV has experienced significant growth, and its applications started to expand until the present in diverse fields like medical diagnostic imaging, factory automation, remote sensing, forensics, autonomous vehicle and robot guidance (Brosnan and Sun, 2004). Regarding the hardware for MV, the acquisition system is usually composed of three main components: a color digital camera, an illumination source, and an image processing software. The lighting source must provide uniform and consistent illumination across the sample to photograph, color temperature is also Theoretical Background 25 considered and usually fixed around 5,000 K to 5,500 K. To ensure uniform illumination conditions, multiple light sources can be employed as long as they are homogeneous (Tarlak et al., 2016; Valous et al., 2009). The camera is usually located at a certain distance so that the measurement does not interfere with the illumination since the resulting images are highly affected by the light (ten Bosch and Coops, 1995). Besides a digital color camera in a single setup, the most used for computer vision, dual cameras setups for stereoscopic image analysis can be used in controlled conditions as was previously described, the stereoscopic setups are used for objects dimension calculation inside the images (Zion, 2012). Multiple color cameras mounted in drones are usually used in outdoors studies (Rumpler et al., 2017), or infrared cameras both indoors and outdoors (Celenk, 1990) are some acquisition systems employed in MV. In the software aspect, the CV is composed of several algorithms and a vast number of variations of them that help the processing of a digital image. Between the most common algorithms are segmentation, thresholding, shape and edge detections, morphological operations, color extraction, among others (Chen, 2015). 2.3.1.1 Computer vision algorithms In CV processing, several algorithms intervene; the main ones are segmentation, thresholding, shape and edge detections. A brief description of these algorithms is presented in the paragraphs below: A. Segmentation Segmentation consists of separate uniform and homogeneous regions of a digital image, generally objects, with respect to some background or other heterogeneous segments of the image (He et al., 1985). Segmentation, for instance, is used to differentiate landscape sections in remote sensing or different tissues in biomedical images and to extract these objects/parts from the background, identifying blood cells in biomedical pictures, detecting machinery parts, among others (Prats-Montalbán et al., 2011). 26 B. Thresholding Thresholding is a type of segmentation. A given threshold value is applied to the pixels in a gray-scaled image, and pixels are classified according to this threshold in a binary image that results from this operation (Sudarsan et al., 2016). A binary image indicates that the pixel value of “1” is where the pixel past the threshold and the pixels with value “0” indicate the segment did not pass the threshold (Promdaen et al., 2014). The binary image is a mask that can be used, which contains less information and is easier to process by the machine (Zhang et al., 2017). For instance, after the thresholding, a morphological operation can be performed in order to fill the possible holes presented in the binary image (Mery and Pedreschi, 2005). C. Shape detections Shape detections can be performed through several algorithms that compare matrices corresponding to shapes, such as circles, squares, triangles, and other more complex figures, with a sample image in which these matrices are compared. When the matrix patterns are found in the image, it means that a given shape is present in the image (Allili and Corriveau, 2007). Object detection is an essential and challenging task; this becomes particularly tricky in images with a strongly cluttered background, and with objects subjected to scale changes, rotation changes, and substantial intra-class variations (Wei et al., 2017). D. Edge detection Edge detection methods are commonly applied to the image to assess any change in the intensity profile of neighboring pixels; therefore, a substantial intensity change between an object and the background corresponds to accurate edge information and yields excellent results (Williams et al., 2014). The main challenges for edge detection are due to pixels noise in the background or another object interfering with the edge of the object of interest (Lu et al., 2010). In addition, there is the concept of regions of interest (ROI). The ROI is assigned on the image in specific coordinates and has a given size, for instance, 60 × 60 pixels at the position x = 200 and y = 200 in a digital image (Sumriddetchkajorn et al., 2014). From this ROI, multiple operations and analysis can be performed inside this region, including edge detection, among others (Losson et al., 2013). Theoretical Background 27 2.3.1.2 Color theory and systems The color is the result of the interaction of certain light wavelengths with an object; this is an attempt to relegate color to the purely physical domain. Instead, it is proper to state that those stimuli are perceived to be of a particular color when viewed under specified conditions (Fairchild, 2013). In terms of computer vision, color is given by the R (red), G (green), and B (blue) tristimulus values. A color image is usually provided by the three values, RGB, at each pixel. In a typical 8-bit image, each component can take values from 0 to 255 (Celenk, 1990). Values of RGB show the total amounts of the three primaries that are required to make a color inside the RGB color system space (Gershon, 2005). In this system, the pure red color is given by (255, 0, 0), white by (255, 255, 255), black by (0, 0, 0) and so on combining the components. A scheme of the RGB system is shown in Figure 4. Figure 4. RGB color system diagram. Purple, amber and light blue color RGB values are described. Purple (224, 64, 251) Amber (255, 160, 0) Light blue (3, 169, 244) B G R 0 255 255 255 28 In CV, RGB color features are important as a mechanism to a rapid and non-destructive inspection of objects (Xia et al., 2016). The color can be used to estimate the quality and maturity of tomatoes, citrus, cranberries, mangos and estimate chemical components associated with indices of quality, such as Brix degrees (Francis, 1995; Jha et al., 2007). Others studies with color images were conducted for meat color and quality (Trinderup et al., 2015), water purity analysis (Andrade et al., 2013), and Earth surface color and texture assessments (Zhao et al., 2016). In these types of studies, a colorimeter is typically the first option for color measurements. However, the colorimeter shows some limitations because of the non-homogeneous color in the surface of the object to analyze, especially in food engineering and research ― colorimeters analyzes points, not the whole surface (Barbin et al., 2016; Girolami et al., 2013; Yam and Papadakis, 2004). In this aspect, the color analysis of the digital images through CV is an advantage. In research, color is frequently represented using the L*a*b* coordinates of the CIELAB color space. This color model is considered of uniform proximity, i.e., the distance between two colors in a linear color space corresponds to the perceived differences between them (Mendoza et al., 2006). The CIELAB system was proposed by the International Commission on Illumination (CIE - by its initials in French, the Commission Internationale de l'Eclairage). In this color space, the positive a* axis points in the direction of red color, the negative axis in the direction of green stimuli; positive b* points in the direction of yellow stimuli; negative b* in the direction of blue stimuli. L* is the luminance; thus a value “0” of lightness indicates the black or non-lightness and “100” represent the white in saturated light conditions (Schanda, 2007), a scheme is presented in Figure 5. The L*a*b* color is device independent, providing consistent color regardless of the input or output device such as digital camera, scanner, monitor, and printer (Yam and Papadakis, 2004). This is the crucial difference between RGB and L*a*b*, RGB is device dependent and the color of a given picture, under equal illumination conditions, leads different cameras to obtain different colors, in fact, vastly different in colors are usually obtained among cameras (Ilie and Welch, 2005; Kim et al., 2012). There are other device dependent color systems such as HSL (for hue, saturation and luminance), HSI (for hue, saturation and intensity) and HSV (for hue, saturation and value), which are mathematically similar and can be calculated Theoretical Background 29 from the RGB; these color systems are commonly used in image processing (Saravanan et al., 2016). Figure 5. CIELAB color space representation. Purple, amber and light blue L*a*b* colors are described. To address the non-uniformity results of digital cameras, a colorchecker is often used. A colorchecker is a color calibrated array of cards that help to tune cameras and counteract the variability among them. Colorchecker have different natural colors (red, blue, green, yellow, and others), grey tones, black and white; and with computer software the image color is modified according to the color cards to balance the image (Girolami et al., 2013; Kirillova et al., 2017; Potočnik et al., 2015). Additionally, colorchecker can also be used to convert the RGB values of the image into CIELAB standards if the manufacturer provides the L*a*b* data of the cards. The RGB to L*a*b* color transformation can be performed through several methods such as equation systems (Barbin et al., 2016), quadratic models, (Tarlak et al., 2016), linear models and ANN (Afshari-Jouybari and Farahnaky, 2011). Between these methods, ANN present remarkable results for this conversion (Afshari-Jouybari and Farahnaky, 2011; León et al., 2006; Pedreschi et al., 2006; +b* -b* -a* +a* L* 128 0 -128 128 -128 100 Purple (58.49, 82.52, -61.83) Amber (73.80, 26.54, 78.22) Light blue (65.69, -9.60, -47.35) 36 2.5 Artificial neural networks An ANN is a mathematical/computational model that is inspired by the structure and functional aspects of biological neural networks. It is a highly accepted technology to alternatively solve complex problems (Bouselham et al., 2017). Comparative surveys show that accuracy of ANN methods is superior to that of traditional statistical methods in dealing with problems, especially regarding nonlinear patterns (Bahrammirzaee, 2010). ANN belongs to machine learning (ML) set of techniques. ML can be described as such algorithms capable of producing learning by the mathematics in the models, and the fine-tuning of several parameters responsible for fitting the input data to the target. The optimization of such parameters could be understood as an unknown black-box function that invokes algorithms developed for such problems (Snoek et al., 2012). Besides ANN, other remarkable ML algorithms are decision trees, random forest, support vector machine, k-nearest neighbors, among others (Maglogiannis, 2007). In the following sections, an overview of ANN is presented. The ANN fundamentals, components, structures, and mathematics behind this technique are described. The review includes the basic units of ANNs, the artificial neuron or perceptron, how the learning is achieved through this technique, to the validation of the models for posterior inferences, classifications, and predictions. 2.5.1 The artificial neuron Animals have complex brains composed of hundreds or thousands of millions of elements called neurons; the number varies according to species. In humans, the brain consists of approximately 1011 neurons – in this text animal, and consequently, human neurons will be referred to as biological neurons. These units are highly interconnected (approximately 104 synapses or connections per neuron), and an electromagnetic signal is passed through them, this signal is information that roves the brain, being processed by the neurons and derives into the diverse forms of actions or reactions (Zurada, 1992). The biological neuron is a highly specialized cell type. Morphologically three major regions can be defined (Figure 7): a cell body (or soma), which contains the nucleus; a variable number of dendrites, which emanate from the soma and Theoretical Background 37 ramify; and a single axon, which extends far from the soma and has numerous terminals to interconnect to other neurons dendrites (Squire et al., 2008). Figure 7. Biological neuron cell scheme. The signal travels from left to right, from the dendrites (input), passing through the body or soma and going out through the axon and its terminals. An artificial neuron is a simple mathematical model inspired in its biological counterpart. Artificial neurons produce an output value in two steps. First, the neuron computes a weighted sum of its signals, input variables for the case, and, in a second stage, applies an activation function to this sum to derive the product as the output (Russell and Norvig, 2016). In 1943, McCulloch and Pitts proposed the first artificial neuron model – today is known as the perceptron, composed of binary threshold activation function (Figure 8). This mathematical neuron computes a weighted sum of its input signals and generates an output of “1” if this sum is above a certain threshold, “0” or another positive number, otherwise, the function returns “0” as a result, this logic is also known as hardlim function (Jain et al., 1996). Dendrites Soma Axon Axon Terminals 38 Figure 8. McCulloch and Pitts artificial neuron. Where xn = inputs and wn = weights. The sum of products of inputs and weights plus the bias is passed to the activation function (σ) to generate the neuron output (𝑦). The perceptron applies a linear combination of its inputs, obtaining the signal in which is applied an activation function to obtain the output signal. Nonlinear functions are the most used ones to give the perceptron nonlinear behavior (Martínez-Martínez et al., 2015). The mathematical model of the perceptron is presented in Eq. 1: Were 𝑦 is the neuron output; 𝜎 is the activation function; 𝑥 is the input vector of 𝑛 elements; 𝑤 is the weight vector and 𝑏𝑖𝑎𝑠 is a value that allows the sift of the activation function. The bias is somehow similar to the constant 𝑏 of a linear function 𝑦=𝑎𝑥+𝑏. The biological neuron and the artificial neuron, the perceptron, are comparable in certain manners, a parallelism from an operational point of view can be described as follows: the weight (𝑤) corresponds to the strength of a synapse, the neuron body is represented by the summation and the activation function, and the neuron output (𝑦) is represented by the signal on the axon (Mindiola et al., 2015). In the perceptron, the analogous of the of electromagnetic signals in the processing of brain are, in deep, mathematical operations that take place in processor and random access memory (RAM) of a computer (Misra and Saha, 2010). 𝑦= 𝜎(∑𝑥𝑖𝑤𝑖+𝑏𝑖𝑎𝑠 𝑛 𝑖=1 ) Eq. 1 𝑥1 𝑥2 𝑥3 𝑥𝑛 𝑏𝑖𝑎𝑠 𝑤1 𝑤2 𝑤3 𝑤𝑛 𝑦 𝑆𝑢𝑚 Σ 𝐴𝑐𝑡𝑖𝑣𝑎𝑡𝑖𝑜𝑛 𝐹𝑢𝑛𝑐𝑡𝑖𝑜𝑛 σ Theoretical Background 39 2.5.2 The artificial neural network The perceptron by itself does not have a lot of processing power, especially when it is compared to an ANN, which is composed of several layers of perceptrons. ANN can be defined as a highly connected array of neurons, which are interconnected with each other; a widely used model called the multi-layered perceptron (MLP) is the most common type of ANN (Haykin, 1999; Park et al., 1991). The MLP consists of one input layer, one or more hidden layers, and one output layer. Each layer employs several neurons, and each neuron in a layer is connected to the neurons in the adjacent layer with different weights, which is used to determine how much one unit will affect the other (Chen et al., 2005). Neural networks are typically represented by a network diagram, which is composed of nodes connected by directed links. Nodes are arranged in layers, and the structure of the most used neural network consists of three layers: an input, a hidden and an output layer of nodes (Figure 9), signals flow into the input layer, pass through the hidden layers, and arrive at the output layer in a unidirectional path in the MLP (Hastie et al., 2009). ANN is a kind of array which can realize a nonlinear mapping from the inputs to the outputs (Rui and El-Keib, 1995). Thereby, the input neurons receive the data, and the inputs and weights products are computed; the signals enter the hidden layer where a sum is performed, a bias is added, and an activation function is applied in each neuron. Then, the signals leaving the neurons in the hidden layer, are multiplied again by weights and enter the neurons in the output layer, where a sum is performed plus the bias addition to generate the output of the ANN. Initially, the McCulloch and Pitts perceptron model presented limitations, and the nonlinear capability of ANN was not a reality until long after McCulloch and Pits pioneer work. It was not until 1962, 20 years after the first proposal, that Rosenblatt implemented hardware with enough computing power to process an array of artificial neurons and the term “perceptron” was coined. He proposed single-layer networks composed of a few processing units. These networks were applied to classification problems, in which the inputs were usually binary images of characters or simple shapes (Bishop, 1995). 40 Figure 9. ANN scheme. Where xn are the input variables that enter in the input layer and ŷn are the outputs from the output layer. The hidden layer contains the activation function. However, in a large number of interesting cases, the neurons were not capable of solving problems. A classic example is the “exclusive OR” (XOR) problem, which was reported by Minsky and Papert in 1969 (Özbay et al., 2007). The report resulted in a lack of enthusiasm and research in ANNs in the computer science community for almost 20 years (Basheer and Hajmeer, 2000). The recession was followed by regeneration of ANNs with the introduction of the ANN models of Hopfield during the 1980s, who popularized the MLP by employing the training algorithm of backpropagation proposed by Werbos in 1974 in his doctoral thesis at Harvard University, and the interest of the scientific community returned (Nastos et al., 2013). This interest reborn was coupled with the rapid growth of computing Input layer Hidden layer Output layer y2 y1 y3 y𝑛 1 2 3 4 5 n 𝑥1 𝑥2 𝑥3 𝑥4 𝑥n Theoretical Background 41 capabilities. Nowadays, the MLP can solve different nonlinear problems, including the XOR case with a quite simple network structure (Singh, 2016). The task of an MLP is to get the desired output to the inputs of certain types. With this goal in mind, the learning of perceptron is performed. The inputs and the desired outputs are loaded to the scheme, and the error of the response is determined. The parameters of the system, the weights in the inter-connections, are changed during the learning phase of the ANN to diminish the difference between the desired and the real output (Negrov et al., 2017). The neural networks can be classified into two classes according to its architecture and interconnection between neurons: feed-forward networks and feed-back (recurrent) networks (Tealab et al., 2017). The most popular ANN in the studies is the multilayer feed-forward network, in which the neurons are organized into series of layers, and information signal flows through the network solely in one direction, from the input layer to the output layer (Tkáč and Verner, 2016). Meanwhile, in an ANN where the outputs of some neurons are feed back to the same neurons or neurons in preceding layers are called recurrent neural networks (RNN). This feed-back enables a flow of information in both forward and backward directions (Basheer and Hajmeer, 2000). This type of neural network is more focused in time series predictions thanks to using as the input data part of the output data from the previous element in the series, such as forecast financial time series (Cavalcante et al., 2016; Li et al., 2018). While MLP is more focused on solving nonlinearly separable problems, classification, and approximate continuous functions (Haykin, 1999), there is another architecture of neural networks based on the number of hidden layers; when the ANN consists of several of perceptrons layers, the term deep neural network (DNN) is employed (Li et al., 2018). These types of neural networks are employed by excellence for image analysis and classifications, including videos (Antipov et al., 2016; Cheng et al., 2018). However, for most of the classification and predictions problems, there is no need to use more than one hidden layer (Heaton, 2008), in some cases augmenting the number of hidden layers can lead to the detriment of performance (de Villiers and Barnard, 1993). In an MLP, adding hidden layers and neurons are generally limited for a few hidden layers. Adding more layers not only increases the complexity of the model and the computational cost, but it also does not ensure more accuracy or better overall performance of the model. Experimental results indicate that ANNs with two hidden layers are prone to perform more inaccurately in comparison to 42 one hidden layer ANNs (de Villiers and Barnard, 1993). Increasing the number of hidden neurons, in forms of a neuron inside the hidden layer or increasing the raw number of hidden layers itself, can dismiss the learning process leading to less accurate models (Karlik, 2011). ANN, especially in complex models with several neurons and layers, have the cost of the training time (Moraes et al., 2013). Although new central processing units (CPUs) are faster than ever, have multicore structures (2 or 4 cores in general) and instructions for high-performance parallel processing (Bergstra et al., 2010; Vanhoucke et al., 2011). Moreover, graphics processing units (GPUs) can be used for faster ANN implementations, due to higher capacity of parallel computing by the multicore system with hundreds of cores (Oh and Jung, 2004; Yang et al., 2011). GPU implementation of ANN can go up to 50 or 60 times faster than standard CPU implementations (Cires et al., 2003; Ramachandran et al., 2015). Regarding ANN frameworks, there are free and open source libraries such as Caffe, Theano, Torch, and TensorFlow, among others, which are widely spread (Abadi et al., 2016; Rampasek and Goldenberg, 2016). The importance of being free and open source lies in the affordability for use in academics and research. TensorFlow, one of the frameworks that most interest has caught in developers in the last few years, is developed by the Google Brain Group, part of the Google Machine Intelligence Research Institute (Qin et al., 2019), and the use of the framework is increasing research since it was released in 2015 (Abadi et al., 2016; Cabañas et al., 2019; Hazan et al., 2018; Kulkarni et al., 2018; Vázquez-Canteli et al., 2019; Zhang and Kagen, 2017). Several ANN studies have been conducted in distinct disciplines with overwhelming results; in the photovoltaic energy industry to forecast the power generation of system under diverse climatic scenarios (Bouselham et al., 2017; Ding et al., 2011; Veerachary and Yadaiah, 2000); in the industry sector to control and monitoring of production process, such as distillation (Singh et al., 2007), lifetime prediction of machinery (Tian, 2012), fault diagnosis in elements of a power installation (Bi et al., 2000) and controlling robotics machinery (Zhao et al., 2014). In medicine, ANNs are useful for the diagnosis of cardiovascular diseases, diabetes, cancer and tumors (Jiang et al., 2010; Singh et al., 2015). In the finance sector, to schedule the energy loads for a more economical consumption and better energy prices (Alanis, 2018; Chen et al., 2001), bankruptcy prediction (Li et al., 2018; Tkáč and Verner, 2016; Zhang et al., 1999) and prices and time series prediction (Abhishek et al., 2012; Göçken et al., 2016; Kaastra and Boyd, 1996; Moghaddam et al., 2016; Theoretical Background 43 Patel et al., 2015); in the telecommunication sector, neural networks can treat and recover digital information in channels when a data loss occurs (Das et al., 2014; Panda et al., 2015), to cite some of the multiple examples of ANN applications. 2.5.3 Activation functions Activations functions are an essential component of ANNs. In a perceptron, the weighted inputs are summed and passed through a limiting function, which scales the output to a fixed range of values. The output of the limiter is then broadcasted out of the neuron to the next layer (Ozturk and Karaboga, 2011). In general, S-shaped functions, such as sigmoid and hyperbolic tangent, are adopted for the activation. Once the activation function is performed, a neuron sends its activated value to the other neurons through the connections (Tsai and Wang, 2001; Zamanlooy and Mirhassani, 2014). The activation function is typically nonlinear, and it ensures that the entire network can estimate a nonlinear function which is learned from the input/output data pair (Moraes et al., 2013). Nonlinear activation function implementation achieve higher accuracy and improves the learning and generalization capabilities of ANNs in contrast to the linear primitive functions (Zamanlooy and Mirhassani, 2014). Usually, neurons in the hidden layer are the ones that perform an activation function; output neurons, in general terms, return a weighted summation of the previous layer output without any transformation. It has been reported that a nonlinear activation function in the output layer failed to improve the performance and accuracy of an ANN (Yonaba et al., 2010). There are several activation functions, with a diverse range of bounded outputs; according to the problem to solve, one or other activation function performs better (Gautam and Ravi, 2015). Some of the main activation functions are described in the subsections below. 2.5.3.1 Hardlim function The hardlim function, also known as a binary step function, was the first activation function presented by McCulloch and Pitts, fathers of the artificial neuron. The function gives binary output values, “0” or “1”, “0” when the signal 44 value is less than “0” and “1” when the signal value is equal or greater than “0”. The formula is presented in Eq. 2. 𝑓(𝑥)=ℎ𝑎𝑟𝑑𝑙𝑖𝑚(𝑥)= { 0 for 𝑥<0 1 for 𝑥≥0 Eq. 2 The function plot is shown in Figure 10. The hardlim function presents two possible stages according to the input value into the activation function, thus the name of the binary step. Regardless of how negative or positive the input value is, the logic of this function will always return “0” or “1”. Figure 10. Hardlim activation function plot. Hardlim is rarely used nowadays since there is no linear relationship between inputs and output patterns in most modern problems of interest (Abhishek et al., 2012). Hardlim activation function was not very effective to solve several problems in comparison with the S-shaped functions in a comparative study using several datasets from the University of California Irvine machine learning repository (Zhang and Suganthan, 2016). - 1.2 - 1.0 - 0.8 - 0.6 - 0.4 - 0.4 0.6 0.8 1.0 1.2 -7 -6 -5 -4 -3 -2 -1 1 2 3 4 5 6 7 Hardlim Theoretical Background 45 2.5.3.2 Sigmoid function The sigmoid activation function is one of the most used in ANNs, and it rapidly replaces the hardlim function. There are some variations of the sigmoid functions, but the one known as logistic sigmoid is often used in studies and applications (Menon et al., 1996). The sigmoid activation function produces positive numbers only between “0” and “1”, and the formula is presented in Eq. 3. This function is also known as the S-shaped function, and its bounded output values make it one of the most useful in training ANNs (Sibi et al., 2013). 𝑓(𝑥)=𝑠𝑖𝑔𝑚𝑜𝑖𝑑(𝑥)= 1 1 + 𝑒−𝑥 Eq. 3 ANN with sigmoid activation functions was used to solve several issues such as inference of rivers water quality (Palani et al., 2008); electricity power demand prediction (Manohar and Reddy, 2008); and solar air-heater modeling (Ghritlahre and Prasad, 2017), among others. The typical S-shaped graph of this function is shown in Figure 11. Figure 11. Sigmoid activation function plot. - 1.2 - 1.0 - 0.8 - 0.6 - 0.4 - 0.4 0.6 0.8 1.0 1.2 -7 -6 -5 -4 -3 -2 -1 1 2 3 4 5 6 7 Sigmoid 52 The training set is the largest dataset, and it is used by the neural network to learn the patterns present in the data. The testing set, ranging in size from 10% to 30% of the training set, is used to evaluate the generalization ability of a supposedly trained network (Kaastra and Boyd, 1996). The data in the test set is used to supervise the training process and avoid memorization or “overfitting”, in other terms, giving good outcomes only for trained data (Dündar and Şahin, 2013). Hence, the training phase can be stopped at the point where the network has the smallest testing error; if the testing error starts to increase while the training error decreases, overfitting is taking place (Zhang and Morris, 1998). The “best” network on training and testing data is the network that best passes the learning phase. However, its performance is finally qualified in the validation data (Zhang et al., 2003). In other words, ANN must be trained, tested, and subsequently validated (Belgrano et al., 2002; Patanaik et al., 2018). Several studies use the test set as a validation set for the final performance assessment of the model, not in the testing phase rightly said (Khashman, 2010; Trejo-Perea et al., 2009; Yang et al., 2009). Other studies do not use all datasets or the test or the validation dataset (Drucker et al., 1993; Kong et al., 2016; Leu et al., 2001). This confusion may be due that ANN is a particular technique that needs these three datasets, for train, test and validate the model, meanwhile, for other predictions methods, just two datasets are usually demanded, a train and a test set (James et al., 2013). In the training phase, the mathematics behind the abstraction capability of ANNs is called the learning algorithm. The most widely used learning algorithm is the backpropagation gradient descent, its popularity is presumably caused by simplicity, universality, and good availability in libraries (Phansalkar and Sastry, 1994; Tkáč and Verner, 2016). There are other learning algorithms such as genetic algorithms and particle swarm algorithms, these algorithms are not as widely implemented by machine learning libraries as the gradient descent but still an alternative method to study (Rios and Sahinidis, 2013). The purpose of the backpropagation training is to iteratively change some parameters of the neurons in a direction that minimizes the error, or as it is defined, the error function ― which basically performs the difference between the desired output and the actual outcome of the ANN across all the training and testing patterns (Örkcü and Bal, 2011). The ANN has unknown parameters, weights, and bias, just as the perceptron does, and these parameters are one of the responsible for the abstraction capability of the model. The objective of the learning phase is to Theoretical Background 53 seek values for the unknown parameters of the weights and bias to make the model fit the training data well. Usually, starting values for weights are chosen to be random values near zero, for instance from “-1” to “1”. Hence the model starts nearly linear and becomes nonlinear as the weights change (Cao et al., 2015; Hastie et al., 2009). The backpropagation procedure computes the gradient of an objective function, the error function, concerning the weights and bias of a multilayer stack of modules. The backpropagation is nothing more than a practical application of the chain rule for derivatives, and it can be defined by Eq. 10: 𝜕𝐸 𝜕𝑤𝑖𝑗= 𝜕𝐸 𝜕𝑦𝑖𝑗 × 𝜕𝑦𝑖𝑗 𝜕𝑧𝑖𝑗 × 𝜕𝑧𝑖𝑗 𝜕𝑤𝑖𝑗 Eq. 10 Where: 𝐸 is the error or cost function. 𝑤𝑖𝑗 is a certain weight of a certain synapsis. 𝑦𝑖𝑗 is a certain output of a neuron. 𝑧𝑖𝑗 is the product of the weight and a input variable. The chain rule is applied in functional dependence relationship ― when a function depends on other function. For instance, the error function (𝐸) depends on the output of a neuron (𝑦), that at the same time depends on the product of the input variables (𝑧), making the link to the final parameter of this chain the weights and bias (𝑤). The above equation describes the error gradient respect a specific weight correction, but it is also valid for a particular bias value since they are adjusted in the training phase in the same way. For the backpropagation, all the individual weights, bias, and errors are computed and added to obtain the cost function gradient of the ANN and adjust the parameters according to the calculus (Rosenbaum and Johnson, 1984). The key insight is that the derivative or gradient of the error function of ANN can be computed by working backward from the gradient concerning the output layer to subsequent layers. The output of the initial network model and corresponding error is computed, then, at each iteration, new weights and biases are determined to add or to subtract a value to them, this value is called learning rate. Again, new outputs and errors are determined, and if the new error is above the previous one, the new weights and biases are rejected, and the fixed value is again added or subtracted according to the gradient (Rusk, 2015; Singh et al., 2015). 54 Training the network involves adjusting the connection weights to correctly map the training set to obtain a desirable output, at least to within some defined error limit, it is basically an optimization problem (Palani et al., 2008). In effect, the network learns from the training set; if the training set is satisfactory and the training algorithm is effective, the network should then be able to correctly estimate the output even for the inputs not belonging to the training set ― this phenomenon is termed as “generalization” (Bishop, 1995; Punitha et al., 2013; Wang et al., 2006). The training phase is a critical part of the use of neural networks. For a given problem, the network training is supervised, given the fact that the target for each input pattern is always known a priori. During the training process, the patterns or examples are presented to the input layer of a network (Zhang et al., 1999). With the gradient descent backpropagation algorithm, a gradient descent search is performed ― it measures the output error and calculates the gradient of the error by adjusting the weights in the descending gradient direction (Behrang et al., 2010; Sharma and K. Venugopalan, 2014). The randomly generated weights and the bias are adjusted by a magnitude called learning rate. The effectiveness and convergence of the training algorithm depend significantly on the value of the learning step, the optimum value of the learning step is system-dependent and varies according to the problem and other ANNs related features. For systems that possess broad minima, a large learning rate value will result in a more rapid convergence. Meanwhile, in a system with a narrow minimum, a small learning rate value is more suitable. There are no general rules to obtain an optimal learning step; typically used values are 0.9, 0.25, and 0.05 (Rui and El-Keib, 1995). Randomly initialized weights achieve much faster learning speeds; that is why the weights and bias are usually randomly generated by the computer (Cao et al., 2018). The backpropagation gradient descent is the rule by which the weights are modified, the reason why it is called the stochastic descent method is that the modifications are performed in an “a priori unknown” sets of weights and bias (Amari, 1993). The backpropagation algorithm is one of the most powerful supervised learning algorithms but if it is not well employed, the ANN will not converge to the minimum error point or will fall into a local minimum, not the global minimum (Figure 15) and the convergence speed can be reduced (Bi et al., 2005; Ding et al., 2011; Örkcü and Bal, 2011). Theoretical Background 55 Other strategies to ensure better performance of the training algorithms are the stopping criteria and the dataset number. The stopping criteria can involve stopping after a certain amount of runs through all of the training data and also stopping when the total target error reaches some low level, for example, less than 2% during the learning phase (Palani et al., 2008). Monitoring the error in the train and the test sets is also a good strategy for early stopping; if the testing error starts to increase in comparison to the training error, early stopping the training prevent the overfitting (Tetko et al., 1995; Zhang and Morris, 1998). Figure 15. Example of the error function with respect to the weight gradient during the backpropagation algorithm execution. Regarding the number of samples in the datasets, generally speaking, a large number of input samples is necessary to avoid overfitting (Ramachandran et al., 2015; Wang et al., 2006; Yosinski et al., 2014). There is not a recommended fixed number of training samples, there is a direct relation between large datasets and proper learning levels of ANNs (Kourou et al., 2015; Krizhevsky et al., 2012), a good indicator can be that the number of training samples should be larger than the number of hidden neurons or neurons in the hidden layers (Huang et al., 2004). Performing proper training is a key for the ANN; avoiding the overfitting of the network is crucial for that purpose. Overfitting affects the generalization capability of ANN, and thus affects the prediction accuracy (Tian, 2012), and in the other hand, early stopping is better in terms of performance-to-cost ratio than Error Weight gradient current (𝑤𝑖𝑗) ideal (𝑤𝑖𝑗) local minimum global minimum 56 conventional stop mechanisms (Iyer and Rhinehart, 2000). The number of samples in the dataset, early stopping by reaching a certain number of iterations, error tolerance on the training set, and monitoring the testing error should be considered together to overcome the overfitting problem and not in isolation (Chen et al., 2005). 2.5.4.3 Validation phase After the training phase has been finalized, ANN performance is evaluated over the validation set (Ahmadi, 2011; Altinay et al., 1997; Jafar et al., 2010). The validation set must be composed by samples not included in the learning phase, i.e., into the train and test sets, for appropriate model evaluation (Basma and Kallas, 2004; Tian, 2012). In terms of the proportion of data needed for each dataset, the division into three parts follows the following relation: 70% for the learning process, 15% for test phase and 15% for validation phase (Bal and Buyle-Bodin, 2013), or 50%, 25%, 25% respectively, can be also an alternative (El Tabach et al., 2007) as much as other proportions in similar ranges. The lack of fit between the observed phenomena, the real values, and estimated values by the ANN during the validation phase indicates that the model should be re-trained, with the proper learning procedure and probably with more massive datasets to ensure more accurate results (Palani et al., 2008; Yang et al., 2009). The cost function, the error algorithm, is used to measure the model accuracy, and it is one crucial algorithm to the network development. In ANNs, the most used cost functions are mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE), and cross-entropy (CE), among others. These functions can be used during the learning and the validation phases, meanwhile, functions such as correlation coefficient (R) and determination coefficient (R2) are most appropriate exclusively for the validation phase (Ali et al., 2017; Ay and Kişi, 2017; Cheng et al., 2018; Moghaddam et al., 2016; Sarkar and Pandey, 2015). The MSE algorithm (Eq. 11) is generally selected because of its excellent performance, relative to the backpropagation algorithm (Zanetti et al., 2007) and it is one the most popular cost function in ANNs (Adya and Collopy, 1998). MSE function appeals to be adequate when dealing with a single variable in a study, since Theoretical Background 57 using multiple variables and dimensions for each one make the error comparison more difficult (Köksoy, 2006). 𝑀𝑆𝐸=1 𝑛∑(𝑦𝑖−𝑦𝑖)2 𝑛 𝑖=1 Eq. 11 Where 𝑛 is the number of samples in the dataset; 𝑦𝑖 is the real data; and 𝑦𝑖 is the estimation of the ANN. The same symbology will be applied to the remaining equations of this section. The RMSE is calculated easily, in the same manner than MSE, and then taking the square root of the same (Eq. 12). The RMSE hence summarizes the overall error of the model, i.e., the precision of the model (Aptula et al., 2005); the main difference between MSE and RMSE is that the RMSE is most useful when large errors are particularly undesirable (Saigal and Mehrotra, 2012) and the small errors tend to be less penalized. 𝑅𝑀𝑆𝐸=√∑(𝑦𝑖−𝑦𝑖)2 𝑛 𝑖=1 𝑛 Eq. 12 The MAE, shown in Eq. 13, is an intuitively appealing measure. The MAE penalizes under and over-prediction respecting to the actual outcome and is the most natural measure of average error magnitude, is an unambiguous measure of average error magnitude and useful for inter-comparisons of model performance with different magnitudes and measure units (Willmott and Matsuura, 2005). While the MAE gives the same weight to all errors, the RMSE penalizes variance ― as it provides errors with larger absolute values more overall weight than errors with smaller absolute values. When both metrics are calculated, the RMSE is, by definition, never lower than the MAE (Chai and Draxler, 2014). 𝑀𝐴𝐸= 1 𝑛∑|𝑦𝑖−𝑦𝑖| 𝑛 𝑖=1 Eq. 13 A modification of the MAE algorithm is the MAPE. This algorithm (Eq. 14) is computed through a term-by-term comparison of the relative error in the prediction concerning the actual value of the variable. Consequently, the MAPE is an unbiased statistic to measure the predictive capability of models (Wang et al., 2009). MAPE is often used because of its intuitive interpretation in terms of relative 58 error, but just viable when the quantity to predict is above zero (de Myttenaere et al., 2016); otherwise, a zero division is produced, and non-result can derive from this formula. 𝑀𝐴𝑃𝐸= 1 𝑛∑|𝑦𝑖−𝑦𝑖| 𝑦𝑖×100 𝑛 𝑖=1 Eq. 14 Regarding the CE function, this error method reflects the similarity between variables from the perspective of probability (Men et al., 2016). This cost function (Eq. 15) is frequently with ANN with softmax activation function in the output layer (Dahl et al., 2013; Liew et al., 2016; Maas et al., 2013). CE has significant advantages over other error functions for probabilities problems, performs better on large and small target values because they tend to result in similar relative errors for both cases. The CE function performs better at estimating low posterior probabilities than squared-error functions (Kline and Berardi, 2005). The output of the softmax layer returns a series of probabilities, zero or positive values, and the CE penalizes all the results which do not match the current probability of the target. 𝐶𝐸= −∑𝑦𝑖×log𝑦𝑖 𝑛 𝑖=1 Eq. 15 The cost functions reviewed so far are eligible for both the learning and the validation phases. Nonetheless, some algorithms are exclusively used for validation of the model, such as the R and R2. The R, as is shown in Eq. 16, investigates the degree of association between two variables, that is, it defines how much a given relationship is fitted by a straight line (Tripepi et al., 2008). 𝑅= 𝑛(∑𝑦𝑖𝑦𝑖)−(∑𝑦𝑖)(∑𝑦𝑖) √[𝑛∑𝑦𝑖2−(∑𝑦𝑖)2][𝑛∑𝑦𝑖2−(∑𝑦𝑖)2] Eq. 16 The R2 is an algorithm (Eq. 17) that is widely used in ANNs to assess the performance of models (Gholami and Fakhari, 2017), including approaches in regression, classification and predictions problems (Ghritlahre and Prasad, 2017; Moghaddam et al., 2016; Palani et al., 2008; Tetko et al., 1995; Turan et al., 2011). Theoretical Background 59 𝑅2= 1 − ∑ [𝑦𝑖−𝑦𝑖]2 𝑛 𝑖=1 ∑ [𝑦𝑖−𝑦𝑚𝑒𝑎𝑛]2 𝑛 𝑖=1 Eq. 17 The cost function and the model evaluation are crucial tasks in the ANN assays. There is not a universal solution for better performance, and most of the work in creating good quality models is to test not only different network architectures, but also cost functions, and the performance in the validation samples. Once that several models are evaluated, it can be settled that the selected ANN is the most pertinent for solving the problem. MATERIALS & METHODS 68 3.2 Artificial neural network for soil color analysis and characterization In this section, the soil characterization and color analysis by ANNs methodology will be detailed. For the experimental work, soil samples and digital color photographs were analyzed through ANNs to find possible relations between color and nutritional contents. The main objective was to develop a tool for soil color classification and analysis in order to develop a fast and inexpensive tool for soil management in agricultural lands. A scheme of the experiments is shown in Figure 20. Figure 20. Soil color analysis and characterization by ANN experiment scheme. The remaining subsections are organized as follows: study area and soil sample methodology (section 3.2.1), image acquisition system for soil sample photographs description (section 3.2.2), the CV software for soil color processing and RGB to L*a*b* conversion (section 3.2.3), Munsell soil color classification and statistical analysis (section 3.2.4), and the ANNs for the soil color characterization (section 3.2.5). RGB L*ab N P K N P K Fertility Characterization Soil Photography ANN Color Processing Munsell Classification Statistical Analysis a ab b Materials and Methods 69 3.2.1 Study area and soil sampling procedure The study area consisted in a 43.01 ha vineyard located between 40⁰00'55.25'' and 40⁰01'24.99'' N to 2⁰56'55.05'' and 2⁰55'34.83'' W in Tarancón, province of Cuenca (Castilla La Mancha, Spain). Previous mapping of apparent electrical conductivity (ECa) was performed for a representative soil sampling. The ECa was measured with a Veris Q2800 at depths of 36 and 90 cm in consecutive parallel paths with 10 m of separation between lines throughout the parcel, the ECa tend to be useful to sort information and identify sets of attributes with similar trends (Officer et al., 2004). The obtained information, one raster file for each depth (Figure 21 a and b), were reclassified using the Jenks natural breaks classification method (Chen et al., 2013) ― 3 classes were defined in each raster. The intersection of the classes in both depths was used to create new segments, and each soil sample was extracted from these new segments (Figure 21c). Hence, 174 sampling points were defined, 87 for each depth in order to get samples of each homogeneous segment of the parcel. The soil samples were analyzed in the laboratory for 13 parameters, such as the OM and total nitrogen (N), expressed in percentage, are measured using combustion autoanalyzer LECO TruSpec; phosphorous (P), potassium (K), calcium (Ca), magnesium (Mg), iron (Fe), copper (Cu), manganese (Mn), sodium (Na) and aluminum (Al), are measured by ICP-OES after acid digestion in nitric acid using microwave assisted digestion, and are expressed in mg/kg of soil; pH and electrical conductivity (EC) were measured in water (1:5 w/v) and are in dS/m. 70 Figure 21. ECa raster map from the Veris data acquisition system at 36 cm (a) and 90 cm (b) depth; Segments for soil sampling generated through the processed raster data (c). 0 200m a) b) c) ECa 0 100 Materials and Methods 71 3.2.2 Image acquisition system Soil digital photos were taken with an image acquisition system consisted of a Panasonic DMC-TZ70EG-S compact digital camera, with 12.1 Mpx, CMOS sensor and Leica DC Vario Elmar 24 mm lens. The illumination set was arranged by a pair of 6,500 K professional photoshoot light sources to provide uniform illumination to the sample (Mendoza et al., 2006) (Figure 22a). The image acquisition system in controlled conditions is the most appropriate manner to take digital photographs, under non-controlled conditions the color can result less accurate and less appropriate for this experiment (Lapins et al., 2013; Potočnik et al., 2015). The camera was mounted on a stand and remotely operated via the “Image App” smartphone application (Figure 22b). The camera aperture speed was set at 1/30s, ISO-80, luminosity f/3.3, and focal distance 4 mm. Soil samples were placed in Petri plates of 55 mm diameter, uniformly filled; 5 repetitions (plates) per soil sample were prepared, a total of 870 experimental units were photographed. Soil samples were dried at ambient temperature to avoid the interference of moisture content in the soil color expression (Sánchez-Marañón et al., 2007). Figure 22. Image acquisition system; light and camera setup (a) and shooter camera app for smartphone (b). a b 72 3.2.3 Computer vision software for color conversion In order to accelerate the process of color measurement of soil samples, a CV application has been programmed to analyze the color in digital photographs and perform the conversion from RGB system to the international standard CIELAB color system. The software and user interface were developed in NI LabVIEW 2018 using the Vision Module. This programming environment has the versatility to work with scripts from other programming languages, such as Python, and it is able to execute ANN’s scripts with the advantage that it allows designing a user interface which enables the user to use an ANN without writing a single line of code. The soil color conversion from RGB to L*a*b* was performed using an ANN following the methodology proposed (León et al., 2006). To generate the model, the “X-Rite ColorChecker Classic” (Figure 23) was used. This colorchecker includes 24 color charts with natural, chromatic, and primaries colors in addition to a greyscale. The color charts are scientifically prepared and verified, see Appendix C for technical specs of the colorchecker. The RGB values from color chambers in digital color photographs were extracted and used as input in the ANN, and the L*a*b* values provided by the manufacturer were the output in the training and testing phases. Figure 23. X-Rite ColorChecker Classic, color calibration card. The ANN was composed of 3 neurons in the input layer, one for each RGB color element; 8 neurons in the hidden layer with the hyperbolic tangent activation Materials and Methods 73 function; and 3 neurons in the output layer, one for each L*a*b* color element ― a schema is shown in Figure 24. The early stop was applied to prevent overfitting. The ANN was programmed in Python 3.6.5 using TensorFlow 1.8 machine-learning library. The obtained model was applied to convert the soil RGB dataset into L*a*b* color. Figure 24. ANN architecture schema, with 3 neurons in the input layer, 8 in the hidden layer and 3 in the output layer. 𝐿∗ 𝑎∗ 𝑏∗ 𝑅 𝐺 𝐵 Input layer Hidden layer Output layer 1 2 3 4 5 7 8 6 74 3.2.4 Munsell soil color classification and description To perform the Munsell color soil classification, a Munsell soil color charts book was photographed in the same image acquisition system described in section 3.2.2. A CV software was programmed in LabVIEW 2018 using the Vision Module to process the color Chart Book and extract RGB median values of each notation color in the book. The RGB mean values of the 5 repetitions of each soil sample were computed to compare them with the RGB obtained from each Munsell color notation. The minimum Euclidean distance match algorithm was performed to compare the colors, as is shown in Eq. 19; where 𝑅𝑠, 𝐺𝑠 and 𝐵𝑠 are the mean of the RGB color values of the soil samples and 𝑅𝑀, 𝐺𝑀 and 𝐵𝑀 are soil chart RGB color of each Munsell notation. Each mean of the RGB values from the soil samples was compared with the RGB values from the Munsell colors, and the notation with the absolute minimum difference was taken as the soil Munsell color of the sample (Stanco et al., 2011). 𝑑=√(𝑅𝑠−𝑅𝑀)2+(𝐺𝑠−𝐺𝑀)2+(𝐵𝑠−𝐵𝑀)2 Eq. 19 Before using the dataset, outlier points were removed based on fertility parameters, the soil samples classified according to their Munsell color notation were analyzed statistically; the Kolmogorov–Smirnov and Levene tests were performed, data normality and equality of variances were discarded according to the tests. The means were compared with the Kruskal-Wallis H test to verify differences between groups and the Mann-Whitney U test were used to compare the groups based on their Munsell hue color classification; all statistical tests were performed in R programing language version 3.4.4. 3.2.5 Soil properties characterization by artificial neural networks For the ANN analysis, two datasets were prepared with the soil color values, one for the RGB values and one for the L*a*b* values, and the laboratory analysis for soil fertility parameters. Each dataset was used to train, test and validate ANNs. The data were normalized to improve the network performance (Gnana Sheela and Deepa, 2013), the Z-score normalization was applied to the datasets as Materials and Methods 75 is shown in Eq. 8; where 𝑥𝑖 is the datum of the variable, 𝜇 is the mean of the variable, 𝜎 is its standard deviation and 𝑧𝑖 is the normalized datum. The ANNs were fully connected feed-forward neural networks, with the RMSE cost function, programmed in Python 3.6.5 language using the TensorFlow 1.8.0 machine learning library. After various configuration attempts, the most straightforward structure ― with less hidden layers and fewer neurons per hidden layers ― and the activation function with the least training and testing errors was selected. The ANNs structures were set with 3 neurons in the input layer, one neuron for each color component; 1 hidden layer with 18 neurons, with hyperbolic tangent activation function; 13 neurons in the output layer, one neuron for each soil property to predict. For training and testing phases, 554 and 62 experimental units were used, respectively. The validation phase was carried out with 25 soil samples not previously shown to the ANN in training or testing phases, the mean color of the 5 repetitions of each soil sample was used to make the 25 data to the validation dataset. 76 3.3 Virtual weather stations for meteorological data estimations In this section, the methodology to develop the VWS is described. This development was planned with the primary objective of evaluate and compare the accuracy of several interpolation algorithms, ANN approaches and the most used interpolation methods, to generate meteorological data and develop the VWS as an option to acquiring accurate data economically and straightforwardly. A scheme of the procedure is shown in Figure 25. Figure 25. VWS development methodology scheme. The remaining subsections of the VWS methodology are organized as follows: the meteorological data source (section 3.3.1), statistical analysis of the dataset (section 3.3.2), the development of the VWS and validation procedure (section 3.3.3), and the description of the interpolations algorithms employed (section 3.3.4). Weather Station Network Virtual Weather Station (VWS) Data collecting and Processing Accuracy measurement VWS Weather Station Estimated data Real data Data for model validation only VWS R2 and RMSE Materials and Methods 77 3.3.1 Meteorological dataset The data was obtained from the InfoRiego metrological station network records of the Agrarian Technological Institute of Castilla y León (Spain), the locations are represented in Figure 26. The meteorological records were downloaded from the file transfer protocol (FTP) server available; a Python script automates the download of the data by dates. The registers consisted of daily data summaries of 53 meteorological stations distributed throughout the territory of the autonomous community of Castilla y León. The records date from July 1st, 2017 to June 30th, 2018 to compile information for a whole year. Figure 26. Meteorological station network of the InfoRiego program of the Agrarian Technology Institute of Castilla y León. The meteorological observation records were grouped into a set of daily precipitation (Precip) in mm of rainfall, reference evapotranspiration (ETo) in mm of water, daily mean air temperature (Mean Temp), maximum registered temperature (Max. Temp) and minimum registered temperature (Min. Temp) in ⁰C, daily mean air relative humidity (Mean RH), maximum registered air relative humidity (Max. RH) 84 observe how the lines overlap each other, thus making the spectral signature of the different microalgae used more comparable (Figure 27 right). The results show that each microalga species had a unique spectral signature: S. almeriensis and C. vulgaris had similar spectral signatures while Nostoc sp. and S. platensis had particular shapes that allow distinguishing one from the other. In the case of Nostoc sp., slight variations in spectral signature were observed over time, especially on days 4 and 5. This variation, especially in the chlorophyll absorbance region, was not due to changes in environmental conditions because these were kept constant under laboratory conditions so that it might be related to changes in the cells’ biochemical composition. Nutrient supply and light conditions are two significant factors that modify the pigment content of a microalga species. Thus, the pigment content is related to the physiological status of the cells (Pancha et al., 2015). Furthermore, the growthcycle phase also modifies the pigment content (Lubián et al., 2000). Other pigment content variations are related to changes in environmental and operational conditions (Pancha et al., 2014; Solovchenko et al., 2013). Despite the difference in days between sample measurements, other factors remained constant, and the absorption spectrum remained uniform. Under different culture conditions, the spectral signature could present variations within the same microalgae species across the measurements. The variation in a microalga species’ spectral signature over time was minor compared to the inter-species signature variation. Similar light absorption spectra, with peaks at 450 nm and 680 nm, have been obtained for other microalgae, such as Chlamydomonas reinhardtii (Isono et al., 2015), Thalassiosira pseudonana (Hewes, 2016) and C. vulgaris (Myers et al., 2013). Absorbance peaks are caused by photosynthetic pigments. Thus, while chlorophylls present two distinct absorption maxima, one between 400 and 500 nm and the other between 600 and 700 nm, the maximum absorption for carotenoids can increase above 500 nm due to spectral shifts caused by different contributing pigments (Holzinger et al., 2016). Results and Discussion 85 Figure 27. Light absorption spectrum of monoalgal cultures as obtained from the colorimeter (left), and relative absorbance conversion (right). 86 4.1.2 Artificial neural network culture discrimination Data from 550 samples were measured to train and test the ANN. Once the training and testing phases were concluded, the obtained model was used to determine whether the validation samples were monoalgal or mixed algal cultures. The number of samples met the requirement of having more training samples than hidden neuron units (Huang et al., 2004). The final ACE during the training and testing phases were 0.520 and 0.487, respectively. For the validation phase, the ACE was 0.527. The comparison between the experimental data and the ANN output is shown in Table 1. Monoalgal cultures (S101 to S104) from the first validation experiment day were correctly identified by the ANN, providing results above 98% purity. Regarding the mixed algal samples, S105 and S106 correspond to not-previously-performed combinations; thus, they are “unknown combinations” for the model. Despite this, the model was capable of approximating the composition of these samples. Therefore, these samples were close to the monoalgal cultures and the model was able to identify the most abundant species and its relative proportion. The ANN considered that small amounts of other species were also present in these samples; nevertheless, it was always capable of identifying the most prevalent species and percentage composition. Regarding the samples with 50% of two different species (S107 and S108), the artificial neural network was, likewise, capable of approximating the composition of these mixed algal samples, identifying the two predominant species and their approximate percentage. In no case, these samples were identified as monoalgal cultures. The same validation protocol was repeated on the second day. In this case, when using monoalgal cultures (S201 to S204), the ANN provided results above 99% purity for these samples. For samples containing 90-95% of one species and 10-5% of another (S205 to S209), these samples were “unknown combinations” for the model. Nonetheless, it was able to identify the predominant algal species present in each sample, with percentages similar to monoalgal cultures. The one exception was the 90% S. platensis and 10% C. vulgaris sample, for which an accurate estimate was given. Similar mixed algal samples to those used for the training process (S210 to S227) provided better results, with none of them being classified as monoalgal. The last sample (S228) also corresponded to an “unknown combination” for the model, and it was, likewise, not classified as monoalgal. Results and Discussion 87 Table 1. Comparison of the results from the experimental and the artificial neural network output of the validation experiment. Day Sample Experimental Model Nostoc sp. S. almeriensis S. platensis C. vulgaris Nostoc sp. S. almeriensis S. platensis C. vulgaris 1 S101 100.00% 0.00% 0.00% 0.00% 100.00% 0.00% 0.00% 0.00% S102 0.00% 100.00% 0.00% 0.00% 0.00% 98.70% 0.00% 1.30% S103 0.00% 0.00% 100.00% 0.00% 0.02% 0.00% 99.98% 0.00% S104 0.00% 0.00% 0.00% 100.00% 0.00% 0.14% 0.00% 99.86% S105 95.00% 5.00% 0.00% 0.00% 99.93% 0.05% 0.00% 0.02% S106 90.00% 10.00% 0.00% 0.00% 98.71% 1.19% 0.00% 0.09% S107 50.00% 0.00% 0.00% 50.00% 60.16% 0.91% 0.00% 38.93% S108 0.00% 50.00% 50.00% 0.00% 0.01% 30.01% 69.36% 0.62% S109 33.33% 33.33% 33.33% 0.00% 47.63% 36.42% 15.29% 0.66% S110 12.50% 50.00% 37.50% 0.00% 5.76% 42.75% 16.39% 35.10% 2 S201 100.00% 0.00% 0.00% 0.00% 99.83% 0.00% 0.17% 0.00% S202 0.00% 100.00% 0.00% 0.00% 0.00% 99.99% 0.00% 0.01% S203 0.00% 0.00% 100.00% 0.00% 0.10% 0.00% 99.89% 0.01% S204 0.00% 0.00% 0.00% 100.00% 0.00% 0.26% 0.00% 99.74% S205 90.00% 10.00% 0.00% 0.00% 99.68% 0.25% 0.01% 0.06% S206 0.00% 90.00% 10.00% 0.00% 0.02% 99.70% 0.21% 0.06% S207 0.00% 0.00% 90.00% 10.00% 0.01% 0.51% 89.19% 10.29% S208 0.00% 5.00% 0.00% 95.00% 0.00% 0.02% 0.00% 99.98% S209 10.00% 0.00% 0.00% 90.00% 0.01% 0.12% 0.00% 99.87% 88 Table 1. Comparison of the results from the experimental and the artificial neural network output of the validation experiment (continuation). Day Sample Experimental Model Nostoc sp. S. almeriensis S. platensis C. vulgaris Nostoc sp. S. almeriensis S. platensis C. vulgaris 2 S210 75.00% 25.00% 0.00% 0.00% 87.42% 12.42% 0.00% 0.16% S211 75.00% 0.00% 25.00% 0.00% 73.57% 0.00% 26.43% 0.00% S212 75.00% 0.00% 0.00% 25.00% 68.98% 0.87% 0.00% 30.15% S213 0.00% 75.00% 25.00% 0.00% 0.09% 82.03% 17.85% 0.03% S214 0.00% 75.00% 0.00% 25.00% 0.00% 79.11% 0.00% 20.88% S215 25.00% 75.00% 0.00% 0.00% 32.55% 67.43% 0.02% 0.00% S216 0.00% 0.00% 75.00% 25.00% 0.00% 0.28% 60.20% 39.52% S217 25.00% 0.00% 75.00% 0.00% 19.39% 0.00% 80.59% 0.01% S218 0.00% 25.00% 75.00% 0.00% 0.01% 19.11% 80.44% 0.44% S219 25.00% 0.00% 0.00% 75.00% 7.92% 0.03% 0.01% 92.04% S220 0.00% 25.00% 0.00% 75.00% 0.00% 17.98% 0.00% 82.02% S221 0.00% 0.00% 25.00% 75.00% 0.06% 0.02% 9.61% 90.31% S222 50.00% 50.00% 0.00% 0.00% 62.62% 37.37% 0.01% 0.00% S223 50.00% 0.00% 50.00% 0.00% 49.55% 0.00% 50.45% 0.00% S224 50.00% 0.00% 0.00% 50.00% 48.69% 0.59% 0.02% 50.70% S225 0.00% 50.00% 50.00% 0.00% 0.04% 38.49% 61.39% 0.09% S226 0.00% 50.00% 0.00% 50.00% 0.01% 53.80% 0.00% 46.20% S227 0.00% 0.00% 50.00% 50.00% 0.00% 0.01% 56.01% 43.98% S228 20.00% 0.00% 40.00% 40.00% 0.79% 0.56% 36.22% 62.44% Results and Discussion 89 According to these results, mixed algal samples with 10% or less contamination were weighted as monoalgal suspensions. Mixed algal samples composition were less accurate predicted in relation to monoalgal suspensions. In a study using flow cytometry with the SYTO9 stain, it was possible to classify C. vulgaris, Scenedesmus obliquus, Chlamydomonas reinhardtii, and Navicula pelliculosa, with errors ranging from 5-10%. However, the method misidentified microalgae cells in mixed algal samples (Peniuk et al., 2016). Given that contamination by non-target microalgae is a severe problem to microalgae cultivation (Wen et al., 2016), our method could be a powerful alternative for supervising algal cultures. Additionally, the method can provide information about the relative composition of a sample, an advantage over traditional methods in which more steps and time are needed to reach similar conclusions. The ANN gives the approximate composition of a sample based on data input processing; each absorbance bandwidth is weighted during the training phase. Monoalgal samples composition were slightly under 100%, probably due to the manner each absorbance bandwidth influences the ANN output. For instance, in S102, a S. almeriensis monoalgal sample, C. Vulgaris obtained a 1.30% prediction compared to 98.70% for S. almeriensis ― both microalgae species exhibited similar spectral signatures (Figure 27 right). Furthermore, C. vulgaris was overestimated in mixed algal samples, whereas S. platensis was underestimated (S110, S216 and S221). This confusion may be due to the manner that the resulting spectral signature bandwidths, in the mixed algal samples, are weighted by the ANN. To better show the accuracy of the developed ANN, a regression analysis of the experimental and modeled values was performed (Figure 28). The results show that, regardless of the microalga species, the ANN fitted the experimental values, with the R2 ranging from 0.951 to 0.970. Microalgae identification using ANN has been previously reported with taxonomic accuracy of up to 99%, using micrographic image analysis (Coltelli et al., 2017), similar accuracy to that achieved in the present study for monoalgal cultures. The method reported here demonstrated its efficiency in discriminating mixed algal cultures, whereas it was less efficient when there were smaller percentages of another species in the samples. If the ANN output indicated 90% or less, there was a high probability that the examined sample was derived from a mixed algal culture; conversely, if it was more than 90%, the culture should be examined to confirm that it was monoalgal. Therefore, this decreases the number of chemical analyses required to monitor the biological composition of microalgae cultures. 90 Figure 28. Correlation between the experimental composition of the samples and that predicted by the artificial neural network developed for the samples contained in Table 1. Re-training the ANN with more samples of mixed algal cultures in a wider variety of relative composition or using a higher spectral resolution could improve the ANN precision. The re-training could also be applied to incorporate more microalgae species into the model and test the model capability to differentiate a more significant number of microalgae. Re-training is a relatively fast process and can take anywhere between 5 minutes to about 1 hour depending on the computer hardware, training the ANN in a GPU is significantly faster than in a CPU only system (Ramachandran et al., 2015). Results and Discussion 91 4.1.3 Final Remarks A rapid methodology to elucidate microalgae species in suspensions has been developed and validated. To do this, microalgae spectral signatures from light absorption measurements of different microalgae species were analyzed through an ANN in order to describe and classify them. Four important species were used: Nostoc sp., Scenedesmus almeriensis, Spirulina platensis and Chlorella vulgaris. Absorbance from monoalgal and mixed algal cultures was the input data for training, testing and validating the ANN. The results show that the ANN was capable of distinguishing between monoalgal and mixed algal cultures, identifying the microalgae species in the monoalgal cultures and providing the approximate composition of mixed algal cultures. These results confirm that the application of spectral signatures with ANN is a suitable method for approximating the biological composition of microalgae cultures. 92 4.2 Artificial neural network for soil color analysis and characterization In this section, the soil characterization and color analysis by ANNs results will be detailed. The remaining subsections are organized as follows: the computer software developed for the color conversion from RGB color system to L*a*b* color system (section 4.2.1), soil samples Munsell color classification and description results (section 4.2.2), the characterization of soil fertility parameters through ANNs (section 4.2.3), and final remarks about the research topic (section 4.2.4). 4.2.1 Computer vision software for color conversion The DigiCIELAB is the CV software developed to perform the soil color experiments. Although the software was initially used for the soil color analysis, it is able to analyze the color of distinct objects such as fruits, plastic, color paints, and other objects. The DigiCIELAB was selected by the University of Valladolid in the “Prometeo” program, 2017 edition, and protected intellectually by the university in consequence of the award (Appendix B). Colorimeters are the most commonly used instrument to perform color measurements in scientific researches; however, this instrument has limitations for assessing some materials, especially those with the irregular color surface (Barbin et al., 2016; Girolami et al., 2013; Yam and Papadakis, 2004). A more economical and versatile alternative was developed using CV and an ANN to assess distinct types of surfaces, including irregularly colored. This application allows accelerating the process of color measurement through digital photographs analysis and the RGB color system conversion to the L*a*b* system, to obtain the CIELAB international standard for color researches. The user interface is shown in Figure 29. In the main screen of the software, two options are displayed; “Color Calibration” and “Colorimeter” as shown in Figure 29a and Figure 29b respectively. The “Color Calibration” option is used to generate the model capable of converting RGB to L*a*b* with a specific camera in a particular light condition. The “Colorimeter” option is used to perform the color measurement of new photographed samples. Results and Discussion 93 Figure 29. The user interface of DigiCIELAB CV software with the “Color Calibration” option (a) to generate the ANN model for the color analysis, and the “Colorimeter” option (b) for the color measurement of samples with the calibrated model. The application automatically detects the colorchecker in a digital color photograph, detects each color sub-frame from the card and extracts the RGB color components. These RGB values and the L*a*b* color values provided by the manufacturer are employed to generate an ANN model capable of converting color systems, from RGB to L*a*b*. To perform the color measurement of new objects, photographed samples must be taken in the same photographic set that the colorchecker used to calibrate the ANN, i.e., same illumination conditions and camera configuration. The workflow diagram of DigiCIELAB is shown below in Figure 30. a b 100 4.2.3 Characterization of soil fertility parameters through artificial neural networks The characterization of soil fertility parameters using soil color and ANNs are shown in Table 4. For the RGB color system, there are significative p values for parameters such as N, P, Ca, Fe, Cu, Mn, pH, and EC. However, just N and EC had an R2 higher than 0.5, with 0.62 and 0.76 respectively. The L*a*b* color system presented similar results; significative p values were obtained for the same variables as in the RGB system; showing R2 higher than 0.5 the N, pH and EC with of 0.57, 0.54 and 0.71 respectively. The color assessment has no response to OM content. Other studies showed higher correlations between color and SOC, R2 between 0.52 to 0.80 using multilinear regression models (Stiglitz et al., 2017; Viscarra Rossel et al., 2006). In contrast, a study of soil color relationship with N, SOC and clay with regression analysis techniques showed R2 values of less than 0.50 for soil samples in a 60 km2 Table 4. Regression analysis between actual and predicted values of soil fertility parameters using RGB and L*a*b* for ANNs. Parameter RGB L*a*b* R2 p R2 p OM 0.03 0.40_ 0.01 0.60_ N 0.62 0.00* 0.57 0.00* P 0.24 0.01* 0.19 0.03* K 0.11 0.10_ 0.10 0.13_ Ca 0.39 0.00* 0.39 0.00* Mg 0.02 0.53_ 0.03 0.40_ Fe 0.36 0.00* 0.36 0.00* Cu 0.48 0.00* 0.41 0.00* Mn 0.31 0.00* 0.23 0.02* Na 0.14 0.07_ 0.04 0.31_ Al 0.13 0.08_ 0.09 0.16_ pH 0.48 0.00* 0.54 0.00* EC 0.76 0.00* 0.71 0.00* * significative for p < 0.05 Results and Discussion 101 area with homogenous climate characteristics, but great diversity in soil parent material, aspect, topography, vegetation and land management (Ibáñez-Asensio et al., 2013). Correlations between soil organic components can be affected due to nature of the OM; soils with a high content of more humified substances show lower lightness (L*) than soils in which fulvic acids predominates upon the same content of SOC (Vodyanitskii and Kirillova, 2016). In the same manner, there is a nonproportionality relationship between the dark soil and the OM, for instance, when the humus content exceeds 6%, the soil color varies only slightly in comparison with soil color variation under this OM threshold (Valeeva et al., 2016). The organic substance present in Fe-coated particles also neutralizes its effect on the soil color (Vodyanitskii and Savichev, 2017). In Figure 33a, it is observed how soils with similar L* values had different OM content – twice as much in comparison; and a lighter soil with more OM than a darker soil as well. In such wise, the association of higher OM content in darker soils (Castañeda and Moret-Fernández, 2013) is not always in this line. Regarding the iron, another standard assumption is that red color is associated with this element in soils (Torrent et al., 2006; Vodyanitskii and Kirillova, 2016). However, the similar phenomena to those that took place with OM were observed for iron (Figure 33b). The color caused by iron concentrations can also be masked by hydrologic conditions and the weathering that can modify the iron color by changing the oxidation state (Maejima et al., 2000). In a study of ANN and soil color to predict 44 parameters, physical and chemical, in a national scale study (Aitkenhead et al., 2013), higher R2 were found (in average 0.40 for predictions to the RGB color system and 0.41 to the L*a*b* color system) compared to those obtained here (in average 0.31 to the RGB and 0.28 to the L*a*b*). In another study of ANN, using Landsat-8 satellite images, clay and OM for the prediction of soil cation exchange capacity, the obtained R2 was 0.80. However, both studies were not clear regarding the dataset used for the validation phase; training, testing and validating, all three phases, must be performed to generate a proper ANN model (de Oliviera et al., 2009), not using all phases can lead to an “apparently good” result product of overfitting or reduced generalization capacity (Piotrowski and Napiorkowski, 2013). 102 Figure 33. Soils samples with their respective lightness (L*) and OM (organic matter) contents (a); and samples with their respective red component (a* and R) and Fe (iron) contents (b). L* 53.22 OM 1.86% a) S083 L* 53.13 OM 0.93% S143 L* 68.62 OM 1.24% S059 L* 38.87 OM 0.64% S039 a* 6.81 Red 52.80 Fe (ppm) 1.54 S070 a* 26.12 Red 110.00 Fe (ppm) 0.88 S088 a* 23.87 Red 97.00 Fe (ppm) 2.20 S035 a* 6.97 Red 52.40 Fe (ppm) 3.87 S145 b) Results and Discussion 103 Despite that ANN has been used in plenty of studies, validation data is frequently not used. In fact, test and validation data is often assumed to be the same such as it is reported (Bagheri Bodaghabadi et al., 2016). The same authors working with ANN for interpolation and extrapolation of soil classes obtained results with errors of 30% for interpolations and 80% or higher for extrapolations in the validation phases. 4.2.4 Final remarks The color is a widely studied attribute in soils and it is routinely used as a parameter for soil description. In this research, the soil color was analyzed to determinate its capability to group soils according to levels of fertility parameters such as OM, N, P, K, Ca, Mg, Fe, Cu, Mn, Na, Al, pH and EC and perform soil properties characterization from color data. For that, soils samples were classified according to the Munsell color notation and the obtained hues were used for statistical analysis; in addition, with the RGB and L*a*b* color, two ANNs were trained, tested and validated to describe the soils. Soil aggrupation based on the Munsell color hue resulted not efficient to separate samples according to fertility levels of the studied variables, and regarding the ANNs approach to describe soils, the obtained models were not capable of providing accurate results. The experiments suggest that soil color does not contain enough information to predict soil fertility parameters. 104 4.3 Virtual weather stations for meteorological data estimations In this section, the VWS result and validation will be described. The interpolations algorithms were evaluated, and contrast to each other for assessing their performance in different seasons of the year. The following subsections are organized as: a statistical summary of the meteorological dataset, including a data breakdown by summer and winter seasons (section 4.3.1), the interpolation methods results comparison (section 4.3.2), seasons effects in the interpolated data quality (section 4.3.3), the VWS operation description (section 4.3.4), and final remarks about the research topic (section 4.3.5). 4.3.1 Statistical summary of the dataset After EDA for removing lecture errors and outliers, the statistical summary of the whole dataset was obtained (Table 5). During the studied period (from July 2017 to June 2018), 18,234 sets of observations were recorded for the 53 meteorological stations. The maximum registered precipitation was 42.85 mm in a day; the mean was 1.24 mm per day during the period. The mean ETo was 2.82 mm, which mean that the overall water balance is negative. Temperatures were 11.01 ⁰C on average, 42.97 ⁰C and -19.75 ⁰C for maximum and minimum registered. RH was 71.06 % on average, mean WS was 1.92 m/s2, and mean TSI was 15.99 MJ/m2. Table 5. Statistical summary of meteorological observations dataset (n = 18,234). Parameter Range Min. Max. Mean St. dev. Precip (mm) 42.85 0.00 42.85 1.24 3.52 ETo (mm) 11.02 0.11 11.13 2.82 1.93 Mean Temp (⁰C) 36.96 -8.95 28.01 11.01 7.19 Max. Temp (⁰C) 42.97 -2.50 40.47 18.25 8.81 Min. Temp (⁰C) 40.50 -19.75 20.75 4.31 6.04 Mean RH (%) 78.74 21.26 100.00 71.06 15.50 Max. RH (%) 63.05 36.95 100.00 92.68 8.55 Min. RH (%) 99.01 0.99 100.00 43.87 20.73 Mean WS (m/s) 17.72 0.01 17.73 1.92 1.22 TSI (MJ/m2) 33.56 0.46 34.02 15.99 8.44 Results and Discussion 105 The mean temperature was similar to the previously described in the same region (del Río et al., 2005), 11.17 ⁰C for a 37 years dataset (from 1961 to 1997) in comparisons with the 11.01 ⁰C obtained here in the 2017-2018 period. Regarding the precipitations, the same authors postulated an average of 664 mm for the same period, the average rainfall registered in the present study was 441.12 mm, less than the minimum of 480 mm registered in 1996 by the authors in the driest year. In a study from 1981 to 2010 in Castilla y León, the mean temperature was 11 ⁰C likewise (Nafría et al., 2013). The statistical summary of the meteorological variables during the summer and winter months are presented below in Table 6. Table 6. Statistical summary of meteorological observations during the summer months and the winter months. Parameter Range Min. Max. Mean St. dev. Summer (n = 4,744) Precip (mm) 42.85 0.00 42.85 0.50 2.69 ETo (mm) 10.19 0.94 11.13 4.89 1.47 Mean Temp (⁰C) 20.76 7.25 28.01 18.98 3.75 Max. Temp (⁰C) 26.25 14.22 40.47 27.94 4.56 Min. Temp (⁰C) 23.36 -2.61 20.75 10.12 3.63 Mean RH (%) 73.13 22.27 95.40 56.63 12.50 Max. RH (%) 60.96 39.04 100.00 87.58 10.59 Min. RH (%) 86.03 1.17 87.20 26.45 11.61 Mean WS (m/s) 17.34 0.39 17.73 1.79 0.95 TSI (MJ/m2) 32.01 1.98 33.99 23.16 5.51 Winter (n = 4,538) Precip (mm) 39.76 0.00 39.76 1.29 3.63 ETo (mm) 3.11 0.11 3.22 0.86 0.43 Mean Temp (⁰C) 23.36 -8.95 14.41 3.24 3.32 Max. Temp (⁰C) 22.67 -2.50 20.17 8.62 3.66 Min. Temp (⁰C) 32.63 -19.75 12.88 -1.46 4.17 Mean RH (%) 67.79 32.21 100.00 83.61 10.37 Max. RH (%) 37.44 62.56 100.00 95.99 4.36 Min. RH (%) 98.93 1.07 100.00 62.31 19.48 Mean WS (m/s) 11.18 0.01 11.19 2.09 1.37 TSI (MJ/m2) 18.62 0.46 19.08 7.24 3.97 106 The average precipitation during the winter is more than twice that in summer, 1.29 mm and 0.50 mm respectively; in this location, the summer is the driest season and the rainiest season is autumn or winter, depending on the specific site. (Nafría et al., 2013). Other variables such as temperatures, ETo, TSI are higher in summer; meanwhile, the RH is generally higher in winter in concordance with precipitations. The extreme temperatures, the minimum, and maximum registered in the entire year occurred during these two seasons. The standard deviation in these two seasons are lower in comparison with those seen in the whole year summary (Table 5), the range measurements of the phenomena were also narrow because of the uniformity of conditions during specific seasons in contrast to ranges of an entire year. The general behavior of the variables during the summer, winter, and the rest of the year is presented in Figure 34. It is observable that for ETo, temperatures, RH and TSI, more extreme lectures were obtained during summer and winter. Meanwhile, the rest of the year presented more intermedium values. The letters above the boxplots show the differences between groups by the Mann-Whitney U test. The variations of the means for precipitation, mean WS and max. RH in the periods were not as notorious as the previously mentioned variables; nonetheless, the statistical test found significant differences in these variables as well. 4.3.2 Interpolation methods comparison Different ANN models as well other interpolation methods were compared in terms results accuracy for meteorological variable estimations. Generally, good agreement between the estimated values and the actual records from the weather stations were observed for the interpolation methods. Results from ANNs approach are shown in Table 7, and results from alternative approaches are shown in Table 8. Between the 5 activation functions in the ANNs, the softsign had the higher R2, 0.84, following by the sigmoid and tanh with an R2 of 0.83, the relu with an R2 of 0.82 and the hardlim with the lowest R2 result, 0.79. The mean and maximum Temp were the most accurate variable to predict, with R2 in the range from 0.96 to 0.98; following by the ETo, TSI and minimum Temp with R2 higher than 0.91. Lower R2 were obtained for mean WS, maximum RH and Prep., with ranges from 0.43 to 0.75. In general rule, the temperature is a more precise variable to interpolate in comparison to the precipitations (Jeffrey et al., 2001). Results and Discussion 107 Figure 34. Boxplots and mean comparison of meteorological observations grouped by the season of the year. c a b a c b a c b a c b a c b c a b c a b c a b c a b a b a 108 Table 7. Analysis of meteorological data interpolations results for the ANNs approach. Parameter Hardlim Sigmoid Tanh Softsign Relu R2 p RMSE R2 p RMSE R2 p RMSE R2 p RMSE R2 p RMSE Prep. 0.62 0.00* 2.22 0.73 0.00* 1.88 0.71 0.00* 2.04 0.75 0.00* 1.82 0.73 0.00* 1.87 ETo 0.92 0.00* 0.55 0.94 0.00* 0.48 0.94 0.00* 0.46 0.94 0.00* 0.48 0.93 0.00* 0.51 Mean Temp 0.96 0.00* 1.45 0.98 0.00* 1.03 0.98 0.00* 1.10 0.98 0.00* 1.04 0.97 0.00* 1.15 Max. Temp 0.96 0.00* 1.86 0.98 0.00* 1.23 0.98 0.00* 1.28 0.98 0.00* 1.28 0.97 0.00* 1.49 Min. Temp 0.91 0.00* 1.83 0.93 0.00* 1.55 0.94 0.00* 1.52 0.93 0.00* 1.57 0.92 0.00* 1.71 Mean RH 0.83 0.00* 6.64 0.87 0.00* 5.87 0.88 0.00* 5.63 0.87 0.00* 5.68 0.86 0.00* 6.05 Max. RH 0.51 0.00* 6.24 0.57 0.00* 5.71 0.60 0.00* 5.50 0.58 0.00* 5.62 0.58 0.00* 5.66 Min. RH 0.81 0.00* 9.18 0.84 0.00* 8.40 0.85 0.00* 8.08 0.85 0.00* 8.06 0.82 0.00* 8.86 Mean WS 0.43 0.00* 1.06 0.47 0.00* 1.01 0.51 0.00* 0.96 0.51 0.00* 0.97 0.50 0.00* 0.98 TSI 0.93 0.00* 2.25 0.96 0.00* 1.81 0.96 0.00* 1.75 0.96 0.00* 1.70 0.96 0.00* 1.79 Mean 0.79 3.33 0.83 2.90 0.83 2.83 0.84 2.82 0.82 3.01 * significative for p < 0.05 Results and Discussion 109 Table 8. Analysis of meteorological data interpolation results for alternative methods. Parameter IDW ISDW MLR RFR R2 p RMSE R2 p RMSE R2 p RMSE R2 p RMSE Prep. 0.63 0.00* 2.16 0.71 0.00* 1.89 0.65 0.00* 2.11 0.72 0.00* 1.95 ETo 0.94 0.00* 0.48 0.95 0.00* 0.44 0.93 0.00* 0.50 0.94 0.00* 0.47 Mean Temp 0.97 0.00* 1.14 0.98 0.00* 0.92 0.98 0.00* 1.02 0.98 0.00* 0.92 Max. Temp 0.97 0.00* 1.47 0.98 0.00* 1.16 0.98 0.00* 1.27 0.98 0.00* 1.12 Min. Temp 0.93 0.00* 1.61 0.94 0.00* 1.43 0.93 0.00* 1.62 0.94 0.00* 1.50 Mean RH 0.83 0.00* 6.52 0.88 0.00* 5.57 0.84 0.00* 6.41 0.86 0.00* 5.97 Max. RH 0.58 0.00* 5.63 0.64 0.00* 5.24 0.52 0.00* 6.02 0.53 0.00* 5.98 Min. RH 0.81 0.00* 9.09 0.85 0.00* 8.07 0.80 0.00* 9.42 0.84 0.00* 8.40 Mean WS 0.50 0.00* 0.98 0.53 0.00* 0.94 0.48 0.00* 1.00 0.52 0.00* 0.96 TSI 0.95 0.00* 1.90 0.97 0.00* 1.54 0.95 0.00* 1.89 0.97 0.00* 1.60 Mean 0.81 3.10 0.84 2.72 0.81 3.13 0.83 2.89 * significative for p < 0.05 Results and Discussion 116 Other intrinsic factors that can alter the quality of the interpolated data are episodes of high-intensity rains, densely cloudy days and frost in winters (Thorsen and Höglind, 2010), temperature inversions (Bailey et al., 2011), and heat waves in summers (Luber and McGeehin, 2008; Meehl and Tebaldi, 2004). In a study using ANN to forecast the TSI, the mean error was different according to the month and season of the year in which the predictions were made, with higher errors in autumn and winter and lower in spring and summer (Kemmoku et al., 1999). The estimation of TSI by ANN using Meteosat-9 images as input was better in clearperiods than rainy or overcast ones; the RMSE was 21.20% against 5.13% for rainy and clear-sky days months respectively (Linares-Rodriguez et al., 2013). 4.3.4 Virtual weather station The algorithms composing the VWS are capable of access the InfoRiego FTP server, other servers address can be configurable into the script in order to access different FTP servers or similar protocols. Once the access is done, the user performs a filtered selection of the files to download information of a specified period. Thereupon, once the data is available in the user’s computer, the interpolations algorithms can be carried out executing the preferred one by the user, introducing as input the XY UTM coordinates in the ETRS89 geodetic system of the weather stations and paring this data with the station lectures. According to the results, the most appropriate methods to perform the data estimation thought interpolations in a given location are the ISDW and the ANN with the softsign function. The interpolations can be made to any given coordinate inside Castilla y León or other areas with a weather station network and accessible data to generate models. The innovative aspect of the VWS lies in the possibility of the user to choose a specific location, and estimated temperatures, RH, ETo, precipitations, TSI, WS with just one method, other studies of interpolations are focused just in a couple of variables using non-automated data access and processing. 4.3.5 Final Remarks Meteorological data is an important series of observations for agricultural activities. The data is generally obtained from automatic weather stations; however, Result and Discussions 117 the data can also be acquired from VWS. A VWS is an integration of algorithms to estimated meteorological data from nearby weather stations observations to other locations with no available stations. To develop the VWS, the performance of different interpolation methods were evaluated to test their accuracy. Daily data from an automatic weather station network were used to perform the interpolations. ANNs with the hardlim, sigmoid, tanh, softsign and relu activations functions were employed, as well as IDW, ISDW, MLR, and RFR to interpolate the daily observations. Additionally, interpolations in the summer and winter months were performed to check the capability of the models during periods with more extreme phenomena registers. The results showed that the interpolation methods have an R2 up to 0.98 for variables such as temperatures for the period of one year. Meanwhile, during the summer and winter, the models presented lower accuracy. From a practical perspective, the methods here described can be an alternative to meteorological data acquisition. CONCLUSIONS 121 5. Conclusions In this thesis, the application of ANN to solve distinct problems associated with agricultural activities were tested, in the microalgae production, soil fertility analysis, and meteorological variables acquisition. In the present section, the general conclusion (section 5.1) and the specific conclusions for each of the three main topics of research (sections 5.2, 5.3, and 5.4) will be detailed. Finally, future research niches based on the obtained result will be proposed (section 6). 5.1 General conclusions While this thesis explores three different topics ― microalgae, soil fertility, and meteorological data ― the ANN approach to solving problems associated with these subjects in the agricultural plane is the common ground in which the present work is substantiated. From the experiments performed, the following general conclusions are proposed: - Firstly, ANNs have proved to be a powerful tool to solve classifications, estimations and predictions problems. The use of ANNs in agricultural related issues is a critical step to find solutions for problems and help the users of this technology to make faster and better decisions in the productive chain, for instance, in the monitoring of microalgae cultures and crop management. By contrast, ANNs are not always capable of mapping the input variables with a target. In these cases, the network configuration must be rethought, dataset size increased and considered the fact that the input variables may not have a relation with the desired output. - Secondly, ANNs can perform and adapt to multiple problems. There is not strictly ANN architecture, activation function, error measure technique, and other parameters to ensure the best performance. Several models for each case were evaluated, and the selected model was the most accurate for the validation set; the dataset in which the model evaluation must be performed to conclude the inquiry. 122 5.2 Monoalgal and mixed algal cultures discrimination by using artificial neural network conclusions - It was demonstrated that microalgae light absorption spectra vary mainly as a function of the microalga species, although minor variations due to environmental and operational conditions can also take place. - When maintaining the cultures under similar conditions, the light absorption spectra can be used to develop an ANN that differentiates monoalgal from mixed algal cultures and identifies the predominant species. In addition, it is useful to be able to approximate the percentage of each species in mixed cultures. - A major advantage of this method is that it does not require much time or chemical analysis; a single light absorption spectrum is enough to quantify the biological composition of the cultures. It can provide a fast and powerful tool for microalgae culture management at the commercial scales. 5.3 Artificial neural network for soil color analysis and characterization conclusions - Although the soil color classification is a frequently used technique in the current days and is it helpful for preliminary classification of soils; in the present assessment, the color did not contain enough information to classify soils according to levels of fertility parameters. - The Munsell color hues analysis was not able to separate heterogeneous soil groups for all the fertility variables studied. For N, Ca, Mg, pH and EC statistical difference were detected, but no differences between all hues for a given variable. For the other eight fertility parameters (OM, P, K, Fe, Cu, Mn, Na and Al), no significative difference was found between Munsell hues. - The ANN approach to predict the fertility parameters in soils based on the color did provide an inaccurate model. The neural network was not able to generate a consistent model from the soil samples due to the nature of the dataset; soils samples did not exhibit a consistent pattern regarding the color and fertility attributes, soils with similar color values drastically differed in the fertility parameter and soils with different colors exhibited identical fertility levels. Conclusions 123 5.4 Virtual weather stations for meteorological data estimations conclusions - In the section, the concept of VWS was introduced. The current state of the weather station networks and the online availability of the records makes possible to get the data and process them to perform estimations of the meteorological observations from the measurements of weather stations to other locations with no stations. - The overall performances of the interpolation methods were accurate to estimate the meteorological variables; the ISDW and the ANN with the softsign activation functions were the most precise approach to perform the interpolations and to use in the VWS as preferred methods. - The occurrence of extreme meteorological events, such as peaks of high and low temperatures or intense rainfalls, had negative impacts on the accuracy of the interpolation models. Therefore, these events should be considered in the model generation process. - The success of the models suggests that they could be used in situations where no weather stations are present, but meteorological data series is needed for various purposes, for instance, the register of the ETo and calculus of crop water requirements. FUTURE WORK