scieee AI-readable full text Open interactive document viewer

A Comparison of Machine Learning Techniques Applied to Landsat-5 TM Spectral Data for Biomass Estimation

López Serrano, Pablito M.; López Sánchez, Carlos A.; Álvarez González, Juan G.; García Gutiérrez, Jorge

Abstract

Machine learning combines inductive and automated techniques for recognizing patterns. These techniques can be used with remote sensing datasets to map aboveground biomass (AGB) with an acceptable degree of accuracy for evaluation and management of forest ecosystems. Unfortunately, statistically rigorous comparisons of machine learning algorithms are scarce. The aim of this study was to compare the performance of the 3 most common nonparametric machine learning techniques reported in the literature, vis., Support Vector Machine (SVM), k-nearest neighbor (kNN) and Random Forest (RF), with that of the parametric multiple linear regression (MLR) for estimating AGB from Landsat-5 Thematic Mapper (TM) spectral reflectance data, texture features derived from the Normalized Difference Vegetation Index (NDVI), and topographical features derived from a digital elevation model (DEM). The results obtained for 99 permanent sites (for calibration/validation of the models) established during the winter of 2011 by systematic sampling in the state of Durango (Mexico), showed that SVM performed best once the parameterization had been optimized. Otherwise, SVM could be outperformed by RF. However, the kNN yielded the best overall results in relation to the goodness-of-fit measures. The findings confirm that nonparametric machine learning algorithms are powerful tools for estimating AGB with datasets derived from sensors with medium spatial resolution.

Full text

A Comparison of Machine Learning Techniques Applied to Landsat-5 TM Spectral Data for Biomass Estimation Pablito M. L´opez-Serrano1,CarlosA. L´opez-S´anchez2,JuanG. ´Alvarez-Gonz´alez3, and Jorge Garc´ıa-Guti´errez4,* 1Ciencias Agropecuarias y Forestales, Universidad Ju´arez del Estado de Durango, Negrete 800, Centro, 34000 Durango, Dgo., Mexico 2Instituto de Silvicultura e Industria de la Madera, Universidad Ju´arez del Estado de Durango, Universidad Ju´arez del Estado de Durango, Negrete 800, Centro, 34000 Durango, Dgo., Mexico 3Departamento de Ingenier´ıa Agroforestal, Universidad de Santiago de Compostela, Avenida Dr. ´Angel Echeverri, s/n. Campus Vida, 15782 Santiago de Compostela, C, Spain 4Departamento de Lenguajes y Sistemas Inform´aticos, Universidad de Sevilla, Reina Mercedes s/n., Sevilla 41012, Spain Abstract. Machine learning combines inductive and automated techniques for recognizing patterns. These techniques can be used with remote sensing datasets to map aboveground biomass (AGB) with an acceptable degree of accuracy for evaluation and management of forest ecosystems. Unfortunately, statistically rigorous comparisons of machine learning algorithms are scarce. The aim of this study was to compare the performance of the 3 most common nonparametric machine learning techniques reported in the literature, vis., Support Vector Machine (SVM), k-nearest neighbor (kNN) and Random Forest (RF), with that of the parametric multiple linear regression (MLR) for estimating AGB from Landsat-5 Thematic Mapper (TM) spectral reflectance data, texture features derived from the Normalized Difference Vegetation Index (NDVI), and topographical features derived from a digital elevation model (DEM). The results obtained for 99 permanent sites (for calibration/validation of the models) established during the winter of 2011 by systematic sampling in the state of Durango (Mexico), showed that SVM performed best once the parameterization had been optimized. Otherwise, SVM could be outperformed by RF. However, the kNN yielded the best overall results in relation to the goodness-of-fit measures. The findings confirm that nonparametric machine learning algorithms are powerful tools for estimating AGB with datasets derived from sensors with medium spatial resolution. R´ esum´ e. L’apprentissage automatique combine des techniques inductives et automatis´ ees pour la reconnaissance des formes. Ces techniques peuvent ˆ etre utilis´ ees avec des ensembles de donn´ ees de t´ el´ ed´ etection pour cartographier la biomasse a´ erienne « aboveground biomass » (AGB) avec un degr´ edepr ´ ecision acceptable pour l’´ evaluation et la gestion des ´ ecosyst` emes forestiers. Malheureusement, des comparaisons statistiquement rigoureuses des algorithmes d’apprentissage automatique sont rares. Le but de cette ´ etude ´ etait de comparer les performances des 3 m´ ethodes d’apprentissage automatique non param´ etriques les plus fr´ equemment rapport´ ees dans la litt´ erature, vis., les machines ` a vecteurs de support « Support Vector Machine » (SVM), les k plus proches voisins « k-nearest neighbor » (kNN) et les forˆ ets al´ eatoires « Random Forest » (RF), avec celle de la r´ egression lin´ eaire multiple param´ etrique (MLR) pour l’estimation de l’AGB provenant des donn´ ees de r´ eflectance spectrale de Landsat-5 Thematic Mapper (TM), des caract´ eristiques de texture d´ eriv´ ees de l’indice de v´ eg´ etation par diff´ erence normalis´ ee « Normalized Difference Vegetation Index » (NDVI) et des caract´ eristiques topographiques d´ eriv´ ees d’un mod` ele num´ erique de terrain « digital elevation model » (DEM).Les r´ esultats obtenus pour 99 sites permanents (pour la calibration/validation des mod` eles) ´ etablis au cours de l’hiver 2011 par l’´ echantillonnage syst´ ematique dans l’ ´ Etat de Durango (Mexique), ont montr´ e que les SVM montrent leurs meilleures performances une fois que le param´ etrage a ´ et´ e optimis´ e. Par ailleurs, les SVM pourraient ˆ etre surpass´ ees par les RF. Cependant, les kNN ont donn´ e les meilleurs r´ esultats globaux par rapport aux mesures d’ajustement. Les r´ esultats confirment que les algorithmes d’apprentissage automatique non param´ etriques sont des outils puissants pour l’estimation de l’AGB avec des ensembles de donn´ ees provenant de capteurs avec une r´ esolution spatiale moyenne. INTRODUCTION Forest biomass plays an important role in the global climate system because forest ecosystems absorb approximately 1/12 *Corresponding author e-mail: [email protected]. of Earth’s atmospheric carbon stocks every year (Malhi et al. 2002), and much of this carbon is stored as aboveground biomass (AGB). The importance of forest biomass has been underlined by the United Nations Framework Convention on Climate Change (UNFCCC), which has identified AGB as an Essential Climate Variable (GCOS 2010). Moreover, quantification of AGB and modeling of the associated dynamics are important to support decision-making models in different fields, including energy and materials provision for human use (FAO 2001, 2006), forest fragmentation (e.g., Malhi and Phillips 2004), and biodiversity conservation (e.g., Bunker et al. 2005). Accurate monitoring of forest biomass and how it changes at local to global scales is, therefore, of critical importance toward a better understanding of these processes (Lu 2006; Hartig et al. 2012; Le Toan and Quegan 2015). The most accurate method of estimating forest biomass is based on field measurements; however, estimating biomass in large areas is not an easy task and is hindered by the high costs (both time and money) associated with fieldwork (Lu et al. 2016). Remote sensing has been shown to be a practical option that helps to overcome these limitations because it enables obtaining forest information in large areas with reasonable effort. This is now the primary data source for large-scale biomass estimation (e.g., Andersen et al. 2011; Lu et al. 2016). Over the past few decades, the so-called passive sensors (i.e., sensors that use the solar radiation reflected or emitted by the objects detected at Earth’s surface) have been used to estimate AGB (e.g., Lu et al. 2012; Frazier et al. 2014). Considering the advantages and limitations of different remote sensing images, the mediumresolution (pixel size, 30 m) Landsat-5 TM sensor is one of the most widely used for biomass estimation (e.g., Agarwal et al. 2014; Pflugmacher et al. 2014; Dube and Mutanga 2015; Zhu and Liu 2015). The advantages of using the Landsat-5 TM sensor over high-resolution sensors, particularly for analysis of large ´ areas, are that numerous historical spatiotemporal archives are available (images since 1972) and the Landsat data is free of cost for users. For a review of Landsat imagery-based AGB estimations, see Wu et al. (2016). Regardless of thept type of sensor used, model accuracy and error estimation vary in relation to a series of factors such as the structure of the field data and the statistical techniques used (Ghosh et al. 2014). The most common model used in estimating forest biomass from remote sensing data is the regression-based model (e.g., Tian et al. 2012; Lu et al. 2012; Næsset et al. 2013); however, the accuracy of estimates obtained with small numbers of sample plots or when there is a weak linear relationship between variables and biomass is rather low (Lu et al. 2016). Nonparametric modeling approaches, which make no assumptions about the statistical distributions of the original data and relationships between predictor and response variables, have also been used to relate AGB and remotely sensed features. Various recent studies have explored the use of nonparametric approaches for estimating AGB with remote sensing data (e.g., Breidenbach et al. 2012; Mutanga et al. 2012; Jung et al. 2013; Fassnacht et al. 2014). Machine learning involves different techniques (mainly nonparametric) that focus on automated and inductive learning to recognize patterns (Cracknell and Reading 2014) in data (e.g., patterns in remote sensing data related to AGB in a set of located plots); once the pattern is learned, it can be applied to yield a prediction or classification in areas where it is not possible to carry out fieldwork to quantify an objective variable (e.g., AGB). In the last decade, various machine learning techniques such as Support Vector Machine (SVM), k-nearest neighbor (kNN) and Random Forest (RF) have been used to develop predictive models of AGB in large areas. Thus, Shataee (2013) showed that kNN performed better than SVMs, RF, and Artificial Neural Networks (ANN) for estimating biophysical variables such as basal area. More recently, Garcia-Gutierrez et al. (2015) showed that SVM models performed best for estimating forest variables from Light Detection and Ranging (LIDAR), while Wang et al. (2016) showed that RF outperformed SVM and ANN for estimating wheat biomass from remote sensing data. For a more complete review of research being carried out to retrieve vegetation biomass from remote sensing data, using machine learning methods, see Ali et al. (2016). The goodness-of-fitpt of models derived from spectral data are usually evaluated by the coefficient of determination (R2) and the root mean square error (RMSE). These measures report the performance of the model in predicting the data used to fit the model; however, because the quality of the fit does not necessarily reflect the quality of the prediction, assessment of their validity is often needed to ensure that the predictions represent the most likely outcome in the real world (Yang et al. 2004). The only method that can be regarded as “true” validation involves the use of a new independent dataset (Pretzsch et al. 2002; Yang et al. 2004); however, the scarcity of such data forces the use of alternative approaches, such as Cross Validation (CV), to enable evaluation of the quality of a particular fitting technique and minimize the risk of overfitting (Molinaro et al. 2005). Unfortunately, most studies involving estimation of AGB do not use CV as part of the model development. For rigorous comparison of the performance of different machine learning techniques, the study should also be accompanied by statistical validation of the results within a statistical framework (i.e., not merely calculating statistics such as R2or RMSE). Although this is well known in the field of machine learning (Garc´ ıa et al. 2010), this type of validation is not common in remote sensing, even though machine learning plays an important role in many biomass estimation studies. This fact might have led to some degree of discordance in the scientific literature, in which we can find examples of kNN, SVM, and RF outperforming each other (Shataee 2013; Garcia-Gutierrez et al. 2015; Wang et al. 2016). The objective of this study was to analyze and statistically compare the performance of 3 nonparametric techniques (SVM, kNN, and RF) and the parametric Multiple Linear Regression (MLR) technique for estimating AGB. The techniques were tested with Landsat-5 TM surface spectral reflectance data, texture features derived from the Normalized Difference Vegetation Index (NDVI), and topographical features derived from a digital elevation model (DEM) in the Sierra Madre Occi- FIG. 1. Geographical location of the study site and sample plots used in the study. dental (state of Durango, Mexico). The results obtained with each technique were compared after application of CV and posterior statistical validation of the mean rankings obtained for each. MATERIAL AND METHODS Study Area The study site is located in the Sierra Madre Occidental, in the north of the state of Durango (Mexico), and covers an area of 1,142,916 ha (Figure 1). The climate is humid temperate, with rainfall in summer (relative humidity, 50.1%). The average temperature ranges from 8 ◦Cto20◦C, and the annual precipitation is from 400 mm to 1200 mm. The average altitude above sea level in this area is 1,900 m. The vegetation comprises pine, oak, Douglas fir, pine-oak, and oak-pine forest, according to the description in the Land Use and Vegetation Cover Chart, scale 1:250,000, Series V (INEGI 2012). The forests are basically mixed and uneven-aged pine-oak stands, with a canopy cover ranging from 32% to 100%. These forests have been subject to selective harvesting for almost a century to provide a mixture of services to local communities. This structure is the result of the management history, which has depended on land ownership and the economic and social changes that have taken place in the state, as well as natural conditions (Wehenkel et al. 2011). Dataset Field Data A network of 99 permanent sampling plots was established during the winter of 2011, following the method described by Corral-Rivas et al. (2009). The plots were located by systematic sampling (with some exceptions to avoid nonforested areas) of a grid of equidistant points separated by 3 km or 5 km, depending on the accessibility, which is limited by the rugged terrain of the study area. In each plot (squares of side 50 m), all species of trees were recorded and the diameters at breast height (cm) and total height (m) of all standing trees were measured. Species-specific individual tree models developed by VargasLarreta (2013) were used to estimate the total AGB of field plots by tree value aggregation. The R2and the RMSE of the models used ranged from 0.87 kg–0.99 kg and 22.8 kg–95.2 kg, respectively. The mean, minimum, maximum and standard deviation of the AGB values per hectare of the sample plots are summarized in Table 1. Spectral Data The spectral data were derived from a satellite image Landsat-5 TM obtained in April 2011 (path 32, row 42) and covering the entire study ´ area.1Landsat-5 TM data have a 1Available from the US Geological Service webpage, at http://glovis.usgs.gov/ TABLE 1 Total biomass statistics expressed in Mg ha−1 No. of Observations Mean Standard Deviation Minimum Value Maximum Value 99 89.03 43.45 2.70 234.03 spatial resolution of 30 m with a revisit period of 16 days. Bands 1, 2, 3, 4, 5, and 7 (level L1T) of Landsat-5 TM were used in the present study; band 6 was not used because of its thermal characteristics, its coarse spatial resolution (120 m), and the low contrast in the forest area (NASA 2011). The satellite images were radiometrically, atmospherically, and topographically corrected by using the ATCOR3 Rmodule (Geosystems 2013), regarded as particularly suitable for mountainous zones. The ATCOR3 Rmodule first calculates the radiance at sensor level (W sr−1m−2) from the image pixel. Several input parameters were required for this calculation and were retrieved from the image metadata (header file): date of acquisition, scale factors, geometry (solar zenith angle and solar azimuth), and other information about the sensor calibration file (“gain and bias”). Other parameters were adjusted by taking into account the characteristics of the input datasets and the conditions of the imagery dates, e.g., visibility (35 km), pixel size of the DEM (15 m), aerosol type (rural), among others. Because the image was cloudless and no suitable water vapor bands were available, dehazing/cloud removal and atmospheric water retrieval settings were kept as “default,” which, in this case, is recommended by the ATCOR3 RUser Manual (Geosystems 2013). The corrections were implemented with the ERDAS RIMAGINE R2013 software. (ERDAS Inc. 2014). A number of vegetation indices were computed from the atmospherically and topographically corrected image bands and included in the biomass estimation models for evaluation as possible regressor features (Table 2). Texture Features The texture features homogeneity, contrast, dissimilarity, mean, standard deviation, entropy, second-order angular moment, and correlation (Haralick et al. 1973) were calculated from the NDVI image based on grey level cooccurrence matrices, with the aim of including information combining the spatial and spectral domain of the remotely sensed imagery in the biomass estimation models. We used NDVI texture features rather than each spectral band of Landsat-5 TM to avoid saturating high biomass values (Mutanga and Skidmore 2004). Because it also becomes more difficult to obtain an optimal subset as the number of attributes increases, we therefore aimed for a compromise between quantity and quality. The features were calculated using PCI Geomatica2013 Rsoftware,2and 3 2PCI Geomatics Inc. 2013 TABLE 2 Features (independent variables) for biomass estimation in comparison of machine learning techniques Abbreviation Variable Reference Vegetation Index NDVI Normalized Difference Vegetation Index Rouse et al. (1974) MSAVI2 Modified Soil-Adjusted Vegetation Index Qi et al. (1994) SAVI Adjusted Soil Vegetation Index Huete (1988) IAF Leaf Area Index Baret and Guyot (1991) ALB Albedo Asrar (1989) Fpar Fraction of Photosynthetically Active Radiation Asrar et al. (1984) FSR Flow Solar Radiation Brutsaerts (1975) Texture (NDVI) HOL Homogeneity Haralick et al. (1973) CO Contrast DI Dissimilarity ME Mean STD Standard Deviation EN Entropy ASM Angular Second Moment CR Correlation Terrain (DEM) Altitude Altitude B Slope TRASP Transformed Aspect Roberts and Cooper (1989) TSI Terrain Shape Index McNab (1989) WI Wetness Index Moore and Nieber (1989) PC Profile Curvature Wilson and Gallant (2000) PLC Plan Curvature CCurvature different scales of operation were considered by using moving window sizes of 3 ×3pixels,5×5 pixels, and 7 ×7pixels (Table 2). Terrain Features Terrain features are directly related to forest species composition, tree height growth, and other forest stand variables, en- abling these to be modeled (McNab 1989; Roberts and Cooper 1989). Firstand second-order terrain features were, therefore, derived from the 5 ×5-pixel low pass filtered DEM of the study area with a spatial resolution of 15 m. The DEM was derived from LIDAR data and corresponds to an array of elevation data interpolated to 15 m resolution from the coordinates of the last return of the pulses emitted (INEGI 2014). The final set of features derived from Landsat-5 TM sensor and from the DEM, which were used as possible predictors (independent variables) for estimating AGB (which played the role of dependent variable), are shown in Table 2. Finally, the sample plots were geopositioned with the aim of extracting the pixel value average with an associated buffer of 25 m for each described feature, to obtain a database with the mean biomass values and the associated features for each plot. The extraction was carried out using R statistical software (R Core Team 2014) and the “raster” package. Comparison Framework Machine Learning Techniques Three nonparametric machine learning techniques and one parametric technique were applied to data from the study area in order to compare their performance: (i) k-Nearest Neighbour (kNN), (ii) Support Vector Machine (SVM), (iii) Random Forest (RF), and (iv) Multiple Linear Regression (MLR). All these techniques were used to estimate AGB, using as possible predictors the variables included in Table 2. The parametric MLR technique is the most commonly used in this kind of study (Fassnacht et al. 2014). Moreover, this type of model is easy to understand and is widely used in most scientific disciplines. However, unlike the nonparametric approaches, MLR relies on certain assumptions, such as the fundamental least squares assumption of independence and equal distribution of errors with zero mean and constant variance, which can be violated by factors such as nonnormality of variables, multicollinearity of variables, and heteroscedasticity of error variance. Nearest neighbor (NN), a well-known machine learning technique used in remote sensing (Shataee 2013), makes a prediction by using the information about the neighbors of the instance to be regressed (Cover and Hart 1967). The NN depends on a parameter, usually called k, which determines the number of neighbors used by the algorithm. The technique is therefore usually called kNN when more than one neighbor is used. Although the idea behind this type of technique is quite intuitive, the resulting model is not easy to interpret because all results depend on a training set. SVMs have been developed from artificial neural networks (Cortes and Vapnik 1995) and have been used in many scientific fields (e.g., Abedi et al. 2012; Bayoudh et al. 2015; GarciaGutierrez et al. 2015). SVM models are developed by a set of vectors (or hyperplanes if greater dimension is requested) that separate instances of different labels (classification) or minimize the mean error (regression). Kernel functions are used to overcome the limitations associated with linear separability in SVM models. Appropriate selection of the kernel function and the kernel regularization parameters is important in relation to the SVM model behavior, which can make this type of technique more difficult to implement for users. As with kNN, the models produced using SVM are more difficult to interpret than those of MLR. RF is not exactly a classification or regression technique, but a combination of other techniques, mainly regression or classification trees (Breiman 2001). The success of this technique is based on the use of numerous trees, developed with different independent variables that are randomly selected from the complete original set of features (e.g., Deschamps et al. 2012; Wang et al. 2016). The number of predictors used by trees and the number of trees are established by the users. WEKA open source software (Hall et al. 2009) was used to implement all of the techniques compared. Thus, linear regression was used for MLR, IBk for kNN, SMOreg with polynomial and Gaussian kernels for SVM, and an adaptation of the RF implementation of WEKA for regression (using M5P as the basic regression technique for the development of this ensemble). Feature Selection, Parameterization and Validation In machine learning, spurious data features must be removed before a model is generated (Hall 1999). Thus, the variables that are potentially most important are selected. Some techniques (e.g., SVM and RF) carry out this selection, but others might be seriously affected by excessively large combinations of variables (e.g., the Hughes effect [Hughes 1968] in kNN and multicollinearity in MLR). This is a common situation in this type of analysis because of the large set of predictor variables that can be calculated from remote sensing data (Packal´ en et al. 2012). Moreover, correct functioning of different machine learning techniques depends on a proper parameterization (set-up of their parameters, i.e., variables that modify the behavior of the machine learning techniques). In this study, both of these steps (feature selection and parameterization) were carried out via a metaheuristic search (Samadzadegan et al. 2012). From the possible metaheuristic techniques (i.e., a method of optimization that provides a near-optimal solution in computationally affordable time), we selected an evolutionary algorithm, which is illustrated in Figure 2. The algorithm starts with a population of random solutions (Initial Population in Figure 2) called individuals and ranks them according to fitness of the individuals (Fitness Sorting in Figure 2). In the present study, the fitness was evaluated by the RMSE obtained with a training set. A new population of individuals is then created by mating parents (random selection of coefficients shown in Figure 2), selected with a probability proportional to their fitness, and later mutating the new individuals with a given probability (in this case, a value will be randomly selected and changed to a new random value, as can be seen in Figure 2). FIG. 2. Description of the evolutionary procedure used to determine the best methods for parameterization and feature selection. TABLE 3 Intervals used by the evolutionary algorithm to search for the different optimal parameters∗ Technique Name Minimum Maximum kNN k 1 20 SVM GAMMA (Gaussiankernel-only) 0.01 2.0 EXP (Polynomialkernel-only) 15 C 1 100 EPSILON 0.0 0.2 RF NT 1 100 NF 1 5 ∗Note: k =number of neighbors; EPSILON =determines the risk of overfitting; GAMMA =controls the transformation produced by the kernel; EXP =kernel’s exponent; C =penalty factor per instance of misclassification in training; NT =number of trees that form each ensemble; NF =number of attributes selected for constructing each tree The general scheme described in Figure 2 was modified slightly according to the specific regression technique. Thus, we used a specific design for MLR (see Garc´ ıa-Guti´ errez et al. 2014) and an adaptation of the genetic algorithm of Huang and Wang (2006) for the nonparametric techniques (kNN, SVM, and RF). In the kNN method, pure selection (coefficients associated with each feature as 1 or 0 depending on whether the predictor is selected or not) was substituted by weighting each attribute (real value between 0.0 and 1.0), which enables better adaptation of the algorithm to the characteristics of kNN (see Mateos et al. 2012). In SVMs, the type of kernel is another parameter to be optimized and had 2 possible values (radial basis function and polynomial). The parameters optimized for each machine learning technique are included in Table 3. For comparison of the different techniques, validation was based on the leave-one-out CV technique. This is a special case of k-fold CV in which kis equal to the number of observations and a prediction is obtained as many times as there are observations in the dataset (Packal´ en et al. 2012). In other words, an observation is excluded (target observation), and a prediction is computed with the other observations (reference observations). The prediction can be evaluated by the target observation. This procedure is repeated for every single observation. The final quality of a technique evaluated with CV is based on the averaged error obtained. A general description of the procedure is provided in Figure 3. Parameterization of each submodel at the different stages of the CV was repeated 5 times for each technique to prevent skew (due to the random nature of the evolutionary algorithms applied to predictor selection and parameterization). The best submodel and the average submodel for the 5 FIG. 3. Description of the leave-one-out CV evaluation of the techniques compared in the text. FIG. 4. Relative frequency of ocurrence (importance) of each attribute in the best models obtained by each technique (in terms of the sum of residuals). executions, ranked in terms of the RMSE reached in the evolutionary procedure, were used to calculate the goodness-of-fit statistics. Statistical Analysis The error of the predictions in the CV was compared for each technique in terms of R2and RMSE. In addition, for statistical analysis of differences between the methods, the absolute errors of the predictions made by each technique throughout the 99 iterations in the CV were compared (the number of iterations is equal to the number of instances in the database, which, in this case, refers to the 99 plots available). In theory, this should be carried out by Analysis of Variance (ANOVA), if the data comply with the underlying assumptions of independence, normality, and homoscedasticity required for parametric tests. These conditions can be tested by, respectively, the Shapiro-Wilk test, Lilliefor’s test, and Levenes’ test. If the data do not comply with these conditions, a nonparametric test such as the Friedman’s (aligned) test (described by Garc´ ıa et al. 2010) should be used. Friedman’s (aligned) test first obtains the mean ranking for each technique by taking into account the position obtained for FIG. 5. Relative frequency of ocurrence (importance) in the averaged models obtained by each technique (in terms of the sum of residuals). FIG. 6. Absolute frequency of relative position achieved by each technique (ranking) with the best parameterization of 5 executions. FIG. 7. Absolute frequency of the relative position achieved (ranking) by each technique with average parameterization. mation.” International Journal of Remote Sensing, Vol. 25: pp. 3999–4014. Næsset, E., Gobakken, T., Bollands˚ as, O.M., Gregoire, T.G., Nelson, R., and St˚ ahl. G. 2013. “Comparison of precision of biomass estimates in regional field sample surveys and airborne LiDAR-assisted surveys in Hedmark County, Norway.” Remote Sensing of Environment, Vol. 130: pp. 108–120. NASA 2011. Landsat 7 Science Data Users Handbook, accessed August 16, 2016, http://landsathandbook.gsfc.nasa.gov/pdfs/Land sat7 Handbook.pdf. Packal´ en, P., Temesgen, H., and Maltamo, M. 2012. “Variable selection strategies for nearest neighbor imputation methods used in remote sensing based forest inventory.” Canadian Journal of Remote Sensing, Vol. 38(No. 6): pp. 557–569. Pflugmacher, D., Cohen, W.B., Kennedy, R.E., and Yang, Z. 2014. “Using Landsat-derived disturbance and recovery history and LiDAR to map forest biomass dynamics.” Remote Sensing of Environment,Vol. 151: pp. 124–137. Pretzsch, H., Biber, P., ¨ Iursk´ y, J., von Gadow, K., Hasenauer, H., K¨ andler G., Kenk, G., et al. 2002. “Recommendations for standardized documentation and further development of forest growth simulators.” Forstwissenschaftliches Centralblatt, Vol. 121(No. 3): pp. 138–151. Qi, J., Chehbouni, A., Huete, A.R., and Kerr, Y.H. 1994. “A modified soil adjusted vegetation index.” Remote Sensing of Environment,Vol. 48(No. 2): pp. 119–126. R Core Team. 2014. R: A Language and Environment for Statistical Computing. Vienna, Austria: R Foundation for Statistical Computing. Roberts, D.W., and Cooper, S.V. 1989. “Concepts and techniques of vegetation mapping.” In Land Classifications Based on Vegetation: Applications for Resource Management, USDA Forest Service GTR INT-257, edited by D. Ferguson, P. Morgan, and F.D. Johnson, pp. 90–96. Ogden, UT: USDA Forest Service. Rouse, J.W., Haas, R.H., Schiell, J.A., Deferino, D.W., and Harlan, J.C. 1974. Monitoring the vernal advancement of retrogradation of natural vegetation. NASA/GSFC Type III, Greenbelt, MD: NASA. Samadzadegan, F., Hasani, H., and Schenk, T. 2012. “Simultaneous feature selection and SVM parameter determination in classification of hyperspectral imagery using ant colony optimization.” Canadian Journal of Remote Sensing, Vol. 38(No. 2): pp. 139–156. Shataee, S. 2013. “Forest attributes estimation using aerial laser scanner and TM Data.” Forest Systems, Vol. 22(No. 3): 484–496. Tian, X., Li, Z., Su, Z., Chen, E., Van der Tol, C., Li, X., Guo, Y., Li L., and Ling, F. 2014. “Estimating montane forest above-ground biomass in the upper reaches of the Heihe River Basin using LandsatTM data.” International Journal of Remote Sensing, Vol. 35(No. 21): pp. 7339–7362. Tian, X., Su, Z., Chen, E., Li, Z., van der Tol, C., Guo, J., and He, Q. 2012. “Estimation of forest above-ground biomass using multiparameter remote sensing data over a cold and arid area.” International Journal of Applied Earth Observation and Geoinformation, Vol. 14(No. 1): pp. 160–168. Vargas-Larreta, B. 2013. Estimaci´ on del potencial de los bosques de Durango para la mitigaci´ on del cambio clim´ atico. Modelizaci´ on de la biomasa forestal. Proyecto: FOMIX-DGO-2011-C01-165681. ITES, EL Salto (M´ exico). Wang, L., Zhou, X., Zhu, X., Dong, X., and Guo, W. 2016. “Estimation of biomass in wheat using random forest regression algorithm and remote sensing data.” The Crop Journal (in press). Wehenkel, C., Corral-Rivas, J.J., Hern´ andez-D´ ıaz, J.C., von Gadow, K. 2011. “Estimating balanced structure areas in multi-species forests on the Sierra Madre Occidental, Mexico.” Annals of Forest Science, Vol. 68: pp. 385–394. Wilson, J.P., and Gallant, J.C. 2000. “Digital terrain analysis.” In Terrain Analysis: Principles and Applications. New York, NY: John Wiley & Sons. Wu, C., Shen, H., Wang, K., Shen, A., Deng, J., and Gan, M. 2016. “Landsat imagery-based above ground biomass estimation and change investigation related to human activities.” Sustainability, Vol. 8(No. 2): 159. Xie, X., Liu, W.T., and Tang, B. 2008. “Space-based estimation of moisture transport in marine atmosphere using support vector regression.” Remote Sensing of Environment, Vol. 112(No. 4): pp. 1846–1855. Yang, Y., Monserud, R.A., and Huang, S. 2004. “An evaluation of diagnostic tests and their roles in validating forest biometrics models.” Canadian Journal of Forest Research, Vol. 34: pp. 619–629. Zhang, Y., Liang, S., and Sun, G. 2014. “Forest biomass mapping of northeastern China using GLAS and MODIS data.” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, Vol. 7: pp. 140–152. Zhao, K., Popescu, S., Meng, X., Pang, Y, and Agca, M. 2011. “Characterizing forest canopy structure with LiDAR composite metrics and machine learning.” Remote Sensing of Environment, Vol. 115(No. 8): pp. 1978–1996. Zhu, X., and Liu, D. 2015. “Improving forest aboveground biomass estimation using seasonal Landsat NDVI time-series.” ISPRS Journal of Photogrammetry and Remote Sensing, Vol. 102: pp. 222–231.