scieee AI-readable full text Open interactive document viewer

Repositorio Institucional de Documentos

Abstract

Los espacios forestales son una fuente de servicios, tanto ambientales como económicos, de gran importancia para la sociedad. La caracterización de estos ambientes ha requerido tradicionalmente de un laborioso trabajo de campo. La aplicación de técnicas de teledetección ha proporcionado una visión más amplia a escala espacial y temporal, a la par que ha generado una reducción de los costes. La utilización de sensores óptico-pasivo multiespectrales y de sensores radar posibilita la estimación de parámetros forestales, si bien el desarrollo de sensores LiDAR, como el caso de los escáneres láser aeroportados (ALS), ha mejorado la caracterización tridimensional de la estructura de los bosques. La disponibilidad pública de dos coberturas LiDAR, generadas en el marco del Plan Nacional de Ortofotografía Aérea (PNOA), ha abierto nuevas líneas de investigación que permiten proporcionar información útil para la gestión forestal. <br />La presente tesis utiliza datos LiDAR aeroportados de baja densidad para estimar diversas variables forestales, con ayuda de trabajo de campo, en masas forestales de Pino carrasco (Pinus halepensis Miller) en Aragón. La investigación aborda dos cuestiones relevantes como son la exploración de las metodologías más adecuadas para estimar variables forestales considerando escalas locales y regionales, teniendo en cuenta las posibles fuentes de error en el modelado; y, además, analiza la potencialidad de los datos LiDAR del PNOA para el desarrollo de aplicaciones forestales que valoricen las áreas forestales como recursos socio-económicos. <br />La tesis se ha desarrollado según la modalidad de compendio de publicaciones, incluyendo cuatro trabajos que dan respuesta a los objetivos planteados. En primer lugar, se realiza un análisis comparativo de distintos modelos de regresión, paramétricos y no paramétricos, para estimar la pérdida de biomasa y las emisiones de CO2 en un incendio, mediante la utilización de datos LiDAR-PNOA y datos ópticos del satélite Landsat 8. En segundo lugar, se explora la idoneidad de distintos métodos de selección de variables para estimar biomasa total en masas de Pino carrasco utilizando datos LiDAR de baja densidad. En tercer lugar, se cuantificó y cartografió la biomasa residual forestal en el conjunto de masas de Pino carrasco de Aragón y se evaluó el efecto de diversas características de la tecnología LiDAR y de las variables ambientales en la precisión de los modelos. Finalmente, se analiza la transferibilidad temporal de modelos para estimar a escala regional siete variables forestales, utilizando datos LiDAR-PNOA multi-temporales. A este respecto, se compararon dos enfoques que permiten analizar la transferibilidad temporal: en primer lugar, el método directo ajusta un modelo para un determinado punto en el tiempo y estima las variables forestales para otra fecha; por otra parte, el método indirecto ajusta dos modelos diferentes para cada momento en el tiempo, estimando las variables forestales en dos fechas distintas. <br />Los resultados obtenidos y las conclusiones derivadas de la investigación indican que la técnica basada en coeficientes de correlación de Spearman y el método de selección por todos los subconjuntos constituyen los métodos de selección de métricas LiDAR más apropiados para la modelización. El análisis de métodos de regresión para la estimación de variables forestales indicó <br />que su idoneidad variaba de acuerdo con el tamaño y complejidad de la muestra. El método de regresión linear multivariante arrojó mejores resultados que los métodos no-paramétricos en el caso de muestras pequeñas. Por el contrario, el método Support Vector Machine produjo los mejores resultados con muestras grandes. El incremento de la densidad de puntos y de los valores de penetración de los pulsos LiDAR en el dosel, así como la presencia de ángulos de escaneo pequeños, incrementó la exactitud de los modelos. De forma similar, el incremento de la pendiente y la presencia de arbustos en el sotobosque implican una reducción en la exactitud de los modelos. En la estimación de variables forestales utilizando datos LiDAR multi-temporales, aunque la utilización del enfoque indirecto arrojó generalmente una mayor precisión en los modelos, se obtuvieron resultados similares con el enfoque directo, el cual constituye una alternativa óptima para reducir el tiempo de modelado y los costes de realización de trabajo de campo. La fusión de datos LiDAR y datos óptico-pasivos ha evidenciado la conveniencia de los métodos aplicados para cuantificar las emisiones de CO2 a la atmósfera generadas por un incendio. Esta metodología constituye una alternativa adecuada cuando no existen datos multi-temporales LiDAR. La estimación de variables de inventario forestal, así como de diversas fracciones de biomasa, como la biomasa total y la biomasa residual forestal, proporciona información valiosa para caracterizar las masas forestales mediterráneas de Pino carrasco y mejorar la gestión forestal<br /> Forest ecosystems provide environmental and economic services of great importance to the society. The characterization of these environments has been traditionally accomplished with intense field work. In comparison, the application of remote sensing tools provides a greater overview over large spatial and temporal scales while minimizing costs. Although optical data and Synthetic Aperture Radar (SAR) allow estimating forest stand variables, the development of LiDAR sensors such as Airborne Laser Scanner (ALS) have improved three-dimensional characterization of forest structure. The availability of two ALS public data coverages for the Spanish territory, provided by the National Plan for Aerial Ortophotography (PNOA), opens new research opportunities to generate useful information for forest management. This PhD Thesis used low-density ALS-PNOA data to estimate different forest variables, with support in fieldwork, in the Aleppo pine (Pinus halepensis Miller) forests of Aragón region. The addressed research is relevant mainly for two reasons: first, the examination of suitable methodologies and error sources in forest stand variables prediction at local (small area) and regional scales (large area), and second, the application of ALS data to the characterization of forest areas as a socio-economic reservoir. This PhD Thesis is a compendium of four scientific papers, which sequentially answer the objectives established. Firstly, a comparative analysis of different parametric and non-parametric models was performed to estimate biomass losses and CO2 emissions using low-density ALS and Landsat 8 data in a burnt Aleppo pine forest. Secondly, we assess the suitability of variable selection methods when estimating total biomass in Aleppo pine forest stands using low-density ALS data. In the third manuscript, the quantification and mapping of forest residual biomass in Aleppo pine forest of Aragón region and the assessment of the effect of ALS and environmental variables in model accuracy were accomplished. Finally, the temporal transferability of seven forest stands attributes modelling using multi-temporal ALS-PNOA data in Aleppo pine forest at regional scale was explored. In this case, the temporal transferability was assessed comparing two methodologies; the direct and indirect approach. The first one fits a model for one point in time and estimates the forest variable for another point in time. The indirect approach adjusts two models in different points in time to estimate the forest variables in two different dates. The results derived from this research indicated that Spearman’s rank and All Subset Selection are the most appropriate methods in the ALS metrics selection step commonly applied in modelling. The suitability of the regression methods depends on the sample size and complexity. Thus, multivariate linear regression outperformed non-parametric methods with small samples while support vector machine was the most accurate method with larger samples. Model accuracy increased with higher point density and canopy pulse penetration, while decreasing with wider scan angles. Furthermore, the presence of steep slopes and shrub reduced model performance. In the case of forest stand variables prediction using multi-temporal ALS data, although the indirect approach produced generally a higher precision, the direct approach provided similar results, constituting a suitable alternative to reduce modelling time and fieldwork costs. The fusion of ALS and passive optical data have evidenced the suitability of this information for quantifying wildfire CO2 emissions to atmosphere, constituting a good alternative when multi-temporal ALS data is not available. The estimation of forest inventory variables as well as different biomass fractions, such as total biomass and forest residual biomass, provided valuable information to characterize Mediterranean Aleppo pine forests and improve forest management.<br /> Domingo Ruiz, Darío; De la Riva Fernández, Juan; Lamela Gracia, María Teresa

Full text

2021 48 Darío Domingo Ruiz Characteriation of Mediterranean Aleppo pine forest using lowdensity ALS data Departamento Director/es Geografía y Ordenación del Territorio De la Riva Fernández, Juan Lamela Gracia, María Teresa © Universidad de Zaragoza Servicio de Publicaciones ISSN 2254-7606 Darío Domingo Ruiz CHARACTERIATION OF MEDITERRANEAN ALEPPO PINE FOREST USING LOW-DENSITY ALS DATA Director/es Geografía y Ordenación del Territorio De la Riva Fernández, Juan Lamela Gracia, María Teresa Tesis Doctoral Autor 2019 UNIVERSIDAD DE ZARAGOZA Repositorio de la Universidad de Zaragoza – Zaguan http://zaguan.unizar.es Darío Domingo Ruiz Directores: Mª Teresa Lamelas Gracia y Juan de la Riva Fernández PhD Thesis Zaragoza 2019 Characterization of Mediterranean Aleppo pine forest using low-density ALS data The author of this PhD Thesis was supported by Government of Spain, Department of Education Culture and Sports under Grant (FPU Grant BOE, 14/06250). Furthermore, the work was supported by FoResBiomALS research project from the Centro Universitario de la Defensa de Zaragoza (2017-09), SERGISAT research project (CGL2014-57013-C2-2-R) from the Spanish National Plan for Scientific and Technical Research and National Innovation Plan (Ministry of Economy and Competitiveness), and Geoforest-IUCA research group (Group S51_17R, financed by FEDER 2014-2020 and Government of Aragon, “Construyendo Europa desde Aragón”). Cover page image: RGB and NIR coloured point clouds (tile 664-4644) from PNOA ©INSTITUTO GEOGRÁFICO NACIONAL DE ESPAÑA ‐ Autonomous Community of Aragón. This PhD Thesis is developed as compendium of papers according to the Doctoral program in Ordenación del Territorio y Medio Ambiente at University of Zaragoza. The PhD student, Darío Domingo, is the first author and responsible for each and every article listed below. The references of the works that constitute the PhD Thesis body are the following: 1. Domingo, D., Lamelas, M.T., Montealegre, A.L., de la Riva, J. 2017. Comparison of regression models to estimate biomass losses and CO2 emissions using low-density airborne laser scanning data in a burnt Aleppo pine forest. European Journal of Remote Sensing, 50 (1), 384-396. doi: 10.1080/22797254.2017.1336067. 2. Domingo, D., Lamelas, M.T., Montealegre, A.L., García-Martín, A., de la Riva, J. 2018. Estimation of total biomass in Aleppo pine forest stands applying parametric and nonparametric methods to lowdensity airborne laser scanning data. Forests, 9, 158-175. doi: 10.3390/f9030158. 3. Domingo, D., Montealegre, A.L., Lamelas, M.T., García-Martín, A., de la Riva, J., Rodríguez, F, Alonso, R. 2019. Quantifying forest residual biomass in Pinus halepensis Miller stands using Airborne Laser Scanning data. GIScience and Remote Sensing, 56 (8), 1210-1232. doi: 10.1080/15481603.2019.1641653. 4. Domingo, D., Alonso, R, Lamelas, M.T., Montealegre, A.L., Rodríguez, F, de la Riva, J. 2019. Temporal Transferability of Pine Forest Attributes Modeling Using Low-Density Airborne Laser Scanning Data. Remote Sensing, 11 (3), 261. doi:10.3390/rs11030261. La presente tesis doctoral se ha elaborado como compendio de publicaciones siguiendo la modalidad ofrecida por el programa de Doctorado en Ordenación del Territorio y Medio Ambiente de la Universidad de Zaragoza. El doctorando, Darío Domingo, figura como primer autor y responsable de todos y cada uno de los artículos publicados. A continuación, se detallan las referencias de los trabajos que constituyen el cuerpo de la tesis: 1. Domingo, D., Lamelas, M.T., Montealegre, A.L., de la Riva, J. 2017. Comparison of regression models to estimate biomass losses and CO2 emissions using low-density airborne laser scanning data in a burnt Aleppo pine forest. European Journal of Remote Sensing, 50 (1), 384-396. doi: 10.1080/22797254.2017.1336067. 2. Domingo, D., Lamelas, M.T., Montealegre, A.L., García-Martín, A., de la Riva, J. 2018. Estimation of total biomass in Aleppo pine forest stands applying parametric and nonparametric methods to lowdensity airborne laser scanning data. Forests, 9, 158-175. doi: 10.3390/f9030158. 3. Domingo, D., Montealegre, A.L., Lamelas, M.T., García-Martín, A., de la Riva, J., Rodríguez, F, Alonso, R. 2019. Quantifying forest residual biomass in Pinus halepensis Miller stands using Airborne Laser Scanning data. GIScience and Remote Sensing, 56 (8), 1210-1232. doi: 10.1080/15481603.2019.1641653. 4. Domingo, D., Alonso, R, Lamelas, M.T., Montealegre, A.L., Rodríguez, F, de la Riva, J. 2019. Temporal Transferability of Pine Forest Attributes Modeling Using Low-Density Airborne Laser Scanning Data. Remote Sensing, 11 (3), 261. doi: 10.3390/rs11030261. Agradecimientos “aprender a caminar es más fácil si alguien te tiende su brazo” En primer lugar, me gustaría expresar mi más sincero agradecimiento a mis directores de tesis, la Dra. María Teresa Lamelas Gracia y el Dr. Juan de la Riva Fernández por su apoyo incondicional y por su buen saber hacer que han permitido que la investigación desarrollada haya llegado a buen puerto. Gracias por haber confiado en mí desde el inicio, por haberme dado ideas y también por haberme dejado abordar las mías propias. Me habéis enseñado a crecer en lo profesional y en lo personal y estoy seguro de que sin vuestro tesón y aliento esto no hubiera sido posible. También quiero agradecer a los coautores de las distintas publicaciones, el Dr. Antonio Luis Montealegre, el Dr. Alberto García Martín, el Dr. Rafael Alonso y el Dr. Francisco Rodríguez, en especial por sus valiosas aportaciones, por sus ideas y por todo el esfuerzo que han dedicado. Asimismo quiero agradecer a todos los miembros del Grupo de Investigación GEOFOREST-IUCA, del Departamento de Geografía y Ordenación del Territorio y al Centro Universitario de la Defensa por acogerme y darme la oportunidad de formarme como doctorando, así como por proporcionarme los recursos económicos y materiales para poder desarrollar las investigaciones. En concreto, el Departamento citado es la casa donde inicié mi formación y ha sido un grato placer poder seguir formando parte de esta familia geográfica de inmejorable calidad humana. Agradezco a los investigadores que me acogieron durante mis estancias en el extranjero por todo el apoyo mostrado, por los conocimientos que me enseñaron y por estar siempre pendientes de mí. Muchas gracias a Warren B. Cohen y a Yang Zhiqiang por acogerme en el Laboratory for Applications of Remote Sensing in Ecology (LARSE) del USDA Forest Service en Corvallis. Así mismo, muchas gracias a Erik Naesset, a Terje Gobakken y a Hans Ole Ørka por el trato magnífico, la dedicación y el buen saber hacer cuando me acogieron en Norwegian University of Life Sciences de Ås. No me puedo olvidar de Antoine y Nemo que hicieron la estancia en Corvallis mucho más amena y gratificante pese a estar lejos de casa. Tampoco puedo olvidarme de Marie-Claude, Ana, Ida, Lennart y Victor por todos los ratos de pingpong y las salidas al campo. También quiero agradecer a los revisores externos de la tesis, a Ole Martin Bollandsås y a Hooman Latifi por su disponibilidad y aportaciones al presente documento. Mi más sentido agradecimiento a todos los miembros del proyecto SERGISAT con los que he pasado intensas jornadas de campo que te hacen crecer en lo personal y laboral, porque el campo une y de qué manera. Gracias a Maite, a Teresa, a Juan, a Paloma, a Alberto, a Pere, a Raúl a Marcos y a Demetrio. Agradezco a todos los compañeros y doctorandos que me han alentado y ayudado durante este periodo. Daniel Borini, Adrián Jiménez, Antonio Montealegre, Xavier Garate, Estela Pérez, Olga Rosero, Daniel Ballarín, Daniel Mora, Aldo Arránz, Ricardo Badía, Samuel Esteban y Andrea Urgilez. Muchas gracias por todos y cada uno de esos cafés juntos. Del mismo modo agradezco a Katalin Varga y a Antonio Montealegre por los buenos momentos que hemos compartido en diversos cursos y congresos. También quiero acordarme de Yago, hace 9 años me dijiste “apúntate a Geografía que te gusta el campo y seguro que te va bien”. Muchas gracias, no fallaste. 1 1. Introduction This chapter describes the main concepts and the conceptual framework in which this PhD Thesis was developed. Firstly, ALS technology and its use for forestry purposes are described. Secondly, the use of ALS data for estimating forest stand variables, at local and regional scales, using different regression algorithms is presented. Then, the research justification, hypothesis and aims are defined. Finally, the chapter depicts the PhD Thesis structure, including the developed research items and their link to the different objectives that form a thematic unity. Introduction 3 1.1. Background 1.1.1. Airborne laser scanning The characterization and quantification of forest resources started in Europe during the late 18th century when society was concerned about wood availability, being the main source of fuel. The estimation of forestry metrics, especially volume, was performed by visual interpretation. The development of forest surveys, inventory tools, sampling and statistical methods and the advance in computers during 19th and 20th century yield great progress in forestry science. The growth of remote sensing tools such as aerial images, optical passive satellite images, and Synthetic Aperture Radar (SAR) data provided a greater overview of forest resources over large areas (Boyd & Danson, 2005). However, the expansion of Light Detection and Ranging (LiDAR) technology has improved the three-dimensional (3D) characterization of forest ecosystems being a suitable tool for forestry variables estimation (Zhao et al., 2018). LiDAR technology measures the distance between a laser transmitter and an object or surface using a monochromatic beam of light, coherent and directional (Andersen et al., 2005). The origin of this technology started in the early 1960s when Theodore Harold Maimam developed the first ruby laser that emits powerful pulses of collimated red light. In the 1980s the profile LiDAR was used for forestry application (Aldred & Bonnor, 1985; Maclean & Krabill, 1986). Further growth occurred in the 90s with the generation of Digital Terrain Models (DTM) and forest inventory variables estimation (Lefsky et al., 1999; Means et al., 1999; Næsset, 1997). During the last two decades the use of LiDAR technology have exponentially grown, being developed diverse hardware, software and applications within the forestry topics. According to the platform used, three main types of LiDAR technologies exist: terrestrial laser scanners (TLS), airborne laser scanners (ALS) and satellite laser scanners (SLS). ALS is one of the most widespread for forestry purposes (Maltamo et al., 2014). Topographic ALS sensors emit their own electromagnetic flux within the infrared wavelengths (900 to 1,064 nm). These wavelength denotes high reflectance values of vegetation and transmissivity of atmosphere (Lefsky et al., 2002). The basic information captured by ALS is denominated point cloud, referring to a dense set of x, y and z coordinates that register the object reflexions reached by the laser light. ALS technology is usually classified in two main types, according to the way it measures distances between the sensor and the object reached by the laser beam: full-waveform systems and discrete return systems. Full-waveform systems completely register the reflected energy. The distance between transmitter and object is measured using phase difference between emitted signal and scattered radiation. Full-waveform sensors provide richer data than discrete return systems, while the processing is more demanding. Although several processing methodologies has been proposed such as voxelization, sometimes the wavelengths are converted into point clouds similar to the ones provided by discrete sensors but with higher number of returns per pulse. On the other hand, discrete return systems capture one to five returns per emitted laser pulse. The distance between the transmitter and the object is measured as a function of time. The generalized use of discrete Characterization of Mediterranean Aleppo pine forest using low-density ALS data 4 return sensors for forestry and topographic purposes may be explained by the higher implementation in the commercial sector (Shan & Toth, 2008). ALS discrete-return systems have been widely used for estimating forest variables. Trees are porous objects from the laser pulse point of view. The laser beam can travel through the canopy and the system is able to capture up to five returns. The first return in a forested area may reach the top of the tree or canopy surface, the intermediate or low returns might be scattered by the tree branches and leaves or even the understory, while the last return might reach the terrain. This capability of ALS sensors provides a reliable 3D representation of forest structure. Canopy penetration pulse varies according to canopy closure, determining the points that reach the terrain. According to Chasmer et al. (2006b) only 50% of last returns in forested areas are backscattered by the terrain. ALS sensors also capture, through a photodiode, the energy reflected by the objects. This information is denominated intensity, registered in 8 or 12 bits. The intensity refers to each laser echo with a footprint varying from 0.2 to 1.0 m in small footprint discrete returns ALS sensors (Andersen et al., 2005; Evans et al. 2009). Intensity values varies according to different parameters such as surface roughness and wetness, flight height, angle of incidence, instrumental characteristics, atmospheric conditions, between others. Consequently, the use of intensity values requires the normalization or calibration of the data in each acquisition. In this sense, although intensity data has been used for some applications, such as species classification (Korpela et al., 2010; Watt et al., 2007), it is still not broadly implemented for forestry purposes. The main components of an ALS discrete sensor are the platform, the laser sensor, the Global Navigation Satellite System (GNSS), the Inertial Measurement Unit (IMU) and the data manager or computer with specific software. The laser scanner includes the laser pulse transmitter, the scanning mechanism and the receptor to record the distance to the objects. The laser scanner emits laser pulses, with a scanning frequency of up to 300 kHz, directed to the terrain surface. The scanning mechanism (oscillating mirror, rotating polygon, palmer scan, fibre scanner, etc.) draws specific scanning patterns to capture the terrain surface in each flight line (Vosselman & Maas, 2010). The scanning angle or sensor field of view (FOV) and the flight height determine the strip width. Accordingly, the overlap between strips modifies the flight time and accuracy (Evans et al., 2009). The GNSS rover unit, located inside of the plane, collect the position from the GNSS satellites. The enhancement of GNSS position accuracy, performed at real time using reference stations or base differential GNSS at the ground level, provide a centimeter nominal accuracy. The IMU includes the Inertial Navigation System (INS), managing the pith, roll and yaw of the plane (Baltsavias, 1999). The use of computers to manage the data collected by the GNSS and IMU with specific software allows providing the x, y and z coordinates for each return pulse along the flight acquisition, constituting the ALS point cloud (Baltsavias, 1999). 1.1.2. Forest variable estimation using remote sensing data ALS data have been proven as a suitable technique for mapping forest vertical and horizontal structure as well as to derive forestry variables (Zhao et al., 2018). The use of structural and Introduction 5 textural information derived from passive optical data have been explored to derive forestry parameters (Dube & Mutanga, 2015; Pfeifer et al., 2012). In this sense, the availability of wide temporal series have provided better results on forestry variables estimation when characterizing disturbance history (Cohen et al., 1996; Coops & Waring, 2001). However, the data captured by passive optical sensors tend to saturate under closed canopy conditions and in dense forests (Lu, 2006). The use of orthophotography to derive forestry parameters such as stand density, height or cover, between others, has also been explored, providing lower accuracies than active sensors (Campbell, 2006) such as SAR or LiDAR. SAR systems allow characterizing forest structure at global scale. This technology uses different wavelengths from the microwave to provide information about leaves, branches and stems (Periasamy, 2018; Tanase et al., 2014). Although SAR data availability have increased and there exist recent advances in processing software, as for example the tools provided by the Copernicus program, difficulties still arise in estimating forestry variables in heterogeneous and dense forests (Hyde et al., 2006). The improvement of structure from motion (SfM) algorithms and photogrammetric techniques provides new insights in the 3D characterization of forest structure. In this sense, unmanned aircraft vehicles (UAV) have been proposed to estimate forestry variables (Giannetti et al., 2018; Kachamba et al., 2017; Puliti et al., 2017), characterize forest fuels (Fernández-Álvarez et al., 2019), detection of canopy gaps (Bagaram et al., 2018) or tree-stump (Puliti et al., 2018) mainly for small geographical areas (Shin et al., 2018). Furthermore, the derivation of 3D point clouds from ortophotography may increase 3D data availability and improve the subsequent estimation of forestry variables, especially those that describe canopy height (Noordermeer et al., 2019). ALS have a broad range of applicability within forest management, as for example estimation of forest inventory variables (Guerra-Hernández et al., 2016a; Montealegre et al., 2016), fire-induced change quantification (McCarley et al., 2017) or characterization of forest structural diversity (Kane et al., 2011). ALS point density refers to the number of points per square meter, denoting the spatial resolution. There are two approaches within LiDAR literature to estimate forest variables: the individual tree-based approach (ITB) and the area-based approach (ABA) (Latifi et al., 2015). The ITB approach generally involves a sequence of tree detection, feature extraction, and estimation of tree variables (Maltamo et al., 2014). Furthermore, it also requires field measurements at tree level. Several algorithms have been proposed for tree segmentation as well as feature extraction (raster-based, point based or multisource-based), providing different accuracies under different forest complexities (Sačkov et al., 2019). The ITB approach generally requires point densities higher than 4-5 points m-2 (Andersen et al., 2006). The ABA approach was created by Næsset (1997) and it is also known as the two phase approach inventory or double sampling inventory (Næsset, 2002). The first phase consist on determining the relationships between ALS metrics and the forest stand variables estimated using field data measurements. The second phase fits models that are subsequently extrapolated to the whole study area. The ABA approach has been widely implemented for the estimation of forestry Characterization of Mediterranean Aleppo pine forest using low-density ALS data 6 variables (González-Ferreiro et al., 2013; Latifi et al., 2010; Noordermeer et al., 2018), being a suitable approach for applications using low point density datasets, as the case of the ALS data from the National Plan for Aerial Ortophotography (ALS-PNOA data) in Spain. The effort made by countries and organizations (i.e.: Finland, Spain, Czech Republic) to provide open low density ALS data at regional and national scales creates new opportunities for forest management. The analysis of the processing methodology of low density ALS-PNOA data in Mediterranean forests has been addressed by Montealegre et al., (2015a and b) who compared several filtering and interpolation routines to assess the most suitable methods to work within Aleppo pine forested areas. The generation of models requires the selection of the most suitable ALS metrics and regression methods. Variable selection, also known as feature selection, constitutes a relevant step in modelling generation. The recently increase in size of datasets, with tens, hundreds or thousands of available variables, increased the research interest on selection techniques (Guyon & Elisseeff, 2003). This growth in the availability of variables generates a dimensionality problem, typical in many fields of science (Mehmood et al., 2012), which refers to the existence of a higher number of variables than samples. These problem is also known as large p small n problem (Martens et al., 1992). The large number of ALS metrics that are derived from the point cloud has substantially increased the number of variables in forestry modelling, being sometimes even higher than the number field plots sampled. Dealing with large feature sets presents several disadvantages such as technical and model decrease of accuracy. Although hardware and software have improved in the last decades, the use of a large number of variables takes too many computational resources and slows down regression algorithms. Furthermore, in accordance with Kohavi & John (1997), there may be a decrease of model accuracy when the number of variables is significantly higher than the optimal. Concerning variable selection, a variety of methods exists in order to reduce the dimensionality problem as well as to deal with large feature sets. According to Guyon & Elisseeff (2003) variable selection process have three objectives: improve the prediction performance of the predictors, reduce time and cost when determining the predictors and provide a better understanding of the generated models. A similar definition was proposed by Peña (2002), who established that this selection process should follow the principle of parsimony; filtering or reducing the predictor variables in order to generate models as simple as possible, while maximizing their power. Variable selection constitutes the first step in model fitting. The regression analysis constitutes the process to estimate or model a relationship between variables. This analysis can be categorized in two main types: parametric and non-parametric regression. The classical regression approach is the parametric one, which assumes the existence of a finite set of parameters. On the contrary, nonparametric methods consider that data distribution cannot be defined by a finite set of parameters. Parametric regression relies on strong assumptions such as normality, homoscedasticity, linearity, independence, no-collinearity and absence of atypical values. In non-parametric models, meeting Introduction 7 these requirements is not necessary, turning them into a more flexible tool. There exists a wide variety of parametric and non-parametric methods; from simple linear regression models to complex neural networks models. According to Hazelton (2015), non-parametric methods are divided in kernel and local polynomial regression, spline-based regression or regression trees. The use of machine learning algorithms, showed good performance in several research fields as data mining. Machine learning algorithms, defined as algorithms and statistical models that do not require specific instructions and whose form of the function is unknown, are not directly associated to the two established types of regressions. Parametric regression have been traditionally used in forestry for stand variable prediction with ALS data (Penner et al., 2013). Recently, the application of non-parametric and machine learning methods to this topic has increased popularity (Bollandsås et al., 2013b; Chirici et al., 2008; Liaw & Wiener, 2002). Model accuracy varies according to several factors such as field and ALS data characteristics, forest complexity, variable selection and regression method used, between others. In this sense, ALS flight configuration determines the final data characteristics, as the quality of the point cloud, affecting model performance. ALS sensors are configured with different scanning patterns, pulse frequencies, scanning frequencies, scanning angles and beam divergence. This configuration, summed up to the flight altitude and speed produce different footprint sizes and point densities. All these settings are normally different from one flight to another. Several authors have explored the effect of some of these settings on the prediction error in forest attributes modelling. Yu et al. (2004) concluded that an increase in flight altitude decreases accuracy in trees detection and tree height prediction, affecting more to deciduous species than coniferous ones. The effect of footprint size has received little attention in the literature, however several authors agreed that, in small footprint applications, the largest footprints generate a higher bias in the prediction of tree heights (Andersen et al., 2006; Hirata, 2004; Roussel et al., 2017). An increase of pulse frequency, defined as the number of pulses per second and expressed in kilohertz (kHz), generates a lower penetration of the pulses through the canopy (Chasmer et al., 2006b; Næsset, 2009), implying a decrease in tree height prediction accuracy (Chasmer et al., 2006a). The effect of scan angle is relatively low up to ~20°, producing considerably higher prediction errors in forestry metrics with higher values (Liu et al., 2018; Montaghi, 2013). According to Andersen et al. (2006), the beam divergence also affect tree height predictions; a decrease in the angle increases accuracy. Furthermore, the analysis of the effect of point density determines that a decrease of this factor generally does not produce a decrease in accuracy (Garcia et al., 2017; Roussel et al., 2017). Finally, some environmental variables such as slope have been considered a source of error in ALS processing in forested areas. The presence of steep slopes generally increase the errors in point cloud filtering and interpolation processes (Montealegre et al., 2015b). Forest structure refers to the size, shape, and horizontal and vertical distribution of leaves, branches and stems. These characteristics vary along time, being affected by natural or anthropic disturbances. Fire is caused by natural factors such as volcanic eruptions or lightning and has historically transformed the landscape. Traditional human activities used fire to manage different Characterization of Mediterranean Aleppo pine forest using low-density ALS data 8 land uses, providing an anthropogenic dimension of wildfires. These disturbances constitute some of the most important socio-environmental hazards in Mediterranean forest ecosystems, that might be enhanced by climate change, since extreme meteorological conditions or long droughts increase the fire risk (González-de Vega et al., 2016; Sebastián-López et al., 2008). Although statistic registers showed a reduction in the number of fire events during the last decade (2001-2010), the occurrence of large fires (>500 ha) in Spain has increased (Rodrigues et al., 2014). Aleppo pine (Pinus halepensis Miller) is a flammable species, frequently affected by wildfires, characterized by a high stand density and a continuous presence of branches along the stem (Pausas et al., 2008). Pine forests have a high resilience to fire, but their regeneration process might fail when fire recurrence is high. Forest fires have important effects in vegetation dynamic and atmosphere, as may emit large quantities of greenhouse gases (GHGs) and represent an important carbon sink (Akagi et al., 2013; van der Werf et al., 2010; Wiedinmyer et al., 2011). The quantification of wildfire carbon dioxide (CO2) emissions is vital for climate regulation policies (Mieville et al., 2010) as well as for highlighting the service that forests provides to societies (Lal, 2008). The estimation of fire GHGs emissions requires pre-fire biomass estimation , the assessment of the fraction of biomass consumed by fire, usually related to fire severity, and, subsequently, the use of conversion factor to estimate GHG emissions (De Santis et al., 2010). The use of passive remote sensing to estimate fire severity have been broadly analysed in the literature (GarcíaLlamas et al., 2019). Thus, different indexes have been proposed to account for fire damage in vegetation as Normalized Burn Ratio (Key & Benson, 2005), Relative delta Normalized Burn Ratio (Miller & Thode, 2007), between others. Furthermore, the use of ALS data to quantify biomass have been tested for different ecosystems (García et al., 2010; Næsset, 2011). The estimation of forestry inventory variables is one of the most relevant application for forestry purposes (Latifi et al., 2010; Montealegre et al., 2016). As mentioned above, forest ecosystems constitute important carbon sinks and play a major role in managing GHGs emissions. In this sense, the estimation of biomass and carbon content has growing interest. Several studies have explored the estimation of aboveground tree biomass using ALS data in different ecosystems (Guerra-Hernández et al., 2016b; Mauya et al., 2015). However, the estimation of some biomass fractions, such as shrub biomass or forest residual biomass, have been less studied (Estornell et al., 2012; Hauglin et al., 2014). The presence of understory in forested areas and the existence of shrubland areas are very common in the Mediterranean basin land cover. Thus, the quantification of these biomass fractions may improve the account of carbon reservoirs. Furthermore, the use of some biomass fractions for bioenergy purposes may reduce the CO2 emissions to the atmosphere produced by other fuels and might help to reach the climate and energy targets of the European Energy Roadmap (Hamelin et al., 2019). Finally, the management of these fractions have several benefits for rural development, as the reduction of wildfire risk or the emergence of new business opportunities for forestland owners (Hauglin et al., 2012). ALS data provides a wide range of applications when multi-temporal data is available. Forest managers could use multi-temporal ALS data for applications such as: characterizing natural or Introduction 9 anthropic changes, determining under sampled areas and selectively add plots for future inventories, applying existing ALS-based models to subsequent acquisitions in forests with similar characteristics, reducing field work (Fekety et al., 2015) and improving the accuracy of long period forestry trend analysis using fusion of active and passive optical data. Despite the great potential of multi-temporal analysis, its application is still limited by the acquisition costs as well as the need of temporal-concomitant field data (Cao et al., 2016; Dubayah et al., 2010; Ferraz et al., 2018). The recent effort made by countries and organizations to provide multi-temporal datasets creates new opportunities for upgrading forest inventories at local and regional scales. Local scale refers to small areas whose stand characteristics are similar while regional scale determine large areas with higher stand and structural variability. In addition, several authors have estimated height growth (Gatziolis et al., 2010; Socha et al., 2017) as well as biomass and carbon dynamics (Hudak et al., 2012; Poudel et al., 2018). The estimation of volume (Næsset & Gobakken, 2005; Poudel et al., 2018; Yu et al., 2008), basal area (Næsset & Gobakken, 2005) and site index (Noordermeer et al., 2018) as well as the quantification of wildfire changes (McCarley et al., 2017) and gaps presence (Vepakomma et al., 2008) or the analysis of defoliation effect (Solberg et al., 2006) have also been performed. Two approaches, direct and indirect, have been proposed to model forest attributes using multi-temporal ALS data over time (Bollandsås et al., 2013). The direct approach fits one model for one point in time and temporally transfers the model to other point in time. The indirect one fits two different models for each point in time. The exploration of these approaches provides useful information to forest managers for the reduction of field data acquisitions (Noordermeer et al., 2018). 1.2. Importance and justification LiDAR technology, and specifically airborne laser scanners, has become a valuable source of 3D information to characterize forest ecosystems. Different products can be derived using ALS data, from the generation of precise DTM and digital surface models (DSM) to the characterization of forest stands variables. Thus, the combination of ALS data with field work as well as with data from other remote sensing sources provides accurate information to quantify and evaluate forest resources. Wildfires constitute a socio-environmental hazard in Mediterranean ecosystems. These events are caused by natural or anthropogenic factors and constitute a relevant source of greenhouse gases emission to the atmosphere. Aleppo pine, being a pyrophyte species, is one of the most affected by fires in Spain. The estimation of fire emissions requires the quantification of pre-fire biomass and the fraction of biomass consumed by the fire. The traditional estimation of pre-fire biomass was performed with field data campaigns, while the estimation of post-fire biomass was based in visual examination or field-based weighting. In this sense, the use of remote sensing tools to determine fire severity, and subsequently, burn efficiency has been widely analysed. The capabilities of ALS data to describe forest structure and quantify forest biomass and the advantages provided by its fusion with optical passive data, might enhance quantification of greenhouse gases emissions sourced from wildfires. 17 2. Study area, materials and methods This chapter describes the study area, which includes four different zones within Aragón region, as well as the material and methods utilized in the research. Firstly, we describe the field inventory data, the allometric equations used for computing the different forest variables and the preprocessing performed to the data. Secondly, the remote sensing information used; the ALS-PNOA data and passive optical images, are presented, including the pre-processes applied. Thirdly, we include the selection and regression methods utilized for modelling different variables at stand level. Then, we define the methods to analyse the temporal transferability of multi-temporal ALS data. Furthermore, the methods utilized to assess the effect of ALS parameters and environmental conditions in model accuracy are presented. Finally, the mapping process and its importance in the generation of information at local or regional scale is addressed. Study area, materials and methods 19 2.1. Study area Aleppo pine is the most broadly distributed species from genus Pinus in the Mediterranean basin. The regions with higher presence are the north of Africa and Spain, which represents ~850,000 ha, according to Cámara (2001). The wide altitudinal and latitudinal gradient of this distribution allows this species to live from sea level up to 1,600 m in the Saharan Atlas (Cabanillas 2010). Although Aleppo pine is a limestone species, it grows in different types of soils such as siliceous, quartzite or granite. Furthermore, the species is heliophilous, thermophile, xerophile and pyrophyte, being adapted to droughts and wildfires. This PhD Thesis studied the Aleppo pine forest of Aragón region. Aragón is an Autonomous Community located in the northeast of Spain. This region occupies 47,720.3 km2 and represents 9.4% of the Spanish territory. Three provinces; Zaragoza, Huesca, and Teruel, constitute this Autonomous Community. Aragón limits to the north with France, to the east with Cataluña and Valencia and to the west part with Castilla-La Mancha, Castilla y León, La Rioja and Navarra. Three main relief units compose Aragón, the Pyrenees, the Ebro Basin and the Iberian Ranges. The Pyrenees are represented by the Axial Pyrenees, the interior mountains, the Intrapyrenean topographic depression and the exterior mountains (pre-Pyrenees). Axial Pyrenees, constituted by granites, quartzite, slates and limestone, presents the highest altitudes such as Aneto (3,404 m) or Maladeta (3,308 m). The interior mountains include calcareous crests adhered to the axial Pyrenees. The Intrapyrenean topographic depression contains several perpendicular river valleys, ending in “San Juan de la Peña” and “Peña Oroel” conglomeratic escarps. The pre-Pyrenees, constituted by calcareous rocks, present heights between 1,500 and 2,000 m (Peña & Lozano, 2004). The “Somontanos” connect the pre-Pyrenees with the left bank of the Ebro Basin. The Ebro Valley was a sea during Mesozoic and Eocene stages. Nowadays it is a topographic depression, covered by tertiary materials and alluvial sediments. The erosion processes have generated different tabular reliefs from 500 up to 800 m in both banks of Ebro River. The Iberian Ranges is a mountain chain of lower altitude than Pyrenees. The “Sierra del Moncayo” constitutes the northwest part, “Puertos de Beceite” and “Gúdar-Maestrazgo” are located in the eastern part, while “Javalambre” and “Albarracín” in the southeastern part. Quartzite and Palaeozoic slate are the main materials present in the higher mountains, while Jurassic and Cretaceous limestone and dolomites constitute the lower relief structures. The climate of Aragón is Mediterranean with continental features, characterized by cold winters, dry summers and irregular and scarce rainfall. The variability in orography modifies the temperatures and precipitations, generating a wide diversity of climatic ambient from semidesertic areas such as Monegros to high mountains in the Pyrenees. According to Cuadrat (2004), Aragón climate is defined by four main characteristics. The Ebro valley presents low annual precipitation values, ~300 mm m2, generated by the shadow effect of Pyrenees and Iberian Ranges; its location in a continental area generates a broad range of temperatures; precipitations are Characterization of Mediterranean Aleppo pine forest using low-density ALS data 20 irregular and the winds come from the northwest in winter and southeast in summer, being frequent and heavy. Aragón includes two Holarctic biogeographical regions: Eurosiberian and Mediterranean. The Eurosiberian region is occupied by forests and grasslands distributed in three altitudinal strata: alpine, subalpine, and montane. The Mediterranean region includes the Ebro Valley, the “Somontanos”, the river valleys located in the right margin of Ebro River and the topographic depression in which Teruel city is situated. Quercus ilex, Pinus halepensis, Pinus nigra and Juniperus sabina constitute the species that dominate Mediterranean forests (Longares, 2004). According to the Spanish Forest Map, the forested area in Aragón represents 1.58 million of ha, of which 259,057.45 ha are occupied by Aleppo pine forests. Concretely, 124,473.12 ha are located in Zaragoza, 37,817.19 ha in Huesca and 96,767.14 ha in Teruel regions. In the whole, semi-natural forests represent 211,013.49 ha and afforested forest 48,043.96 ha. As mentioned before, four study zones were delimited to provide answers to the specific research objectives (Figure 1). The study zones are described below:  Zone A is located in “Las Cinco Villas” region, northwest of Aragón. The relief is characterized by a topographic depression close to the pre-Pyrenees. Elevations range from 430 to 1150 m above sea level and slopes from 0° to 39°. The climate is Mediterranean with continental features with an annual average precipitation of 525 mm. The Zone A includes two different test areas: “Luna” wildfire and field inventory 3 (Figure 3). Luna wildfire was caused by agricultural machinery on 4th July 2015. The fire scorched 14,263 ha, 3,390.4 ha covered by woodland. Field inventory 3 is located in an unburned area close to the wildfire, with similar environmental, climatic and forest characteristics. The unburned Aleppo pine forest is heterogeneous from the structural point of view, being accompanied by an evergreen understorey with species such as Quercus ilex subsp. rotundifolia, Quercus coccifera, Juniperus oxycedrus, Buxus sempervirens and Juniperus phoenicea.  Zone B is located in the Ebro Basin, Northeast Spain. This zone includes two different test areas: inventory 2 (Figure 3) and inventory 3. Both areas are representative of Aleppo pine Mediterranean forest, occupying 11,400 ha. The inventory 3 area was described above. The majority of Inventory 2 plots are located in the military training center “San Gregorio” (CENAD), located in the north of Zaragoza city. The area presents a hilly topography with elevations ranging from 300 to 750 m and slopes from 0° to 39°. Climate is Mediterranean with continental features, while the average annual precipitation is lower than 350 mm. Most of the pine stands are semi-natural, while the stands located in the south eastern part were planted approximately forty years ago. The evergreen understorey is characterized by xerophiles species such as Quercus coccifera, Juniperus oxycedrus, Rosmarinus officinalis and Thymus vulgaris. Study area, materials and methods 21 Figure 3. Location of forest inventory campaigns. High spatial resolution orthophotography from Spanish National Plan for Aerial Orthophotography spatial data infrastructure (SDI) included as backdrop. Characterization of Mediterranean Aleppo pine forest using low-density ALS data 22  Zone C is located in the Ebro Basin and Iberian Ranges, including a broad part of Aleppo pine forest in Aragón region, except from the forested areas close to the pre-Pyrenees and the stands located in the south of Teruel. This zone includes three different test areas: inventory 1 (Figure 3), inventory 2 and most of the plots from inventory 4. Inventory 2 was described previously. Inventory 1 is located in Daroca Municipality in the Iberian Range. The area includes two forests denominated “Dehesa de los enebrales” and “Valdá y Carrilanga”, covering 1,102 ha. Both forests were afforested from 1908 to 1979, being occupied by a monospecific Aleppo pine forest with scarce understory. The area presents a hilly topography, with elevations ranging from 860 up to 980 m. Inventory 4 includes several stands from the Middle Ebro Basin up to the Iberian Ranges. The stands were afforested approximately forty to sixty years ago, keeping a low presence of hardwood species. The different environmental sample conditions include a great variability of Aleppo pine forests.  Zone D represents 197,951.24 ha of the Aleppo pine forested area of Aragón, including all the species distribution range. The Zone D includes four different test areas: inventories 1 to 4 (Figure 3), which have been previously described. The broad area shows a variability of geomorphological forms from topographic depressions up to mountains that reach more than 2,000 m above sea level. Consequently, temperature changes across the altitudinal gradient and the annual precipitation range from less than 350 mm up to 1,000 mm. This variability and the presence of semi-natural and afforested stands are characteristic of Aleppo pine forest at Aragón region. 2.2. Materials and methods 2.2.1. Field inventory data Field inventories provide an accurate quantification of forest variables within a small fraction of the study area. Field plot information is the ground-truth evidence, constituting the dependent variables that are estimated for a broad area. In this sense, field plot data must be representative for the study area, capturing the maximum variability to minimize extrapolation errors. Sampling design is a relevant factor to capture forest stand variability. Several sampling types exist, such as systematic sampling, random sampling or stratified random sampling. Traditional inventories, performed using field campaigns, require the division of forest stands into homogeneous strata, considering factors such as tree species, site productivity, forest development stage, latitude, elevation and stand structure. Thus, performing inventories using ALS data is easier as this technology provides information about some of those factors at a stand level. This PhD Thesis applied a stratified random sampling, considering terrain slope, canopy height and canopy cover variability, according to Næsset & Økland (2002), in monospecific Aleppo pine forest. Field data was acquired in 192 plots in four campaigns performed during 2013, 2014, 2015 and 2016 (hereinafter mentioned as first, second, third and fourth campaign, respectively). Field data were related to ALS data using an area-based approach (Næsset & Økland, 2002). The Study area, materials and methods 23 number of plots allowed the estimation of forest stand variables at local and regional scales. The specifications of each inventory are described below:  The acquisition of field data from the first campaign was performed from June to July 2013, in 53 circular plots within the Master Thesis of Jesús Cabrera (Cabrera, 2013). The center point of each circular plot with 15 m radius was positioned using a Leica VIVA® GS15 CS10 real-time kinematic GNSS with a planimetric accuracy of 0.30 m. We used a diameter tape, with millimeter precision, for measuring tree diameter at breast height (dbh) in those trees with a dbh larger than 7.5 cm, which is the standard dbh for inventoried trees in Spain. A Suunto® hypsometer was used for measuring green crown height and tree height of up to 4 randomly selected trees within each plot. The selection of the sample trees, from 7.5 cm up to 42.5 cm, considers the diametric classes defined as representative for the study area in the third national forest inventory. The height for those trees not measured in the field plots was predicted by using a height-diameter model developed from the sampled trees (equation 1). The model performance of the height model gave a RMSE of 1.36 m and R2 of 0.63. Normality, homoscedasticity and independence or no auto-correlation in the residuals were verified for the fitted model. ℎ𝑡=0.776·𝐺0.179·𝑑𝑏ℎ𝑖0.660·1.009 (1) where ht is tree height (m), dbhi is the diameter at breast height (cm) and G is field plot basal area (m2 ha-1).  The second and third field campaigns use the same inventory methodology. The second campaign includes 43 circular plots acquired from July to September 2014 (Montealegre et al., 2016). The third field plot campaign sampled 45 plots from June to July 2015, being used for objectives 1, 2 and 4. The same GNSS instrument used for the first campaign positioned the 30 m diameter plots, obtaining a planimetric accuracy of 0.15 and 0.18 m in 2014 and 2015, respectively. A Haglöf Sweden® Mantax Precision Blue diameter calliper allowed measuring the dbh of those trees with a dbh larger than 7.5 cm. We used a Haglöf Sweden® Vertex instrument to measure the green crown height and the height for all trees in the plot. Furthermore, the percentage of shrub canopy cover and the average height of the different shrub species that represent the understory were measured (used for objective 2).  The fourth campaign inventoried 51 field plots in April 2016 being carried out by föra forest technologies within the project RF-64079 (Rural Development Program of Aragón 2014–2020). A Trimble submetric GNSS was used to position the center of each plot with a submetric accuracy in planimetry. A variable plot radius was selected (5.6 m, 8.5 m, 11.3 m, and 14.10 m), in order to obtain data from a similar number of trees in each plot, due to the difference in stand density. A Haglöf Sweden® Mantax Precision Blue diameter calliper allowed measuring those trees with a dbh larger than 7.5 cm. A Haglöf Sweden® Vertex was used for measuring the green crown height and the height of up to 6 trees, the nearest to the plot center. The sample was completed to achieve 100 dominant stems ha-1, Characterization of Mediterranean Aleppo pine forest using low-density ALS data 24 considering those with larger dbh. The height for those trees not measured in the field was estimated by using a height-diameter model developed from the sampled trees (equation 2). The model performance of the height model gave a RMSE of 0.80 m and R2 of 0.93. ℎ𝑡=(1.32.5511+(𝐻02.5511−1.32.5511)·1−𝑒𝑥𝑝(−0.025687·𝑑𝑏ℎ) 1−exp(−0.025687·𝐷0))12.5511 ⁄ (2) where ht is tree height (m), dbh is the diameter at breast height (cm), Ho is the Assmann dominant height (m) and Do is the Assman dominant diameter (cm). 2.2.2. Estimation of forest stand variables The estimation of forest variables at tree level requires the use of allometric equations. These equations generally need a destructive sampling, drying and subsequent weighting of a representative sample. The majority of tree forest species have an allometric equation. However, the accessibility of shrub allometric equations in Spain is more limited, being available for the estimation of some specific variables such as biomass (Montero et al., 2013). This PhD Thesis estimated nine stand variables using allometric equations as a ground-truth. The estimated variables are: stand density (N), basal area (G), squared mean diameter (Dg), dominant diameter (Do), dominant height (Ho), timber volume over bark of stem (V), above ground biomass (Wag), total tree biomass (W) and forest residual biomass (FRB). In addition, the equations determined by Montero et al. (2013) for different shrub formations defined in the Spanish Forest Map (MFE) were used to calculate shrub biomass. Finally, total biomass including aboveground tree biomass and shrub biomass (TW) was computed by summing up total tree biomass and shrub biomass. Tree equations The equations 3 to 16 were used to estimate the above mentioned tree variables. 𝑁 (𝑠𝑡𝑒𝑚𝑠 ℎ𝑎−1)=𝑠𝑡𝑒𝑚𝑠 𝑎 (3) 𝐺 (𝑚2 ℎ𝑎−1)=∑𝜋 4∙𝑑𝑖2𝑛 𝑖𝑎 (4) 𝐷𝑔 (𝑐𝑚)=100·√4∙𝐺 𝜋∙𝑁 (5) 𝐷𝑜=∑𝑑𝑖 𝑘 𝑖 𝑎/100 (6) 𝐻𝑜=∑ℎ𝑖 𝑘 𝑖 𝑎/100 (7) 𝑣𝑖=𝜋 40000∫[(1+1.121163·𝑒(−10.23293·ℎ𝑖 ℎ𝑡))·0.696362·𝑑𝑖 ℎ𝑡 0·((1−ℎ𝑖 ℎ𝑡)1.266261−(0.003553∗𝐸)−1.865418·(1−ℎ ℎ𝑡))]2𝑑𝑖ℎ𝑖 𝑉 (𝑚3ℎ𝑎−1=10000·∑𝑣𝑖 𝑛 𝑖𝑎 (8) Study area, materials and methods 25 Wag (kg ℎ𝑎−1)=𝐶𝐹∙𝑒𝑎∙𝑑𝑖𝑏 𝑎∙10,000 (9) Ws (kg)=0.0139 ∙ 𝑑𝑖2∙ℎ𝑖 (10) Wb7 (kg)= [3.926∙(𝑑𝑖−27.5)]∙Z; (11) If 𝑑𝑖≤27.5 cm then Z=0;If 𝑑𝑖>27.5 𝑐𝑚 𝑡ℎ𝑒𝑛 𝑍=1 Wb2−7 (kg)=4.257+0.00506∙𝑑𝑖2∙ℎ𝑖−0.0722∙𝑑𝑖∙ℎ𝑖 (12) Wb2+n (kg)=6.197+0.00932∙𝑑𝑖2∙ℎ𝑖−0.0686∙𝑑𝑖∙ℎ𝑖 (13) Wr (kg)=0.0785∙𝑑𝑖2 (14) W (kg)=𝑊𝑠+𝑊𝑏7+𝑊𝑏2−7+𝑊𝑏2+𝑛 (15) FRB (kg)=𝑊𝑏7+𝑊𝑏2−7+𝑊𝑏2+𝑛 (16) where d is the normal diameter in cm; a is the area of the plot expressed in m2; ∑n refers to the number of trees inside of a plot; ∑k refers to the number of k trees being k the thickest trees; hi is the height of the trees; ht is the total tree height; E is the tree slenderness (ht /di); vi is the volume of each tree in m3; CF is a correction factor (𝐶𝐹=𝑒𝑆𝐸𝐸2/2) being e the Euler number and SEE the standard error (0.151637); a is (−2.0939) and b (2.20988) are the specific parameters for Aleppo pine; Ws is the biomass weight of the stem fraction, Wb7 is the biomass weight of the thick branch fraction (diameter larger than 7 cm), Wb2−7 is the biomass weight of medium branch fraction (diameter between 2 and 7 cm), Wb2+n is the biomass weight of the thin branch fraction (diameter smaller than 2 cm) with needles, and Wr is the biomass weight of the roots. Shrub equations The shrub equations used for each formation are presented below (equations 17 to 21):  Shrub hedges, borders, galleries, etc.: ln(W𝑠)=0.494×ln (CC) (17)  Quercus coccifera and Pistacia lentiscus: ln(W𝑠)=−2.892+1.505×ln(hm)+0.462× ln (CC) (18)  Leguminosae aulagoideas and related shrubs: ln(W𝑠)=−2.464+0.808×ln(hm)+0.761× ln (CC) (19)  Labiatae and Thymus formations: ln(W𝑠)=−1.877+0.643×ln(hm)+0.661× ln (CC) (20)  General shrub biomass: ln(W𝑠)=−2.560+1.006×ln(hm)+0.672×ln (C) (21) where Ws is the biomass weight for each species in tons/ha, hm is the average shrub height at plot level and CC is the percentage of shrub canopy cover at plot level. Equation 17 was applied for Crataegus monogyna, Rhamnus lycioides and Rosa canina; equation 18 was used for Quercus coccifera; equation 19 was applied for Genista scorpius; equation 20 was used Characterization of Mediterranean Aleppo pine forest using low-density ALS data 32  GPS time indicates the date and time of GNSS laser point registration when the plane is flying. The time is expressed in seconds. Table 4. Technical specifications of LiDAR-PNOA plan for Aragón region from the first and second coverage. Characteristic Description First coverage Second coverage Sensor Leica ALS-50 Leica ALS-80 Geodetic reference system ETRS89 Cartographic projection Universal Transversal Mercator (UTM) 30 and 31 Geoid model EGM2008-REDNAP Field of view (FOV) The maximum allowed FOV is 50° Scanning frequency Minimum of 70 Hz and up to 40 Hz with a FOV of 50° Pulse frequency Minimum of 45 kHz with a FOV of 50° and up to 3,000 m Point density 0.5 points m-2 implying a spacing between points ≤1.41 m Radiometric resolution for multiple intensities Dynamic range with at least 8 bits Ability to capture multiple returns for the same pulse Up to 4 returns per pulse when the vertical distance is higher than 4 m GNSS navigation system Double frequency GNSS with at least 2 Hz Inertial system (IMU/INS) Data register frequency ≥200Hz and drift <0.1° h-1 Plane velocity when LiDAR data capturing Variable Flight height Variable Maximum flight line length 90 km Horizontal and vertical precision after processing The horizontal global precision at nadir have a RMSEx,y <30 cm (1 sigma), while the vertical global position at nadir have a RMSEz <20cm (1 sigma). Errors up to 3xRMSE in dense vegetation or steep slopes might occur. The edge of flight line might present an error up to 2xRMSE Altimetry precision and maximum error ≤0.40 m for 95% of the cases. Points cannot have an error higher than 0.60 m Altimetry discrepancies between flight lines ≤0.40 m Point cloud format *.las 2x2 km tiles LAS file classification Automatically classified (ground, vegetation, buildings, overlap) Non classified Study area, materials and methods 33 Data from both ALS-PNOA coverages are utilized in this PhD Thesis. In this sense, more than 4,000 *.las files from the first coverage were downloaded from the CNIG webpage. Furthermore, the Geographic Institute of Aragón (IGEAR) and CNIG provided 147 tiles from the second coverage before being included in CNIG platform. The flight technical specifications, the point cloud characteristics and the file nomenclature is included in Table 4, Table 5 and Figure 5, subsequently. Table 5. Characteristics of LiDAR-PNOA data according to the study area. Characteristics Zone A Zone B Zone C Zone D Acquisition date October 2010 (first coverage) October 2010, January 2011 and February 2011 (first coverage) July 2010 to February 2011 (first coverage) July 2010 to February 2011 (first coverage) September to November 2016 (second coverage) Sensor Leica ALS50 Leica ALS50 Leica ALS50 Leica ALS50 (first coverage) Leica ALS80 (second coverage) Average flight height (m) 3146 3022 3240 3138 (first coverage), 2943 (second coverage) Average plane velocity (km/h) 241 241 184 184 (first coverage), ~240 (second coverage) Number of tiles 42 690 3800 147 (first coverage) and 147 (second coverage) Figure 5. Nomenclature of LiDAR-PNOA files. Characterization of Mediterranean Aleppo pine forest using low-density ALS data 34 2.2.5. ALS pre-processing and metric computation Quality control of ALS point clouds is required for accurate pre-process the data. The ALS data provided by the PNOA is not raw data, but users should check the quality of downloaded point clouds before performing any further analysis. The first ALS processing step considered was the removal of noise return hits. ALS-PNOA data is classified and noise returns are denoted as class 7. These points were removed using LAStools software implemented in ArcGIS 10.5. Overlapping strips might generate horizontal and vertical discrepancies in laser scanning surveys playing and important role in quality control (Vosselman, 2012). Two approaches could be considered when dealing with overlapping strips: removing those discrepant strips or using methods to compensate for discrepancies between two datasets in the overlapping areas (Latypov, 2002). In this PhD the validity of overlapping returns was verified by visualizing the 3D point clouds, generating reports about ALS characteristics and analysing the generated DEMs. The analysis performed determined that some tiles from the south of Aragón, named as “ARA_SUR” present vertical and/or horizontal displacements. These displacements are classified with the number 12 and, consequently, were removed using the same procedure as noise return removal. The ALS point clouds in forested environments contain returns from several surface objects, such as shrubs, trees, electrical wires and buildings that should be separated from ground returns (Montealegre et al., 2015a). The process of separating the ground and non-ground points, performed prior to DEM generation, is called filtering or classification. This is a key process in forestry applications, allowing normalizing point clouds to determine aboveground return heights. There are several filtering algorithms which can be classified in four types (Meng et al., 2010; Sithole & Vosselman, 2004): interpolation-based, slope-based, segmentation-based and morphological methods. Most filtering methods provide accurate results in flat and non-complex areas while steep slopes and complex forest stands increases classification errors (Sithole & Vosselman, 2004). Previous research developed by Montealegre et al. (2015a) compared seven filtering methods implemented in open software in Aleppo pine forested areas with moderate to steep slopes. The study determined that Multiscale Curvature Classification (MCC) method, developed by Evans & Hudak (2007), was the most accurate to filter ALS-PNOA point clouds in these Mediterranean environments. Accordingly, in this PhD MCC algorithm, implemented in MCC-LiDAR software, was selected for classifying the point clouds. MCC is an iterative-interpolation-based filter. The algorithm calculates an interpolated surface by using a thin-plated spline and discards the ALS returns that exceed a threshold curvature. MCC creates three scale domains defining three processing window sizes. The algorithm iterates in the three scale domains until the number of remaining returns changes by less than 1%, less than 0.1%, and less than 0.01%, respectively (Evans & Hudak, 2007). Two parameters must be defined for running MCC: the scale parameter (s) and the curvature threshold (t). The scale depends on object sizes and ALS point spacing, while the curvature threshold is related to curvature tolerance for a Study area, materials and methods 35 scale domain. In this research, both parameters were determined according to Montealegre et al. (2015b). The scale was set to 1 m and the curvature threshold to 0.3. ALS data create a random sampling of the terrain surface, being necessary to apply interpolation processes to generate a continuous surface (Vosselman & Maas, 2010). The selection of the appropriate interpolation algorithm and DEM spatial resolution is relevant as may constitute a source of inaccuracy in vegetation metrics prediction (Aguilar et al., 2006). Several interpolation methods, such as natural neighbour, Triangulated Irregular Network (TIN) to raster, kriging, point to raster, are commonly implemented in Geographical Information System (GIS) software. Montealegre et al. (2015b) compared the suitability of six interpolation routines over a range of terrain roughness in Aleppo pine forested stands. The analysis determined that the TIN to raster interpolation method produces the best result when generating a DEM with 1 m resolution. Consequently, the TIN to raster routine, implemented in ArcGIS 10.5 software, was selected for interpolating and generating the DEMs in this research. TIN to raster method includes two procedures. Firstly, the interpolation method build a terrain surface using Delaunay irregular triangles by connecting ALS returns. The elevation is recorded for each triangle node, while elevations between nodes can be interpolated to generate a continuous surface. Secondly, the TIN structure is transformed to a raster structure using natural neighbour interpolation from previous triangles nodes to generate a value at the center of each raster cell. The statistical metrics derived from ALS data were related to field variables to fit regression models. The computation of ALS metrics requires the clipping of the point cloud to the spatial extent of each field plot, being performed using the “ClipData” command of FUSION LDV 3.60 open source software (McGaughey, 2009). Furthermore, the height of the point cloud returns were normalized using the DEM to compute a wide range of statistics using the “Cloudmetrics” command (Table 6). These statistical metrics are commonly used as independent variables in forestry (Evans et al., 2009). The computation of ALS metrics normally requires the application of a threshold value to remove ground and/or understorey laser hits. In this PhD Thesis, several tests were performed to determine the more suitable thresholds according to the type of estimated stand variable and application. A threshold value of 2 m height was selected according to shrub height agreeing with Nilsson (1996) and Næsset & Økland (2002) for predicting the analysed forestry metrics except from total biomass, that required the selection of a different threshold to include understorey laser hits. The explored thresholds ranging from 0 up to 1 m determined that 0.2 is the most suitable one being related to the ALS sensor precision in coordinate Z, which has a root mean square error below 0.2 m. Table 6 describes the computed ALS metrics structured in three main groups: metrics related with canopy height, metrics associated to canopy height variability and canopy density metrics. The Characterization of Mediterranean Aleppo pine forest using low-density ALS data 36 computed metrics might be selected to be used in modelling and all of them present a coherent relationship with vegetation structure (Evans et al., 2009; McGaughey, 2009). Table 6. Derived metrics from ALS point clouds, where xi is the height value of the return, N is the total number of observations, ri is the return, and p is the pulse. Metric Description Canopy height metrics (CHM) Percentiles of the return heights 1, 5, 10, 20, 25, 30, 40, 50, 60, 70, 75, 80, 90, 95 and 99 (P01, P05, P10, etc.) The percentiles were computed according to the following methodology: (𝑁−1)𝑃=𝐼+𝑓{ 𝐼 𝑖𝑠 𝑡ℎ𝑒 𝑖𝑛𝑡𝑒𝑔𝑒𝑟 𝑝𝑎𝑟𝑡 𝑜𝑓 (𝑁−1)𝑃 𝑓 𝑖𝑠 𝑡ℎ𝑒 𝑓𝑟𝑎𝑐𝑡𝑖𝑜𝑛𝑎𝑙 𝑝𝑎𝑟𝑡 𝑜𝑓 (𝑁−1)𝑃 where N is the number of observation and P is the percentile value divided by 100. 𝑖𝑓 𝑓=0 𝑡ℎ𝑒𝑛 𝑃𝑒𝑟𝑐𝑒𝑛𝑡𝑖𝑙𝑒 𝑣𝑎𝑙𝑢𝑒=𝑥𝑖+1 𝑖𝑓 𝑓>0 𝑡ℎ𝑒𝑛 𝑃𝑒𝑟𝑐𝑒𝑛𝑡𝑖𝑙𝑒 𝑣𝑎𝑙𝑢𝑒=𝑥𝑖+1+𝑓(𝑥𝑖+2−𝑥𝑖+1) where xi is the observation value considering that observations are ranked in ascending order. Minimum elevation 𝑥𝑖 𝑚𝑖𝑛𝑖𝑚𝑢𝑚 Mean elevation ∑𝑥𝑖 𝑁 𝑖=1 𝑁 Mode elevation xi value more frequent in the plot Elevation quadratic mean (1 𝑁∑𝑥𝑖2 𝑁 𝑖=1 )1 2 Elevation cubic mean (1 𝑁∑𝑥𝑖3 𝑁 𝑖=1 )1 3 L moments (λ1 to λ4) λ1=1𝐶1 𝑛∑𝑥(𝑖) 𝑁 𝑖=1 λ2=121𝐶2 𝑛∑( 𝐶1− 𝐶1 𝑛−1𝑖−1 )𝑥(𝑖) 𝑁 𝑖=1 λ3=131𝐶3 𝑛∑( 𝐶2−2 𝐶1 𝑖−1𝑖−1 𝐶1+ 𝐶2 𝑛−1𝑛−1 )𝑥(𝑖) 𝑁 𝑖=1 λ4=141𝐶4 𝑛∑(𝐶3−3 𝐶2 𝑖−1𝑖−1 𝐶1+3 𝐶1 𝑖−1𝑛−1 𝐶2− 𝐶3 𝑛−1𝑛−1 )𝑥(𝑖) 𝑁 𝑖=1 where x(i), i=1, 2, …, n, are sample values ranked in ascending order and 𝐶𝑘=(𝑚 𝑘) 𝑚=𝑚! 𝑘!(𝑚−𝑘)! is the number of combinations of any k items form m items and is equal to zero when k > m. Maximum elevation 𝑥𝑖 𝑚𝑎𝑥𝑖𝑚𝑢𝑚 Study area, materials and methods 37 Canopy height variability metrics (CHVM) Standard deviation of point height distribution √∑(𝑥𝑖−𝜇)2 𝑁 𝑖=1 𝑁 Variance of point height distribution (σ2) ∑(𝑥𝑖−𝜇)2 𝑁 𝑖=1 𝑁 Coefficient of variation of point height distribution 𝜎 𝜇100 Skewness of point height distribution ∑(𝑥𝑖−𝜇)3 𝑁 𝑖=1 (𝑁−1)𝜎3 kurtosis of point height distribution ∑(𝑥𝑖−𝜇)4 𝑁 𝑖=1 (𝑁−1)𝜎4 Interquartile distance of point height distribution [𝑃75(𝑥)−𝑃25(𝑥)] Average Absolute Deviation of point height distribution ∑(𝑥𝑖−𝜇) 𝑁 𝑖=1 𝑁 L moment coefficient of variation of point height distribution (𝜏2) λ2 λ1 0< 𝜏2<1 L moment skewness of point height distribution (𝜏3) λ3 λ2 -1< 𝜏3<1 L moment kurtosis of point height distribution (𝜏4) λ4 λ2 14(5τ32−1)≤τ4<1 Canopy density metrics (CDM) Percentage of first returns above a height-break, above the mean or the mode ∑𝑟𝑖 𝑓𝑖𝑟𝑠𝑡 𝑟𝑒𝑡𝑢𝑟𝑛𝑠>ℎ𝑒𝑖𝑔ℎ𝑡𝑏𝑟𝑒𝑎𝑘 𝑁 𝑖=1 ∑𝑟𝑖 𝑓𝑖𝑟𝑠𝑡 𝑟𝑒𝑡𝑢𝑟𝑛𝑠 𝑁 𝑖=1 100 Percentage of all returns above a height-break, above the mean or the mode ∑𝑟𝑖 >ℎ𝑒𝑖𝑔ℎ𝑡𝑏𝑟𝑒𝑎𝑘 𝑁 𝑖=1 𝑁100 Canopy relief ratio 𝜇−𝑥𝑖 𝑚𝑖𝑛𝑖𝑚𝑢𝑚 𝑥𝑖 𝑚𝑎𝑥𝑖𝑚𝑢𝑚−𝑥𝑖 𝑚𝑖𝑛𝑖𝑚𝑢𝑚 All returns above a heightbreak, above the mean or the mode x 100 ∑𝑟𝑖 𝑎𝑙𝑙 𝑟𝑒𝑡𝑢𝑟𝑛𝑠>ℎ𝑒𝑖𝑔ℎ𝑡𝑏𝑟𝑒𝑎𝑘 𝑁 𝑖=1 ∑𝑟𝑖 𝑎𝑙𝑙 𝑟𝑒𝑡𝑢𝑟𝑛𝑠 𝑁 𝑖=1 100 2.2.6. Optical data and spectral indices Although the PhD Thesis mainly focus on the use of ALS-PNOA data for forestry applications, the combination with optical data allowed to estimate fire severity and, subsequently, determine the CO2 emissions. Characterization of Mediterranean Aleppo pine forest using low-density ALS data 38 Wildfires constitute a socio-environmental hazard in pyrophyte Aleppo pine Mediterranean ecosystems. Fire severity usually reaches high levels, generating important changes in vegetation vigour, colour, water content as well as forest composition, structure and density. Optical remote sensing data have been widely used to characterize fire severity (García-Llamas et al., 2019) considering infrared spectral regions. Landsat program is the longest middle resolution remote sensing program and provides coherent and continuous global data since 1972. Specifically, Landsat 8 was launched in 2013 carrying on board the Operational Land Imager (OLI) sensor. The mission presents a polar low earth orbit and 16 days of temporal resolution. The spatial resolution of the 9 optical spectrum bands is 30 m, except from the panchromatic band with 15 m resolution. The most valuable bands to characterize fire severity are number 5, which corresponds to the near infrared (NIR), and 7, which refers to the short wave infrared (SWIR). NIR is linked to foliar structure and morphology while SWIR is related to vegetation and soil water content (Soverel et al., 2010). In this research, two images were selected to characterize pre and post-fire conditions and estimate fire severity. The pre-fire image was acquired on June 30 (path 200, row 31) 2015 and the post-fire image on July 9 (path 199, row 31) 2015. The selection of these images was conditioned by the need for temporary proximity between images in order to capture the variability of vegetation spectral response. The images were provided by the United States Geological Survey (USGS) and downloaded using Earth Explorer platform. Both images present a high processing level. Images were geometrically and radiometric corrected. Specifically, the Landsat Surface Reflectance Code (LaSRC) was applied according to Vermote et al. (2016). The algorithm uses a radiative transfer model, climate data from MODIS and the coastal aerosol band to perform those corrections. The use of spectral indices, derived from the combination of different spectral bands, provides relevant information about surface properties (Chuvieco, 2010). The Normalized Burn Ratio (NBR) index, developed by Key & Benson (2006), relates Landsat band 5, from NIR, with band 7 from SWIR. NBR values range from -1 to 1. Surfaces with no presence of vegetation or low vegetation vigour present negative values while photosynthetically active areas show positive values. In this PhD Thesis, NBR was calculated for pre-fire and post-fire images according to equation 27. Then, the Differenced or Delta NBR index (∆NBR) was computed to provide a quantitative measure of the burned area (equation 28). NBR= (ρNIR − ρSWIR )/ (ρNIR + ρSWIR) (27) ∆NBR= NBRprefire – NBRpostfire (28) where  NIR (near infrared) and  SWIR (short wave infrared) refer to bands 5 and 7 Landsat 8 OLI reflectance, respectively. The ∆NBR values were multiplied by 1,000 to provide a valid continuous range of values. Negative values are associated with fast regrowth from herbaceous, while positive values are related to different degree of fire severity. Study area, materials and methods 39 2.2.7. Wildfire biomass loss and CO2 emission estimation Wildfires constitutes a relevant source of carbon monoxide emissions (Pétron et al., 2004) in the Mediterranean basin, which yearly records an average of 45,000 fires (Oliveira et al., 2012). The account of carbon dioxide emissions is relevant for understanding carbon cycling and provides information to develop climate regulation policies (Mieville et al., 2010). Thus, different approaches have been tested to estimate the emissions for wildfire as the one proposed by the Intergovernmental Panel on Climate Change (IPCC) (equation 29): 𝐿𝑓𝑖𝑟𝑒=𝐴×𝐵×𝐶×𝐷×10−6 (29) where Lfire refers to the quantity of GHG released due to fire (tonnes of GHG), A is the burned area (ha), B is the mass of “available” fuel including biomass, ground litter and dead wood (tons ha-1), C is the combustion efficiency or fraction of combusted biomass (dimensionless), D is the emission factor (g kg-1 of dry matter burnt). A similar approach, proposed by De Santis et al. (2010), adapted the IPCC one to estimate fire GHG emissions, using remote sensing techniques. In this sense, four steps were required: the delimitation of the burned area; the estimation of pre-fire biomass; the assessment of the fraction of biomass consumed by the fire or burning efficiency, associated with fire severity and; the use of conversion factors to estimate GHG emissions. Traditionally the estimation of pre-fire biomass was performed by using field data and allometric equations, while post-fire biomass was assessed either by visual examination (Roy et al., 2005) or field-based weighting (Sá et al., 2005). However, the implementation of remote sensing techniques provided new methodological approaches. In this research, ALS data (used to determine pre-fire biomass) and passive optical data (applied to estimate fire severity) were combined with the final objective of estimating biomass losses and, subsequently CO2 emissions. The following five steps were conducted in this approach: i. Estimation of Aleppo pine pre-fire biomass using ALS data captured previously to the occurrence of the wildfire. ii. Estimation of fire severity. ∆NBR index is computed using two Landsat 8 OLI images from June 30th and July 9th 2015. iii. Mapping of pre-fire Aleppo pine forest. The Spanish National Forest Map as well as a canopy height model derived from the ALS data were used to delimit the burned stands. iv. Selection of burning efficiency factors and biomass losses quantification. Most approaches consider that forest biomass is completely consumed by fire (French et al., 2004). However, more accurate approaches from our point of view consider different burn severity levels, related to different biomass losses. Thus, in this research three burning efficiency factors were applied to three different severity levels (De Santis et al., 2010). Key & Benson (2006) generic severity ranges were reclassified to match the three burning efficiency factors (Table 7). Low burning efficiency refers to low consumption of leaves and very low consumption of branches. Moderate burning efficiency denotes intermediate consumption Characterization of Mediterranean Aleppo pine forest using low-density ALS data 40 of leaves and moderate consumption of small branches. High burning efficiency corresponds to a complete consumption of leaves and high loss of small branches and twigs. Table 7. Severity levels, ΔNBR generic ranges defined by Key & Benson (2006) and burning efficiency factors used to estimate biomass losses following De Santis et al. (2010). Severity level ΔNBR range Burning efficiency factors Unburned −100 to +99 0.00 Low severity +100 to +269 0.25 Moderate–low severity +270 to +439 0.42 Moderate–high severity +440 to +659 0.57 High severity +660 to +1300 v. Conversion of biomass losses to carbon content and, subsequently, to CO2 emissions. Biomass can be converted to carbon content by applying conversion factors. The most common factor is 0.5, but, in this PhD Thesis, a value of 0.499 was set for Aleppo pine, in accordance with Montero et al. (2005). These authors determined several conversion factors for Mediterranean species based in field work. Secondly, the conversion from carbon to CO2 was performed according to Trozzi et al. (2002) equation. This equation includes the variables proposed by Seiler & Crutzen (1980), while modifying the specific factors to better represent Mediterranean environments (equation 30) CO 2 = ε*δ* C (30) where ε is the fraction of total carbon emitted as CO 2 (0.888); δ is the factor of conversion from the emissions in ton of carbon to the emissions in ton of CO 2 (44/12); and C is the carbon content. 2.2.8. Modelling: variable selection, regression methods and validation Variable selection Variable selection includes a variety of methods to reduce the dimensionality problem as well as deal with large feature sets, improving model performance, reducing time and costs and generating understandable models (Guyon & Elisseeff, 2003). This PhD Thesis explored the performance of five selection processes to estimate forest stand variables: Spearman’s rank correlation, Stepwise selection, Principal component analysis (PCA) and Varimax rotation, Last absolute shrinkage and selection operator (LASSO) and All subset selection. These methods are described below. Spearman’s rank correlation coefficient is generally denoted by the Greek letter p (rho) (Spearman, 1904). This coefficient determines the strength and direction of the relationship between two Study area, materials and methods 41 variables being carried out using “cor.spearman” function in R environment. Although other correlation coefficients as Pearson or Kendall exist, Spearman’s rank coefficient has been widely used for forestry applications, showing a uniform power for linear and non-linear relationships. The rho coefficient varies between +1 and -1. The value +1 indicates a perfect positive degree of association between two variables while the value -1 denotes a perfect negative degree of association. The weaker the correlation, the closer to 0 is the value. Thus, the correlation value of 0 indicates a null relationship between the analysed variables. In this PhD, the selection of the ALS variables was made considering a minimum positive and negative rho value, ranging from 0.2 up to 0.5. Depending on the higher or lower strength of the relationship between the variables the range was changed to better reduce variable redundancy and multicollinearity. Stepwise selection, or also called stepwise regression, allows selecting predictive variables in an automatic procedure (Efroymson, 1960). The method iteratively adds or drops variables at several steps to determine the best subset of variables, which denotes the best model performance and lower prediction error. The predictor, that is added or dropped in each step, is based in the Akaike information criterion (AIC). There exist three ways of performing stepwise regression: forward selection, backward selection and bidirectional stepwise selection. Forward selection begins with no variables in the model, iteratively selects the variables that contribute the most and stops adding predictors when the improvement is no longer statistically significant. Backward selection begin with all variables in the model, iteratively drops the one that contribute the least and stops dropping predictors when all the selected variables are statistically significant. Bidirectional stepwise selection combines forward and backward selections. Firstly, this method begins with forward selection and then drops those variables no statistically significant using backward selection. The stepwise selection was carried out using “step” function in R environment. PCA is a dimension-reduction method that reduces large sets of predictors to a smaller size, keeping the majority of the information. Although is a classic statistical method that has been broadly used in many research fields, it is not a common approach in ALS metric selection (Silva et al., 2016). PCA was computed using R package “lattice” and specifically with the function “prcomp”, which uses the singular value decomposition to examine the covariance and correlations between individuals. According to the Kaiser Criterion, in his research the generated components with greater values than 0.1 were retained. Then, a Varimax rotation, which maximizes the sum of the variance, was applied to better interpret the PCA results (Darlington & Horst, 1966; Kaiser, 1958). LASSO is a method that performs two processes: regularization and feature selection. The method was popularized and improved by Tibshirani (1996). This technique applies a shrinking or regularization process to penalize some of the coefficients to zero by using a penalty term, which refers to the sum of the absolute coefficients. Those variables that have a non-zero coefficient after regularization will be selected to be part of the model, minimizing prediction errors and generating interpretable models. LASSO was computed in R using the “glmnet” package. All subset selection methods allow a suitable group of metrics being selected, while disregarding Characterization of Mediterranean Aleppo pine forest using low-density ALS data 48 The indirect approach requires higher cost of modelling and data capturing. Thus, we found different results of the evaluation of the two approaches in the literature (Næsset & Gobakken, 2005; Noordermeer et al., 2018; Zhao et al., 2018). Some authors get slightly better performance of the direct approach in the prediction of biomass and carbon fluxes (Bollandsås et al., 2013; Cao et al., 2016; Skowronski et al., 2014) while Meyer et al. (2013) and Zhao et al. (2018) achieved better results with the indirect approach. This PhD Thesis explores the benefits of the two mentioned approaches in the prediction of seven forestry attributes at regional scale: stand density, basal area, squared mean diameter, dominant diameter, tree dominant height, timber volume and total tree biomass. We performed a comparison of the direct and indirect approaches to assess temporal transferability. Firstly, following the indirect approach, two different models were fitted for the available ALS-PNOA data (2011 and 2016), estimating the stand attributes for each point in time, using different ALSmetrics and model parameters. In this approach inventory data was updated using single-treegrowth models to generate concomitant information to the ALS-PNOA years (see Section 2.2.3). Secondly, the direct approach was tested. The models, fitted for one point in time, were extrapolated to the other point in time using the same variables and model parameters. The process was performed two times: the 2011 models were extrapolated to 2016, and inversely the 2016 models were extrapolated to 2011. The modification of the original methodology tested the validity of multi-temporal ALS data to perform future or retrospective analyses. 2.2.10. Variable influence assessment ALS sensors have different configurations that determine the final data characteristics and quality of the point clouds. These configurations are normally different from one flight to another, varying according to the specific sensor capabilities, and might have an effect in model accuracy. In this sense, the effect of three ALS characteristics (point density, scan angle, canopy penetration pulse) and two environmental conditions (slope and shrub presence) on the prediction of forest residual biomass have been addressed. The effect of point density has been explored in the prediction of different forest inventory attributes such as height, basal area or volume (Roussel et al., 2017; Gobakken & Næsset, 2008). However, less studies have considered point density effect on biomass prediction (García et al., 2017; Singh et al., 2015; Ruiz et al., 2014). The effect of scan angle on tree height prediction has been analysed by Disney et al. (2010), Holmgren (2004), Liu et al. (2018) and Montaghi (2013). Disney et al. (2010) proposed to minimize the use of data collected at scan angles greater than ∼15°. According to Holmgren (2004) and Liu et al. (2018), the prediction of canopy closure is affected by the scan angle, being necessary to avoid off-nadir angles from 23 to 38°. Similar results were concluded by Montaghi (2013) in the prediction of forestry metrics using an area based approach, determining that scan angles higher than 20° had a great effect in forest parameters predictions. The structural characteristics and ALS flight settings modify canopy pulse penetration (CPP), decreasing DTM accuracy (Cowen et al., 2000; Hollaus et al., 2006; Hyyppä et al., 2000). The higher Study area, materials and methods 49 the density of the forest, the lower CPP rates are found (Hollaus et al., 2006). Secondary effects of densely covered areas would be a decrease in CPP in lower strata (Chasmer et al., 2006b; Wasser et al., 2013) and the decrease of DTM accuracy (Cowen et al., 2000; Hollaus et al., 2006; Hyyppä et al., 2000). In contrast, leaf off conditions in deciduous forest improves CPP rates (Hill et al., 2009; Wasser et al., 2013). The presence of steep slopes reduces the accuracy in tree height prediction (Breidenbach et al., 2008; Clark et al., 2004; Ørka et al., 2018), as well as in tree diameter, basal area, number of stems and volume (Ørka et al., 2018). Furthermore, decreases the ability to detect tree tops (Khosravipour et al., 2015). Although ALS predictions are affected by the increase of slope, the detected effect was not severe. Slope effect might be partially explained by the lower accuracy of DTMs in these areas, considering that filters have more difficulties on determining ground points on steep slopes (Montealegre et al., 2015a). The effect of shrub presence has not been previously analysed. The hypothesis that the presence of shrub might affect the prediction of forest residual biomass is associated to the lower CPP in those areas and, consequently, a lower number of ground returns, which may affect DTM generation. Two methodological approaches were tested to analyse the effect of the five abovementioned variables in forest residual biomass prediction. Firstly, a graphical assessment using boxplots, including the average mean prediction error per defined class and variable was applied. In addition, several statistical tests, described below, were applied to determine whether the differences between model performances were statistically significant. The defined classes for each analysed variable are the following:  Point density. Plots were classified into two classes: up to 1 point m-2 and higher than 1 point m-2 according to Montealegre et al. (2015b). Jakubowski et al. (2013) and García et al. (2017) also concluded that this threshold allows to have relatively high accuracies in forest parameters modelling.  Scan angle. Three scan angle classes were tested. The first class was set close to nadir with average scan angle of up to 5°. The second class was defined as plots with an average scan angle between 5 and 15°. The third class included those plots with an average scan angle higher than 15°. The breakpoint between the second and third classes was defined by the maximum average scan angle established by the PNOA mission in order to reach the minimum nominal point density specified by the project.  Canopy pulse penetration. The influence of four categories were analysed: 0%-25%, 25%- 50%, 50%-75% and 75%-100% following Montealegre et al. (2015b). The proportion of the pulses that penetrate canopy and reach the ground was calculated using a ground tolerance of 2 m.  Terrain slope. The effect of the terrain slope was assessed using two categories: smooth slopes of up to 15% and steep slopes higher than 15%.  Shrub presence was binary categorized in plots with and without understory. Characterization of Mediterranean Aleppo pine forest using low-density ALS data 50 The mean predicted error (MPE) was selected to analyse the differences between the established classes. MPE was obtained for each field plot and, afterwards, the average mean per class was computed (equation 36). Furthermore the percentage of mean predicted error (%MPE) respect to the observed mean value was computed according to equation 37. 𝑀𝑃𝐸=∑(𝑦𝑖−𝑦𝑖) 𝑛 𝑖=1𝑛 (36) %𝑀𝑃𝐸=𝑀𝑃𝐸 𝑦×100 (37) where 𝑦𝑖 is the observed value for plot i; 𝑦𝑖 is the predicted value for sample plot i; 𝑛 is the number of plots and 𝑦 is mean observed value for all plots. Previous to significance analysis, normality and homogeneity test were computed. After considering logarithmic and square root transformation we concluded that variables were not normally distributed. In this sense, non-parametric Mann-Whitney and median tests were applied for analysing differences between two categories. Kruskal Wallis test was applied for those analyses with more than two classes. The Mann-Whitney test is considered the main non-parametric alternative to the independent sample t-test. This test is designated to compare two populations. The null hypothesis of the test establishes that both samples come from different populations, while the alternative to the null hypothesis indicates that both samples come from the same population. The median test is a non-parametric one used to determine whether two or more samples differ in their central tendency or median value, consequently it can be inferred whether the samples are from the same population or not. The null hypothesis establishes that both samples come from the same population. Kruskal Wallis test is considered the main non-parametric alternative to the One Way ANOVA test. This method is a rank-based test that determines whether the medians of two or more groups are different. The null hypothesis considers that the samples are from the same population, while the alternative hypothesis establishes that at least one of the samples come from a different population. 2.2.11. Mapping of forest variables The mapping of forest stand variables have been traditionally performed by linking tree metrics to the estimated stand variable using allometric equations and, subsequently, extrapolating these estimates at stand-level or regional scale (Boudreau et al., 2008). The use of remote sensing tools provides a greater overview of large areas and with higher temporal resolution (Castro et al., 2003). The use of optical data and, specially, the characterization of disturbance history with temporal series, have been proposed for estimating forestry variables (Cohen et al., 1996; Coops & Waring, 2001). However, the information captured by passive optical sensors tends to saturate under closed canopy conditions and in dense forests (Lu, 2006). Although active SAR remote Study area, materials and methods 51 sensing improved the accuracy on forestry predictions (Tanase et al., 2014), there are still difficulties in heterogeneous and dense forests (Hyde et al., 2006). ALS data have been proven as the most suitable technique for mapping 3D structure (Zhao et al., 2018). However, the use of ALS data at regional scales is still limited by the acquisition costs, the absence of global coverage and the huge data volumes that need to be processed. The use of ALS data for sampling some areas or by creating strips and the subsequent fusion with passive optical data have been explored for estimating forest stand variables at regional scales (e.g.: Matasci et al., 2018; Pflugmacher et al. 2014). Recently, the increase of ALS data at regional scale or country level, as the case of Spain, has opened new opportunities. In this sense, although computational effort is required, the improvement in hardware and software has allowed processing time reduction. In this sense, this PhD Thesis explored the use of ALS data at regional scales, providing cartographic information to forest managers. In this research, the generation of cartography implied the prediction of a dependent variable for the whole study area by applying a regression model previously generated with the above mentioned methodology. Pixel size is one of the most relevant factors to be considered in mapping. The determination of pixel size should consider field plot size and ALS point density, but might be modified to account for specific requirement of forest managers. Commonly, pixel size is similar to field plot size. An increase in pixel size might decrease mapping accuracy when the number of ALS returns per pixel is low. Furthermore, independent variables that are included in regression models should be converted to raster format. In this sense, FUSION software, specifically designated to work with ALS data for forestry purposes, was used to generate the raster files using “Gridmetrics” and, subsequently “CSVGrid” commands. The generation of the raster for the dependent variable was performed in an R environment following three steps: running the best selected model; reading the independent variables in raster format; spatializing the results. 53 3. Research contributions The papers that constitute the PhD Thesis body are entirely included in this chapter as required by compendium PhD Thesis type. The papers provide different forest applications using ALS-PNOA data and field surveys: the estimation of CO2 emissions generated by wildfires, several biomass fractions and inventory variables. Furthermore, the papers compare several selection and regression methods within different conditions, analysing the effect of ALS characteristics and environmental variables in modelling error and assessing temporal model transferability. Research contributions 55 3.1. Comparison of regression models to estimate biomass losses and CO2 emissions using low density airborne laser scanning data in a burnt Aleppo pine forest Comparación de modelos de regresión para estimar las pérdidas de biomasa y emisiones de CO2 utilizando datos de escáner láser aeroportado de baja densidad en bosques de Pino carrasco afectados por el fuego RESUMEN El conocimiento de la pérdida de biomasa forestal producida por un incendio puede ser de utilidad para la estimación de las emisiones de los gases de efecto invernadero a la atmósfera. Este estudio se centra en la estimación de la pérdida de biomasa y las emisiones de CO2 por la combustión de masas forestales de Pino carrasco en un incendio ocurrido en el municipio de Luna (España). La disponibilidad de datos de escáner láser aeroportado (ALS) de baja densidad permitió estimar la biomasa arbórea pre-fuego. Se realizó una comparación de nueve modelos de regresión con objeto de relacionar la biomasa, estimada en 46 parcelas de campo, con distintas métricas extraídas de los datos ALS. El método de regresión linear multivariante seleccionado como óptimo, incluyó entre las variables independientes el porcentaje de primeros retornos sobre 2 m y el percentil 40 de la altura de los retornos. El modelo se validó utilizando una técnica de validación cruzada “leave-one-out-cross-valiation” (RMSE: 6.1 ton ha-1). Las pérdidas de biomasa se estimaron utilizando una aproximación en tres fases: (i) la severidad del incendio fue obtenida utilizando el índice de diferencia normalizado (ΔNBR), (ii) los pinares de Pino carrasco se delimitaron utilizando el Mapa Forestal Nacional y datos ALS y, (iii) tres factores de eficiencia de combustión se aplicaron considerando los niveles de severidad. La biomasa post-fuego se transformó en emisiones de CO2 (426.754,8 ton). Este estudio evidencia la utilidad de los datos ALS de baja densidad para estimar de forma precisa la biomasa pre-fuego y evaluar las emisiones de CO2 en masas forestales mediterráneas de Pino carrasco. Research contributions 57 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 64 Research contributions 65 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 66 Research contributions 67 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 68 Research contributions 69 Research contributions 71 3.2. Estimation of total biomass in Aleppo pine forest stands applying parametric and nonparametric methods to lowdensity airborne laser scanning data Estimación de biomasa total en bosques de pino carrasco aplicando métodos paramétricos y no parámetros mediante datos de escáner láser aeroportado de baja densidad RESUMEN La cuantificación de la biomasa total es de utilidad para la evaluación de las políticas de regulación climáticas desde escalas locales a escalas globales. Esta investigación estima la biomasa total, incluyendo la biomasa arbórea y arbustiva, en bosques de Pino carrasco localizados en la región de Aragón (España), utilizando datos de escáner láser aeroportado (ALS) y trabajo de campo. La comparación de cinco métodos de selección y cinco modelos de regresión se realizó con objeto de relacionar la biomasa total, estimada en 83 parcelas de campo mediante ecuaciones alométricas, con diversas variables extraídas de la nube de puntos ALS. Para el cálculo de las variables ALS se utilizó un umbral de 0.2 m. La muestra se dividió en entrenamiento y validación componiéndose de 62 y 21 parcelas de campo, respectivamente. El modelo con menor error cuadrático medio después de la validación (15,14 tons ha-1) fue el modelo de regresión linear multivariante. Dicho modelo incluyó tres variables ALS: el percentil 25 de la altura de retornos, la varianza y el porcentaje de primeros retornos sobre la media. El estudio confirma la utilidad de los datos ALS de baja densidad para estimar con exactitud la biomasa total, y por consiguiente mejorar la cuantificación de biomasa disponible y contenido de carbono en Pinares mediterráneos de Pino carrasco. Research contributions 73 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 80 Research contributions 81 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 82 Research contributions 83 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 84 Research contributions 85 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 86 Research contributions 87 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 88 Research contributions 89 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 96 1 Research contributions 97 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 98 Research contributions 99 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 100 Research contributions 101 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 102 Research contributions 103 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 104 Research contributions 105 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 112 Research contributions 113 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 114 Research contributions 115 Research contributions 117 3.4. Temporal Transferability of Pine Forest Attributes Modelling Using Low-Density Airborne Laser Scanning Data Transferibilidad temporal de modelos de variables forestales utilizando datos de escáner láser aeroportado de baja densidad RESUMEN Este estudio evalúa la transferibilidad temporal de modelos generados utilizando datos de escáner láser aeroportado (ALS) y adquiridos en dos fechas diferentes. La estimación de siete variables forestales (densidad de pies, área basal, diámetro cuadrático medio, diámetro dominante, altura dominante, volumen maderable y biomasa arbórea) se realizó utilizando un enfoque basado en áreas en masas forestales mediterráneas de Pino carrasco. Los datos ALS de baja densidad se adquirieron en 2011 y en 2016, mientras que las 147 parcelas de campo se muestrearon en 2013, 2014 y 2016. La generación de datos de campo para las fechas de los vuelos ALS se realizó mediante la aplicación de modelos de crecimiento de árbol individual. Cinco métodos de selección y cinco modelos de regresión fueron comparados para relacionar las observaciones tomadas en campo respecto a las métricas ALS. La selección de los mejores modelos de regresión ajustados para cada variable forestal, y separadamente para los años 2011 y 2016, se realizó utilizando un enfoque indirecto. El ajuste de los modelos y la transferibilidad temporal de los mismos se analizó extrapolando los mejores modelos ajustados para 2011 a 2016, e inversamente de 2016 a 2011. El modelo de regresión no paramétrico support vector machine con kernel radial mostró los mejores resultados. La diferencia en el error cuadrático medio en porcentaje de los modelos ajustados y extrapolados fue de 2,13% para los modelos de 2011 y 1,58% para los de 2016. Research contributions 119 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 120 Research contributions 121 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 128 Research contributions 129 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 130 Research contributions 131 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 132 Research contributions 133 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 134 Research contributions 135 Characterization of Mediterranean Aleppo pine forest using low-density ALS data 136 Research contributions 137