Advanced optical technologies for phytoplankton discrimination: Application in adaptive ocean sampling networks
Abstract
Memoria de tesis doctoral presentada por Ismael Fernández Aymerich para obtener el título de Doctor en Ciencias del Mar por la Universitat Politècnica de Catalunya (UPC), realizada bajo la dirección del Dr. Jaume Piera Fernández del Institut de Ciències del Mar (ICM-CSIC) y del Dr. Albert-Miquel Sánchez Delgado.-- 118 pages
Full text
Universitat Polit` ecnica de Catalunya Doctoral Thesis Advanced optical technologies for phytoplankton discrimination: Application in adaptive ocean sampling networks Author: Ismael Fern´andez Aymerich Supervisors: Dr. Jaume Piera Dr. Albert Miquel A thesis submitted in fulfilment of the requirements for the degree of Doctor of Philosophy in Telecommunications Engineering in the Universitat Polit`ecnica de Catalunya (UPC) Department of Signal Theory and Communications Barcelona, November 2015
“Cree a aquellos que buscan la verdad. Duda de los que la encuentran.” ”Croyez ceux qui cherchent la v´erit´e, doutez de ceux qui la trouvent.” Andr´e Gide
Abstract There is a lack on ocean dynamics understanding, and that lead oceanographers to the need of acquiring more reliable data to study ocean characteristics. Oceanographic measurements are difficult and expensive but essential for effective study oceanic and atmospheric systems. Despite rapid advances in ocean sampling capabilities, the number of disciplinary variables that are necessary to solve oceanographic problems is large. In addition, the time scales of important processes span over ten orders of magnitude, and due to technology limitations, there are important spectral gaps in the sampling methods obtained in the last decades. Thus, the main limitation to understand these dynamics is an inaccurate measurement of the process due to undersampling. But fortunately, recent advances in ocean platforms and in situ autonomous sampling systems and satellite sensors are enabling unprecedented rates of data acquisition as well as the expansion of temporal and spatial coverage. Many advances in technologies involving different areas such as computing, nanotechnology, robotics, molecular biology, etc. are being developed. There exist the effort that these advantages could be applied to ocean sciences and will prove to extremely beneficial for oceanographers in the next few decades. Autonomous underwater vehicles, in situ automatic sampling devices, high spectral resolution optical and chemical sensors are some of the new advances that are being utilized by a limited number of oceanographers, and in a few years are expected to be widely used. Thanks to new technologies and, for instance, utilization of data assimilation models coupled with autonomous sampling platforms can increase temporal and spatial sampling capabilities. For instance, studies of phytoplankton dynamics in the water column, or the transportation and aggregation of organisms need a high rate of sampling because of their rapid evolution, that is why new strategies and technologies to increase sampling rate and coverage would be really useful. However, other challenges come up when increasing the variety and quantity of ocean measurements. For instance, number of measurements are limited by costs of instruments and their deployment, as well as data processing and production of useful data products and visualizations. In some studies, there exists the necessity to discriminate and detect different phytoplankton species present in sea water, and even track their evolution. The use of their optical properties is one of the approximations used by some of them. Acquiring optical properties is a non-invasive and nondestructive method to study phytoplankton communities. Phytoplankton species are then organized thanks to presenting similar optical characteristics. Fluorescence spectroscopy has been used and
found as a really potential technique for this goal, although passive optical techniques such as the study of the absorption can be also useful, or even their combination can be studied. Specifically speaking about fluorescence, the majority of the studies have centered their effort in discriminating phytoplankton groups using their excitation spectra because the emission spectra contains less information. The inconvenient of using this kind of information, is that the acquisition is not instantaneous and it is necessary to spend some time (over a second) exciting the sample at different wavelengths sequentially. In contrast, the whole emission spectra can be acquired instantaneously. Therefore, the aim of this thesis is to explore new and powerful signal processing techniques able to discriminate between different phytoplankton groups from their emission fluorescence spectra. This document presents important results that demonstrate the capabilities of these methods.
Resum Existeix una falta de coneixement sobre les din`amiques dels oceans, i aix`o porta als ocean`ografs a la necessitat d’adquirir dades m´es fiables per tal d’estudiar les caracter´ıstiques dels oceans. Les mesures oceanogr`afiques s´on dif´ıcils i costoses d’adquirir, per`o s´on essencials per estudiar de manera efectiva els sistemes oce`anics i atmosf`erics. A causa dels r`apids aven¸cos a l’hora de mostrejar aquest medi tan hostil, ´es necessari que diverses disciplines treballin juntes per tal de solucionar el gran nombre de problem`atiques que es poden trobar. A m´es, els processos que s’han d’estudiar poden perdurar fins i tot deu ordres de magnitud, i per culpa de les limitacions tecnol`ogiques, existeixen importants manques en els m`etodes de mostrejar que es porten utilitzant en les ´ultimes d`ecades. Per tant, la principal limitaci´o per entendre aquestes din`amiques ´es la impossibilitat de mesurar els processos correctament com a conseq¨u`encia d’una baixa freq¨u`encia de mostreig. Per sort, aven¸cos recents en plataformes oce`aniques i sistemes de mostratge aut`onoms, junt amb dades de sat`el·lit estan millorant molt aquestes freq¨u`encies d’adquisici´o, i per tant augmentant la cobertura temporal i espacial d’aquests processos. Actualment hi ha disciplines com computaci´o, nanotecnologia, rob`otica, biologia molecular, etc. que estan protagonitzant uns aven¸cos tecnol`ogics sense precedents. La intenci´o ´es aprofitar tot aquest esfor¸c i aplicar-ho en oceanografia. Vehicles aut`onoms sota l’aigua, sistemes autom`atics de mostreig, sensors `optics o qu´ımics d’alta resoluci´o s´on algunes de les tecnologies que es comencen a utilitzar, per`o que per culpa del seu cost encara no estan esteses i s’espera que ho puguin estar en els pr`oxims anys. Gr`acies a algunes d’aquestes tecnologies, com per exemple la utilitzaci´o de models d’assimilaci´o de dates conjuntament amb plataformes aut`onomes de mostreig, es pot incrementar la capacitat de mostreig, tant temporal com espacial. Un exemple clar d’aplicaci´o ´es l’estudi de les din`amiques del fitopl`ancton, aix´ı com el transport i agregaci´o d’organismes dins la columna d’aigua. No obstant aix`o, no tot s´on aspectes positius, altres reptes sorgeixen en augmentar la varietat i quantitat de mesures oceanogr`afiques. El nombre de mesures queda doncs limitat pels costos dels instruments i les campanyes, i a m´es s’han d’estudiar nous sistemes per processar i extreure informaci´o ´util d’aquestes dades, ja que els m`etodes coneguts fins ara potser no s´on els m´es adients.
La detecci´o i discriminaci´o de diferents esp`ecies de fitopl`ancton al mar ´es molt important en certs estudis cient´ıfics. Alguns d’aquests estudis es basen en extreure informaci´o de les seves propietats `optiques, per qu`e ´es un m`etode no invasiu ni destructiu. Espectroscopia a partir de la resposta de fluoresc`encia del fitopl`ancton s’ha fet servir en molts experiments i s’ha demostrat que ´es una t`ecnica amb gran potencial, tot i que l’estudi dels espectres d’absorci´o o d’altres t`ecniques basades en m`etodes passius tamb´e es poden fer servir, incl´us combinar-les. Centrant-se en la fluoresc`encia, la majoria dels estudis s’han centrat en discriminar grups de fitopl`ancton a partir dels espectres d’excitaci´o per qu`e els espectres d’emissi´o contenen menys informaci´o. El desavantatge ´es que el temps necessari per adquirir una mostra pot estar entorn al segon, per qu`e es necessita estimular la mostra a diferents longituds d’ona seq¨uencialment. En el cas dels espectres d’emissi´o, amb els aven¸cos actuals en sensors `optics, les respostes espectrals poden ser adquirides gaireb´e instant`aniament. Per aquest motiu, l’objectiu principal d’aquesta tesi ´es explorar noves t`ecniques de processat capa¸ces de discriminar diferents grups de fitopl`ancton a partir dels seus espectres d’emissi´o de fluoresc`encia. Aquest document presenta doncs importants resultats que demostren la capacitat de discriminaci´o d’aquest tipus d’informaci´o en combinaci´o amb t`ecniques de processat adients.
Acknowledgements Tots aquells que han arribat a aquest moment de presentar una tesi doctoral saben que ´es un cam´ı llarg i dur, on el treball b`asicament ´es individual. Per`o durant aquest temps hi ha persones que t’ajuden i et donen suport. A totes aquestes persones els hi dono les gr`acies per qu`e sense la seva aportaci´o, per petita que sembli, no hagu´es estat possible terminar aquesta tesi. Sense menysprear el suport de ning´u, voldria destacar en aquestes l´ınies a algunes persones, sabent que aix`o sempre pot comportar que m’oblidi d’alg´u, i que demano perd´o d’avantm`a. Evidentment, el meu primer agra¨ıment ´es pel meu director de tesi, en Jaume Piera, que juntament amb l’Albert Miquel, codirector, m’han ajudat a tirar endavant aquest projecte i poder acabar la tesi. Tanmateix, voldria tamb´e agrair a en Sergi Pons i l’Elena Torrecilla el suport mutu que ens hem donat durant aquesta etapa. Ara queden molt lluny aquelles llargues jornades, compartint dubtes i treballant fins a les tantes de la nit. Encara recordo aquelles converses que ten´ıem, Sergi, quan ens qued`avem sols al despatx i gaireb´e a tot l’edifici, que ens servien per desconnectar encara que fos una estona. A les persones de l’ICM i la UTM amb qui he compartit el dia a dia, en especial a en Rub´en, la N´uria i en Marc, amb qui he compartit dinars, converses, discussions, molts bons moments i d’altres no tan bons, per`o sempre animant i aportant en positiu. About my internships in the TU Berlin and the IST Lisboa, I would also like to thank Prof. Obermayer and Prof. Bioucas for letting me work in their research group. It was an enormous pleasure to work with them and their colleagues. Their knowledge and expertise were really valuable and priceless for this thesis. Tamb´e mereixen un agra¨ıment aquelles persones que m’ha donat suport fora de l’`ambit acad`emic. Gr`acies a tots els meus amics, especialment als meus amics Ra´ul, David, Javi, Oscar, Lu´ıs, Jaume, Carlos, que sempre m’han ajudat a tenir moments de desconnexi´o de la tesi, totalment necessaris per afrontar la feina amb nous `anims i forces per seguir treballant. Ara, per fi, tenen un document que poden llegir per tal de saber a qu`e li dedicava tant de temps, i entendre en qu`e consistia la meva feina. Tamb´e hi han hagut persones i fets que han contribu¨ıt a qu`e aquest cam´ı hagi estat m´es dur i llarg, per`o que incl´us tamb´e mereixen aquest reconeixement, per qu`e sense elles no estaria tampoc ara aqu´ı, ni seria la persona que s´oc ara. viii
List of Figures xv 4.5 Hierarchical Clustering of phytoplankton culture spectra. ............... 84 4.6 Hydrolight-Ecolight Radiative Transfer schema ..................... 87 4.7 Triangular concentration profile used for simulacions. .................. 89 4.8 Example of variations of Rand Kdwith depth. ..................... 91 4.9 Example of variations of Rand Kdwith concentration and depth ........... 91 4.10 Kdphytoplankton endmembers at different depths ................... 94 4.11 Resulted abundances using LSMM of mixtures 1, 2 and 3 ............... 95 4.12 Resulted abundances using LSMM of mixtures 4, 5 and 6. ............... 96 4.13 Resulted abundances using LSMM of mixtures 7, 8 and 9. ............... 97 4.14 Performance error retrieving abundances using Kd................... 98
List of Tables 2.1 Simple example of a binary confusion matrix. ...................... 22 2.2 Phytoplankton species used in this chapter and number of samples acquired in each experiment. ......................................... 26 2.3 Example of a confusion matrix. Classification of excitation spectra from a random selection of training and test samples. .......................... 29 2.4 Confusion matrix. Classification behavior using excitation spectra. Samples from the stable growth stage used for training, and samples from the exponential growth stage used for testing. ................................... 30 2.5 Example of a confusion matrix. Classification of emission spectra from a random selection of training and test samples. .......................... 32 2.6 Confusion matrix. Classification behavior using emission spectra. Samples from the stable growth stage used for training, and samples from the exponential growth stage used for testing. ....................................... 33 2.7 Example of a confusion matrix. Classification of first-derivative emission spectra from a random selection of training and test samples. .................. 34 2.8 Confusion matrix. Classification behavior using first-derivative emission spectra. Samples from the stable growth stage used for training, and samples from the exponential growth stage used for testing. ........................... 35 2.9 Example of confusion matrix obtained from excitation spectra: (a) P-SVM, (b) SOM 37 2.10 Example of confusion matrix obtained from emission spectra: (a) P-SVM, (b) SOM 38 2.11 Example of confusion matrix obtained from derivative excitation spectra: (a) PSVM, (b) SOM ....................................... 39 2.12 Example of confusion matrix obtained from derivative excitation spectra: (a) PSVM, (b) SOM ....................................... 40 3.1 Algorithms of the four-step signal-processing chain. ................... 47 3.2 Taxonomic groups of the five cultures. .......................... 54 3.3 RMSE between the covariance of the original data and the covariance of the smoothed data at 684 nm. ....................................... 58 3.4 Kappa indices obtained with the three classification techniques and the four normalization methods when denoising with the WMA method (Gaussian window with α2). ............................................. 62 3.5 Kappa indices obtained with the three classification techniques and the four normalization methods when denoising with the Savitzky-Golay method and n=13. ..... 62 3.6 Kappa indices obtained with the three classification techniques and the four normalization methods when denoising with the wavelet method (the letter denotes the use either soft or hard threshold). ............................... 62 xvi
List of Tables xvii 3.7 Kappa indices obtained with the three classification techniques and the three transformation methods, considering the WMA (Gaussian window with α2) as a denoising method and the modified SBN as a normalization method. ............... 63 3.8 Computational cost expressed in terms of execution time (in seconds), considering the WMA (Gaussian window with α2) as a denoising method and the modified SBN as a normalization method. ................................ 63 3.9 Averaged confusion matrix obtained with the SOM classification method when using the Savitzky-Golay and the Min-Max methods to denoise and normalize. Results of the confusion matrix have been averaged due to the five-fold cross-validation. . . . . 64 3.10 Averaged confusion matrix obtained with the SOM algorithm when using the wavelet (soft threshold) and the Min-Max methods to denoise and normalize. Results of the confusion matrix have been averaged due to the five-fold cross-validation. ...... 64 3.11 Kappa indices obtained with the three classification techniques and the four normalization methods when denoising with the WMA method (Gaussian window with α2). ............................................. 65 3.12 Kappa indices obtained with the WMA method to denoise, the SNV method to normalize, the PCA method for a dimensional reduction and the k-neighbors to classify, by using the degraded samples with four different variances (σ) of the noise Gaussian distribution. ................................... 66 4.1 Phytoplankton species merged for mixtures. The abbreviation is neccesary to follow the discussion and to understand table 4.2. ....................... 81 4.2 Summary of mixtures acquired in the laboratory. It is indicated for each mixture the phytoplankton species mixed and the abundance in volume percentage of each one. ............................................. 82 4.3 List of species used in this chapter experiment grouped by Class. .......... 83 4.4 List of species grouped by major algal functional groups following the proposal stated by Beutler in [1]. ...................................... 83 4.5 List of species grouped by spectral similarity using Hierarchical Clustering Analysis (HCA). ........................................... 84 4.6 Phytoplankton groups under study working with hyperspectral Irradiance Reflectance (R) and Diffuse Attenuation Coefficient (Kd) simulations. ............... 88 4.7 Performance results of similarity for different mixed profiles. Combinations of concentrations between the first and second peak (bold for Kd). .............. 92 4.8 Performance results of similarity for different mixed profiles after derivative analysis. Combinations of concentrations between the first and second peak (bold for Kd). . . 92 4.9 Phytoplankton groups using Diffuse Attenuation Coefficient (Kd) for unmixing using Linear Spectral Mixing Model (LSMM) ......................... 93 4.10 Summary of mixtures simulated in Hydrolight-Ecolight to be unmixed. It is listed the abundance of each constituent. ............................ 93
Abbreviations ANN Artificial Neural Networks AOP Apparent Optical Property AUV Autonomous Underwater Vehicle BMU Best Matching Unit CCA Convex Cone Analysis DWT Discrete Wavelet Transform EMR ElectroMagnetic Radiation FCLS Full Constrained Least Squares FN False Negative FP False Positive FPR False Positive Rate FWT Fast Wavelet Transform GCS Growing Cell Structures GSM Growing Spectra Modeling HCA Hierarchical Clustering Analysis HE HydroLight-Eco-Light ICA Independent Component Analysis IDWT Inverse Discrete Wavelet Transform IFA Independent Factor Analysis IOP Inherent Optical Property LMM Linear Mixing Model LSE Least Squares Error LSMM Linear Spectral Mixing Model MLE Maximum Likelihood Estimator xviii
Abbreviations xix MWV Max-Wins Voting NCLS Nonegatively Constrained Least Squares NLMM Non-Linear Mixing Model PCA Principal Component Analysis P-SVM Potential-Support Vector Machines RMSE Root Mean Square Error SBN Scale-Based Normalization SMA Spectral Mixture Analysis SNR Signal to Noise Ratio SNV Standard Normal Variate SOM Self-Organizing Maps SVD Singular Value Decomposition SVM Support Vector Machines TN True Negative TP True Positive TPR True Positive Rate VCA Vertex Component Analysis WMA Weighted Moving Average
Dedicada a la meva familia per estar sempre al meu costat. xx
Chapter 1 Introduction Oceans flow over nearly three quarters of the Earth, and holds almost the 97% of the planet’s water. Oceans are also responsible of more than a half of the oxygen produced and released to the atmosphere, as well as the absorption of the most carbon dioxide from it. Oceanic processes like El Ni˜no change weather patterns. About half of the world’s population lives within the coastal zones, and their economy are direct or indirectly related with ocean-based businesses, contributing to the world’s economy. Oceans are in so many ways really crucials for us, and it is not surprising the importance of studying and understanding them [2]. Oceanography is an example of interdisciplinary science. Diverse scientific disciplines are required to successfully solve the wide variety of oceanographic problems. There is a lack on ocean dynamics understanding, and that lead oceanographers to the need of acquiring more reliable data to study ocean characteristics. Oceanographic measurements are difficult and expensive but essential for effective study of the oceanic and atmospheric systems. Ocean’s complexity and variability at spatial and temporal scales are spanning over ten orders of magnitude (Figure 1.1) and environmental adversity makes it one of the most challenging environments of science. Further, episodic events such as tsunamis, hurricanes, typhoons, submarine volcanic eruptions, earthquakes, harmful algal blooms (HABs), oil leakages, etc., which are difficult to include in time-space diagrams such as Figure 1.1, present especially great sampling challenges. Oceanographers study such a richly diverse spectrum of interesting problems that it is difficult to focus on a single example. However, from several relevant studies, as Dickey and Bidigare pointed 1
Chapter 1. Introduction 2 Figure 1.1: Example of different ocean processes and how they span in time and horizontal space (Figure adapted from [3]). out in [3], four interdisciplinary problems can be extracted as the most important and challenging, and this thesis is focused in one of them. •Biogeochemical variability and global climate change: Understanding of biogeochemical variability, ocean circulation and processes, and global change will require an effort to obtain more accurate measurements using a higher temporal and spatial resolution. The influence of changing oceanic conditions on climatic time scale phenomena and vice versa should be studied, and it is especially challenging. •Impacts of hurricanes and typhoons on the ocean: Their impacts on the open and coastal ocean have remained largely documented, but there is not much knowledge about the ocean key processes. Especially regarding how they affect the distributions of physical, biological, chemical, and geological variables. Recent studies have provided new insights into oceanic effects on currents, mixing, biological productivity, gas exchange, and sediment resuspension
Chapter 1. Introduction 3 [4,5]. However, hurricane or typhoon responses over broad regions have not been observed because satellites can just infer surface processes. •Oceans and human health: It is also important to mention that the application of biotechnology to the marine biosphere can provide for example new drugs and processes that serve a broad array of sectors, including human health or environmental advances. The biodiversity of the subtropical ocean, for instance, is really high, and mostly unknown. Recent studies have shown that certain microorganisms possess novel compounds with anti-bacterial and anti-fungal activities as well as anti-tumor activity [3]. Tropical coastal waters remain also uncharacterized with regard to conditions, which are responsible for higher transmission of diseases in the use of these waters. It is then necessary to continue exploring this field in order to make great advances in human health. •Harmful algal blooms (HABs): The fourth topic highlighted by Dickey in his review is the impact of harmful algal blooms, or commonly called red tides [6], on human health, ocean ecosystems and economic consequences for coastal communities. The technologies required to understand and model HABs and some other processes are quite diverse. The identification of HABs and other species is necessary in many studies. Since HABs generally have their greatest impacts in coastal environments, the complexity of the coastal ocean comes into play. In particular, HAB processes can span from hours to decades, for this reason it is important to adequately set the sampling rate. Spatial coverage is also another important parameter since these processes can also cover a wide distance range. Although there are still lots of unsolved oceanographic problems limited by technological barriers (as summarized in Table 2 of [3]) and their complex and unfavorable environment, powerful techniques applied to other sectors of science and engineering can be applied. Interdisciplinary work is really profitable, and in the case of oceanography needed. It is important to take advantage of this collaboration that is accelerating the understanding of the ocean environment. One of the main concerns due to the temporal and spatial variability of the ocean processes is the adequate sampling strategy. Sampling the ocean is complicated and expensive, for this reason it continues undersampled. Over the last years, technological advances have accelerated this progress. Advances in sensors and systems for measuring chemical, bio-optical or bio-acoustical variables, as well as technological achievements have led oceanographers to be able to improve their studies,
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 10 2.1 Introduction High dimensional data is being widely used for many studies since computation machines became more powerful and also due to the arrival of new sensors providing a huge amount of information. In the case tackled in this thesis, hyperspectral optical sensors provide hundreds of characteristics for each measurement, increasing the dimensionality of the data to be analyzed. Specific processing techniques are needed to deal with such quantity of information. Traditional processing methods were designed to be used with data that usually lays in a low-dimensional space. However, there are many applications where data is of considerably higher dimensionality, and these methods have important limitations when dealing with this data and need to be modified or simply replaced by other methods specifically developed for this goal. That is the case, for instance, of the two main classifications methods used in this chapter, neural networks or machine learning, techniques that have their potential application dealing with high dimensional data. Hyperspectral optical sensors provide hundreds of bands or characteristics of the measured signal. Next section will introduce the type of data used to discriminate phytoplankton, its characteristics, and how it is acquired. As mentioned, data acquired with a hyperspectral optical sensor lays in a high dimensional space, which makes it very difficult to analyze and visualize, except for some techniques like Self-Organizing Maps (SOM). Self-Organizing Maps, which is a method helpful to discover low-dimensional manifolds in such dimensional data, is a type of artificial neural network (ANN) that has been successfully applied for extracting interpretable patterns from large and complex data sets, for example, in satellite remote sensing [29]. It has not been widely applied to oceanographic data, but in recent years several studies have shown its good performance in pattern recognition and classification [30–32]. In our case, SOM is part of a solution to the problem that involve other steps and techniques. As it is shown later, the first approximation uses this cluster analysis method to determine possible classes and it is a preliminary step to achieve discrimination results. It is also shown, how with an appropriate treatment of the data, using some pre-processing techniques, it is possible to improve the final results. However, SOM is not always the best technique. In some cases it cannot find a low-dimensional manifold (probably because it does not exist) and cannot converge. In those cases there is still the possibility of using other techniques, such as Support Vector Machines (SVM), which searches for a separation plane, not in a low-dimensional space, but in a much higher space than the analyzed
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 11 data dimension. Although it can seem more computationally expensive, it is not. The properties of SVMs make them well-suited to tackle the problem of hyperspectral classification since it can: a) handle large input spaces efficiently; b) deal with noisy samples in a robust way; and c) produce sparse solutions. Thus, this chapter shows the results with SOM and a comparison between SOM and a variant of SVM called Potential-Support Vector Machines (P-SVM). About the chapter organization, next section will briefly introduce several processing techniques used in the data analysis chain followed for this study, focusing the attention on those more important such as Self-Organizing Maps or Potential-Support Vector Machines. First of all, a couple of transformations that have been proven really powerful in several studies will be introduced: derivative analysis and wavelet denoising. Next, the clustering and classification techniques explored in this chapter are explained, as well as a performance an evaluation index called kappa. This index helps to evaluate the performance of the classification analyzing the confusion matrices generated. And once the theoretical background has been set, the results of this study are presented, together with some conclusions. 2.2 Data Analysis Figure 2.1 shows a block diagram of the processing steps followed in this chapter. Once data have been obtained, two actuations can be done: either attempt directly the classification or adapt the data to extract as much information as possible in order to increase the classification performance. For this reason, in this work the results with and without this pre-processing step have been compared, which in the block diagram appears in two separated blocks: denoising and derivative analysis. Figure 2.1: Schematic of the proposed processing chain. On the other hand, the classification step includes the study of phytoplankton discrimination using the two techniques mentioned above: SOM and P-SVM, and finally the performance evaluation.
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 12 2.2.1 Transformations Several transformations useful to adapt the measured data are described in this chapter. Transformations are used to better extract as much information as possible from the available data, and make its classification easier. 2.2.1.1 Derivative Analysis Until relatively recently, studies using optical information have centered their attention in multispectral data. However, thanks to new technology such as the above mentioned hyperspectral sensors, with higher resolution, signals can be thought as spectrally continuous. Typical multispectral analysis methods treat each spectral band as an independent variable, a reasonable assumption for multispectral data but it might not be really appropriate for hyperspectral data. Till nowadays, few researchers have tried to manipulate data as truly spectrally continuous data [33–35]. Derivative spectroscopy, for instance, is a promising technique to be used with hyperspectral data, including fluorescence. In remote sensing, some researchers have already addressed applications using spectral derivatives [35–37]. One concern about the derivative analysis is the derivative order. While some of these studies have used high order derivatives, others such as [38,39] use first and second order derivatives. For example, [40] describes the application and utility of derivative analysis of absorption spectra in conjunction with spectral similarity analysis to discriminate the presence and dominance of G. breve in natural phytoplankton assemblages. In this experiment, derivatives are estimated using a finite divided difference scheme. The advantage is that the derivatives can be computed according to different finite band resolutions (band separations) to extract special spectral features of interest at different spectral scales. The first derivative can be estimated as follows 2.1: ds dλ|i≃s(λi)−s(λj) ∆λ(2.1) where s is the original spectrum and ∆λis the band separation.
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 13 2.2.1.2 Denoising Spectra are increasingly analyzed using methods such as derivative analysis. These techniques require smooth spectra because they are extremely sensitive to noise. There is then the need of smoothing algorithms that fulfill the requirement of preserving local spectral features while simultaneously removing noise. Minimizing random noise is often the first pre-processing step after acquisition. Although the signal can be affected by different noise sources, in this specific case, data are basically affected by instrumental noise. In this chapter, wavelet denoising has been applied and it is briefly introduced in the next section. However, there exist several denoising techniques that can be used, and actually, in the next chapter some other techniques are explained and tested. Wavelet Denoising The wavelet denoising [41] is a more refined method that separates the frequency content of the original signals into different data structures. The low-frequency components (approximation coefficients) keep the global features of the signal, while the high-frequency components (detail coefficients) retain the local features. For discrete data, it can be computed as: ˜x(λ, k) = ∞ ρ=−∞ x(ρ)1 √2λΨρ−k2λ 2λ(2.2) being Ψ the mother wavelet function, ˜xthe discrete wavelet transform (DWT), and ka location parameter. A fast algorithm to compute the discrete wavelet transform is presented in [41]. Soft and hard threshold techniques [42,43] can be used to reduce the noise, and the threshold level is selected as described in [42], following equation 2.3: thr =ξ2log(n) (2.3) where nis the number of samples and ξis a rescaling factor estimated from the noise level present in the signal. The estimation of the noise level can be based on the first level of the detail coefficients (D1) as [44]:
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 14 ξ=median(|D1|) 0.6745 (2.4) Finally, by applying the inverse wavelet transform, a smoothed version of the original signal is recovered. The advantage of this method relies on a denoising procedure that does not affect the sharp structures of the original data, which can contain important information. Figure 2.2 shows an example of a three-level wavelet decomposition. First, the original signal x yields one series of approximation coefficients A3and a set of three distinct detail coefficient signals D1,2,3. Then, either a soft or a hard threshold methodology is applied on the detail coefficients. In a soft threshold (Figure 2.2a), coefficients smaller than the threshold thr are suppressed while the rest of the coefficients are shrunk an equivalent of the threshold value [41]. In a hard threshold (Figure 2.2b), coefficients smaller than the threshold thr are set to 0 while the rest of the coefficients remain intact. The denoised profile xis finally recovered from the transformed coefficients by applying the inverse discrete wavelet transform (IDWT). Figure 2.2: Schematic diagram of the three steps of the wavelet method: multilevel decomposition, thresholding, and multilevel reconstruction. Thresholding is obtained via (a) soft threshold techniques or (b) hard threshold techniques.
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 15 2.2.2 Clustering and classification techniques Nowadays there exist a wide range of classifiers that are employed in numerous applications, from face recognition to speech processing. The results of these studies are really successful, however, there does not exist a classifier that can reliably outperform all the others on a given data set [45]. Thus, choosing a classifier is still a process of trial and error. The accuracy of a particular classifier on a given data set will clearly depend on the relationship between the classifier and the data. Classification algorithms can be grouped into parametric and non-parametric techniques. For parametric classifiers, such as the Maximum Likelihood Classifier, the data is assumed to follow a statistical distribution. For this reason, the major drawback of parametric classifiers is their high dependence on assumptions related to statistical distribution of the data. Furthermore, these algorithms are more likely to suffer from the problem of the curse of dimensionality or Hughes phenomenon [46] in hyperspectral classification. Non-parametric classifiers, such as those based on neural networks and decision trees, are then often used to classify hyperspectral data. Although neural networks have the advantage of a good performance with complex data sets, they are slow in the training phase. On the other hand, Support Vector Machines, based on machine learning algorithms, have been proposed that can overcome some of the limitations of other non-parametric classifiers, such as neural networks [47]. Therefore, two non-parametric approaches have been tested: an artificial neural network (SelfOrganizing Maps) and a machine learning method (Potential-Support Vector Machines). The main characteristics of these two techniques are briefly presented below. 2.2.2.1 Neural Networks. Self-Organizing Maps Artificial neural networks are a processing technology that has been widely studied over the last few decades. Inspired by neuroscience, they are trained to behave like biological neural networks, emulating how they process the data. One of their advantages is that they are more robust in handling noisy and missing data than traditional methods. How neural networks work depends on the interconnectivity between the neurons. As it is mentioned in [48], there are three categories of neural networks, each one based on a different philosophy. In feedforward networks sets of input signals are transformed into sets of output signals, and this transformation is determined by externally adjusting several parameters. In feedback networks, the parameters are changed iteratively
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 16 from an initial state until the desired outcome is obtained. Finally, in competitive, unsupervised, or self-organizing networks, neighboring cells in a neural network compete and interact to correctly match (represent) the input space. SOMs are commonly used for clustering high dimensional data, but also for pattern recognition and visualization of complex data sets in a variety of environmental science applications. It is a useful tool for multivariate data analysis because it is both a projection method, mapping high-dimensional data to a low-dimensional space, and a clustering method, mapping similar data patterns onto neighbouring SOM output nodes [49]. SOMs have been used in a wide range of studies, from meteoroloical and climatology applications [50–54] and to oceanographic studies. In this sense, it has been increasingly used since Richardson showed in [29] the use of the SOMs to the wider oceanographic community [55,56]. SOMs have also been used for sea surface temperature studies [29,57] and ocean colour and chlorophyll studies [58,59], among others [60,61]. Self-Organizing Maps (SOM) The Kohonen self-organizing maps [48] are a type of artificial neural network based on unsupervised learning, which means that the network learns only based on the input training data. In contrast, supervised learning needs the pairs of input/output training patterns in order to approximate the input data. The SOM projects high-dimensional input data, usually onto a two-dimensional map, a feature that is useful for the visualization and classification of high-dimensional data. Also, the algorithm is topology-preserving, which means that similar input data will be mapped to spatially close areas on the map, and elements which are spatially close on the map should have similar input data. A SOM output map consists of neurons organized on a regular low-dimensional grid, each one holding a weight vector (Wij) (figure 2.3). This weight vector has exactly the same length as the dimension of the input data, and the lattice of the nodes can be either hexagonal or rectangular. Once the weight vectors are initialized, the SOM training algorithm adapts them so that the neurons span across the data cloud. At the end of the training phase the map is organized such that neighboring neurons in the grid have similar weight vectors. Two auxiliary matrices are generated to help in the visualization of the resulting clusters, the U-matrix and the hit-matrix. The training algorithm can be summarized in two steps:
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 17 Figure 2.3: (a) Example of weight vectors, codebook, once the neural network has been trained. Neighbor neurons have a similar weight vector. and schematic neural adaptation. (b) Visual representation of the adaptive step, in which the BMU and its neighborhood learn and change their weight vectors. The data point of the training set driving the adaptation of neurons is represented as an X. The BMU moves into this position in the feature space. Due to the neighboring function definition, neighboring neurons are moved in the same direction. •Finding the Best Matching Unit: During each training step, one input sample xis randomly chosen from the training set. The distances between this input sample and the weight vectors of all neurons are then computed (typically Euclidean distance). The neuron that has the minimum Euclidean distance between the input vector and its weight vector is the winning neuron and is called the best-matching unit (BMU). •Adapting the Weight Vector: Once the BMU has been chosen and the input vector has been assigned to the winning neuron, it is time to learn. The BMU and its neighboring neurons update their weight vectors to make them similar to the input vector as follows (Eq. 2.5). Wij(t+ 1) = Wij(t) + α(t)×hc(t)×[x(t)−Wij(t)] (2.5) where x(t) is the input data vector, hc(t) is the learning neighborhood function (typically a Gaussian bell-shaped one), and α(t) is the learning rate. The neighboring function (hc(t))
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 18 defines the region of influence that the input sample has on the SOM, and both αand hc decrease with time, performing a fine tuning at the end of the training. At each learning step, all the neurons within the neighborhood (Nc) are updated, whereas cells outside (Nc) are left intact. The neighborhood function is often taken to be Gaussian (Eq. 2.6): h(t) = exp −ρ2(t) 2σ2(2.6) where σ2is the variance parameter specifying the spread of the Gaussian function, and ρ(t) is the radius of the neighboring function centered at the BMU. The learning rate denotes the regularization parameter of the adapting procedure (Fig. 2). Once the SOM training has finished, the U-matrix is constructed, representing the distances between the neurons of the output map, for example, as gray values. For a network of P×Qneurons, the Umatrix has (2P−1)×(2Q−1) distances between neurons or values [62]. It is used in order to obtain an initial idea of the cluster distribution [63]. Clusters are characterized in this representation as a homogeneous area of dark gray values separated by edge-wise elongated areas of light gray values (as an example, figure 2.4 shows the resulted U-matrix clustering 5 different species of phytoplankton). Once the output map has been trained, the data set is applied once again in order to obtain the winning neuron for each sample. This information is accumulated, and the most-frequent winning values can be considered as the most representative ones. The result, presented in a two-dimensional histogram, is the so called hit-matrix, used in the classification step (figure 2.5). In this study, the somtoolbox [64] for Matlab was used for the presented result computation. 2.2.2.2 Kernel Methods applied on Potential-Support Vector Machines The basic idea of kernel methods is that they apply two consecutive mappings to data. The first one maps the points of the input space (the input data) into an intermediate space called feature space. This step transforms the original nonlinear problem into a linear one. If a problem in the original representation can only be solved by nonlinear approaches, its transformed version in the feature space can be solved using linear methods. The theoretical background of applying nonlinear mapping to transform a nonlinear problem into a linear one comes from the Cover’s theorem on the separability of patterns [65]. Cover’s theorem states that a multidimensional space may be transformed into a new feature space where the patterns are linearly separable with high probability
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 19 Figure 2.4: Complete classification schema using Self-Organizing Maps. Diagram of the steps followed by the SOM classification method used in this thesis. Excitation spectra are used in this example, acquiring the different emission fluorescence spectra at 680nm. Figure 2.5: Example of two fuzzy hit-matrices from different cultures; (a) Thalassiosira weissflogii (Thwi) and (b) Alexandrium minutum (Amin), extracted from excitation spectra analysis. This information is used in the classification step.
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 26 peak were evaluated. However,in this thesis an interpolation of the two neighboring samples to avoid this effect is used because it obtains better results in combination with SOM. Table 2.2: Phytoplankton species used in this chapter and number of samples acquired in each experiment. Species Class Abbreviation Number samples Experiment n. 1 Number samples Experiment n. 2 Alexandrium minutum Dinophyceae Am 22 41 Thalassiosira weissflogii Bacillariophyceae Thwi 18 38 Dunaliella Chlorophyceae Duna 20 40 Isochrysis galbana Prymnesiophyceae Iso 10 30 Pleurocrysis elongata Prymnesiophyceae Pl 21 42 Synechococcus sp. Cyanophyceae Syn 15 - Ostreococcus sp. Prasinophyceae Ost 15 - Figure 2.8: Classification performance curve at different excitation wavelengths to select the best excitation wavelength to work with. 2.3.2 Discrimination using Self-Organizing Maps In this section, the first results obtained using SOM and published in [24] are shown. The classification using SOM can be explained in three steps. First, the network is trained using the training data set. The network adapts its properties to the input data and then the distances between neurons are calculated. The result of this process is the U-matrix (figure 2.4), which shows the distances between neighboring neurons. The light gray is the edge of the clusters, where these
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 27 distances are higher. Afterwards, as explained above, a variant of the hit-matrix called the fuzzy hit-matrix [62] is computed. This matrix is obtained for each culture. Figure 2.5 shows an example of these matrices. The gray value of the matrices represents a particular neuron’s membership of the culture class. The matrices are then labeled. Lastly, we wish to assign a label to each sample to classify, so the best-matching unit of each sample is found. Then, the different membership values, which are extracted from the fuzzy hit-matrices, are compared. The winning class, whose label is then taken to classify the sample into one culture or another, is that with the largest BMU membership value. As explained above, two different approaches have been followed. First of all, the results using excitation spectra are shown, while those obtained with emission fluorescence spectra will be presented later. Finally, derivative analysis has been applied to emissions spectra in order to increase classification performance. The data used in this section is the one obtained in the experiment n. 1 (table 2.2). 2.3.2.1 Classification using Excitation Spectra Discrimination results among seven different phytoplankton cultures using SOM are now presented. In this step only the responses at 680 nm emission wavelength with a range of 200–600 nm excitations were used. In this case, randomly selected training and test data sets were constructed to evaluate the method. An example of a training data set is shown in figure 2.9. Once the neural network has been defined with 8x15 neurons, it is trained, and a label can be assigned to each neuron in the SOM by first computing the hit-matrix over the whole training set. The label of each neuron corresponds to the label of the class with a maximal number of hits in each neuron. In this label representation, the discrimination of the method can be appreciated (figure 2.10). The neurons have changed their properties to better characterize the input data. The different cultures should appear classified, grouping all the samples of the same culture, but there are some mixed samples. The samples that have a similar spectrum appear closer. For instance, Alexandrium minutum and Synechococcus sp. have similar excitation spectra because they appear close one from each other, whereas Thalassiosira weissflogii differs because it is in a separated region of the network. Once the neural network has been trained, the labeled output map can be used to classify the test data set, using the membership matrices. Based on this classification, the confusion matrix is
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 28 Figure 2.9: An example of a training data set working with excitation spectra. Fluorescence excitation spectra acquired at 680 nm. 60 samples randomly selected representing the seven cultures. The excitation range is 200–600 nm, every 10 nm. Figure 2.10: An example of the results obtained from randomly chosen training and test samples using excitation spectra.
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 29 computed in order to evaluate the performance of the classification methodology. The confusion matrix summarizes in a table how well the classifier is working. It shows the number of samples correctly classified and how many have been mistakenly classified. The Kappa, the TPR, and the FPR indexes are calculated. Ten different validation runs, in which the training and test data sets had been randomly constructed iteratively, were undertaken. Thence the results were averaged in order to make the performance evaluation as independent as possible from the selected training set. An example of a confusion matrix is presented in table 2.3. Table 2.3: Example of a confusion matrix. Classification of excitation spectra from a random selection of training and test samples. Predicted Class Ost Syn Thwii Duna Pl Am Iso Sum TPR FPR True Class Ost 40 0 2 0 0 2 8 0.5 0 Syn 0 40 0 2 2 0 8 0.5 0 Thwii 0 0 81 0 1 10 0.8 0.88 0.018 Duna 0 0 0 81 0 1 10 0.8 0.057 Pl 0 0 0 0 56 0 11 0.45 0.058 Am 0 0 0 0 0 11 0 11 1 0.157 Iso 0 0 1 0 0 0 45 0.8 0.052 Even though the discrimination does not seem to be good enough, the average TPR and the average FPR over cultures and validation runs were 0.7344 and 0.0508, respectively. Also, the Kappa index value was quite good, 0.6839, with 1 denoting a perfect classification without any mistake, and 0 denoting the result of a random classification. The greatest problems arose with Pleurochrysis elongata, which is not properly classified. This means that the spectra from this culture are very similar, in this case, to those of Alexandrium minutum. Some studies carried out by Margalef [22] pointed out that pigment composition changes during the life of phytoplankton. This statement introduces another variable that makes the classification even more challenging. For this reason, another configuration of training and test data has been proven. In this case, the system has been trained with the samples from the stable growth stage of the culture, while the test samples have been taken from the exponential growth stage (figure 2.11), the performance of the method using the stable samples as training data is then evaluated. The resulting U-matrix is shown in Figure 2.4. Pigment composition variations play here an important role in the classification. Testing samples from the exponential growth stage were classified to form the confusion matrix (table 2.4). The Kappa index obtained was 0.4974 (TPR = 0.5787, FPR =
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 30 0.0981). Again, the greatest classification problems arose with Pleurochrysis. Although initially the U-matrix seems to be better, the Kappa value was below 0.5. This result could be a consequence of the pigment changes that phytoplankton cultures suffer during their growth stage. Figure 2.11: As an example, Alexandrium minutum’s growth curve is shown. It has been made computing the fluorescence emission at 680 nm every day with 490 nm excitation wavelength. Table 2.4: Confusion matrix. Classification behavior using excitation spectra. Samples from the stable growth stage used for training, and samples from the exponential growth stage used for testing. Predicted Class Ost Syn Thwii Duna Pl Am Iso Sum TPR FPR True Class Ost 40 0 2 0 0 4 10 0.4 0.013 Syn 0 30 0 0 7 0 10 0.2 0 Thwii 1 0 92 1 0 0 13 0.692 0 Duna 0 0 0 11 1 3 0 15 0.666 0.028 Pl 0 0 0 1 213 0 16 0.187 0.086 Am 0 0 0 0 0 17 0 17 1 0.348 Iso 0 0 0 1 0 0 45 0.8 0.049 Using this table, the Kappa index is computed as follows: nr k=1 Xkk = 86 ×(4 + 3 + 9 + 11 + 2 + 17 + 4) = 86 ×50 = 4300 r k=1 Xk+X+k= (10 ×5) + (10 ×3) + (13 ×9) + (15 ×17) + (16 ×4) + (17 ×40) + (5 ×8) = 1236 K= (nr k=1 Xkk −r k=1 Xk+X+k)/(n2−r k=1 Xk+X+k) = (4300 −1236)/(862−1236) = 0.4974 2.3.2.2 Classification using Emission Spectra The idea of working with emission spectra comes up from the need of a rapid acquisition method. There are some physiological and trophic processes that may be constrained by physical processes operating over spatial scales of a few centimeters and temporal scales of seconds to minutes [15].
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 31 The importance of detecting these processes emphasizes the need to develop new techniques aimed at increasing the number of samples and improving then horizontal and vertical resolution. Acquiring excitation spectra takes almost a second, and the speed of a free falling vertical profiler is near 10cm/s. In contrast, acquiring emission fluorescence spectra is almost instantaneous and the number of samples could be increased. Using the emission spectra, in situ acquisition would be faster and a higher vertical and horizontal resolution could be achieved. The main problem is that fluorescence information is not as high in emission spectra as in excitation spectra. This means that there are less differences between fluorescence responses of different phytoplankton classes, and their discrimination is more difficult. Having established the good performance of the SOM with excitation spectra, its performance using only emission spectra was evaluated. The results are described in the following paragraphs. Figure 2.12: An example of a training data set working with emission spectra. Fluorescence emission spectra excited at 490 nm. 60 samples randomly selected representing the seven cultures. The emission range is 535–735 nm, every 1 nm. The EEMs acquired were used again, but now using the training and test data sets as described above: emission spectra (520–735 nm) with 1 nm resolution (216 features) excited at 490 nm and randomly sub-sampled. An example of training samples for this case is presented in figure 2.12. The discrimination can be observed in figure 2.13, and table 2.5 represents the confusion matrix obtained. The averaged TPR index over ten validation runs is 0.7046, the average FPR is 0.0588, and the Kappa index is 0.6343. Although a good classification performance is obtained, there are again some classes that appear mixed. For example, Synechococcus sp. is clearly distinguishable,
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 32 Figure 2.13: An example of the results obtained from randomly chosen training and test samples using emission spectra. Table 2.5: Example of a confusion matrix. Classification of emission spectra from a random selection of training and test samples. Predicted Class Ost Syn Thwii Duna Pl Am Iso Sum TPR FPR True Class Ost 60 0 0 0 0 2 8 0.75 0 Syn 0 80 0 0 0 0 8 1 0 Thwii 0 0 90 0 0 0 9 1 0.037 Duna 0 0 1 54 0 0 10 0.5 0.038 Pl 0 0 1 0 73 0 11 0.63 0.176 Am 0 0 0 0 5 60 11 0.54 0.035 Iso 0 0 0 2 0 0 35 0.6 0.035 while the other classes appear closer and mixed. The reason is that these classes have very similar spectra and they are more difficult to discriminate from emission fluorescence. Repeating the same procedure as in the previous section, we now focus our attention on the evaluation of the performance using the stable samples for training and then classifying samples from the growth stage. Surprisingly, using emission spectra, the results (table 2.6) are better than using excitation spectra stable samples; the Kappa index is equal to 0.5985. From this result and that obtained with excitation spectra, it seems that the pigment changes have a greater effect on the excitation spectra (K=0.4974) than on the emission spectra (K=0.5985), making emission spectra
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 33 more robust to these changes. Several studies use derivative techniques to enhance minute differences between similar signals [35,38]. These techniques have proven to be a powerful tool that is commonly used, for example, in the analysis of hyperspectral data. However, the derivative spectroscopy used to explore these minute features in spectral data is notoriously sensitive to noise [78]. To remove this noise from the hyperspectral data, smoothing techniques are commonly used [79]. It is worth noting that there must be a trade-off between noise removal and the ability to resolve fine spectral details [80]. The following section is devoted to the analysis of the SOM method classification, using as input data the derivative of the spectra. Table 2.6: Confusion matrix. Classification behavior using emission spectra. Samples from the stable growth stage used for training, and samples from the exponential growth stage used for testing. Predicted Class Ost Syn Thwii Duna Pl Am Iso Sum TPR FPR True Class Ost 50 1 0 0 0 4 10 0.5 0 Syn 0 10 0 0 0 0 0 10 1 0 Thwii 0 0 10 2 1 0 0 13 0.77 0.041 Duna 0 0 1 90 5 0 15 0.6 0.028 Pl 0 0 0 0 214 0 16 0.125 0.014 Am 0 0 0 0 0 17 0 17 1 0.275 Iso 0 0 1 0 0 0 45 0.8 0.049 2.3.2.3 Classification applying derivative analysis to Emission Spectra In typical multispectral analysis, each spectral band is considered as an independent variable, a reasonable assumption for multispectral data, but this is not suitable for hyperspectral data. Due to the huge number of bands of the hyperspectral data, it can be treated as spectrally continuous data, and some methods particularly developed for this type of data can be applied. Among these methods, derivative analysis is particularly promising dealing with hyperspectral signals. Derivative techniques enhance minute fluctuations in spectra and separate closely related features. But derivatives are notoriously sensitive to noise. Thus, smoothing or otherwise minimizing random noise is a major issue. In an attempt to increase the performance of the classification using emission fluorescence spectra, derivative analysis was applied to the emission spectra in order to enhance the differences between the fluorescence spectra of each algae. The noise of the signals was reduced using a wavelet denoising
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 34 Figure 2.14: An example of a training data set working with derivative emission spectra. Derivative fluorescence emission spectra excited at 490 nm. 60 samples randomly selected representing the seven cultures. The emission range is 535–735 nm, every 1 nm. applied to both data sets before using derivative analysis. Once obtained the first-order derivative spectra, SOM has been applied in order to evaluate the performance using the derivative preprocessing step to enhance their subtle differences. First of all, the samples were also randomly sub-sampled in order to evaluate the classification indexes for different training and test data sets. An example of the results obtained using the derivative training data is presented in figure 2.14. The discrimination this time was higher. The neurons of the network were properly distributed and the indices obtained (table 2.5) were slightly better than those obtained using excitation spectra: TPR = 0.7574, FPR = 0.0481, and Kappa = 0.7109. Table 2.7: Example of a confusion matrix. Classification of first-derivative emission spectra from a random selection of training and test samples. Predicted Class Ost Syn Thwii Duna Pl Am Iso Sum TPR FPR True Class Ost 8 0 0 0 0 0 0 8 1 0 Syn 0 8 0 0 0 0 0 8 1 0 Thwii 0 0 6 2 1 0 0 9 0.66 0.019 Duna 0 0 0 9 1 0 0 10 0.9 0.057 Pl 0 0 1 1 6 3 0 11 0.54 0.059 Am 0 0 0 0 1 10 0 11 0.9 0.059 Iso 0 0 0 0 0 0 5 5 1 0 In the case of using stable samples for training, the results obtained were TPR = 0.7638, FPR = 0.0611, and Kappa = 0.6803 (table 2.8). As can be clearly seen in the confusion matrices, the
Chapter 2. Phytoplankton Fluorescence Spectra Discrimination 35 greatest problems arose with Pleurochrysis elongata (Pl). The fluorescence properties of Alexandrium minutum and Pl are very similar for this method. Both species like coastal areas where they can form blooms [81–84]. It has also been noticed that coastal species in coccolithophores share pigments unexpected by their phylogeny [85], and if these two cultures are joined into one group, the performance increases considerably. For example, for the confusion matrices found in the last two approximations studied above (using derivative emission fluorescence spectra), if these two species are grouped the Kappa indices are 0.8105 and 0.8439, respectively. Table 2.8: Confusion matrix. Classification behavior using first-derivative emission spectra. Samples from the stable growth stage used for training, and samples from the exponential growth stage used for testing. Predicted Class Ost Syn Thwii Duna Pl Am Iso Sum TPR FPR True Class Ost 10 0 0 0 0 0 0 10 1 0 Syn 0 10 0 0 0 0 0 10 1 0 Thwii 0 0 9 0 4 0 0 13 0.692 0 Duna 0 0 0 10 0 5 0 15 0.666 0.014 Pl 0 0 0 0 3 13 0 16 0.187 0.057 Am 0 0 0 0 0 17 0 17 1 0.26 Iso 0 0 0 1 0 0 4 5 0.8 0 2.3.3 Potential-Support Vector Machines for Phytoplankton Discrimination: Comparison with Self-Organizing Maps In this section Potential-Support Vector Machines is tested for phytoplankton fluorescence spectra discrimination. The results of P-SVM are presented and compared with SOM. This technique has already been introduced above, but just comment that, in this work, it has been used with C-Support Vector classification and a radial basis function kernel. The optimal cost parameters, C, and gamma parameters, g, are determined by performing 8-fold cross validation. Although SVMs were originally developed to perform binary classifications, multiclass classification (more than two classes), is frequently applied, for example, in remote sensing applications. A number of methods to generate multiclass SVMs from the binary SVMs have been proposed. More information on multiclass classification with their advantages and disadvantages can be found in [75]. In this thesis, SVMs binary classifiers are created for all the possible pairs of classes. Thus, the multiclass approach to the problem has been implemented using one-against-one approximation, following the max-wins voting strategy (MWV). In that approach, one binary classifier is constructed for
Chapter 3 Chlorophyll a Fluorescence Peak Analysis The need for covering large areas in oceanographic measurement campaigns and the general interest in reducing the observational costs open the necessity to develop new strategies towards this objective, fundamental to deal with current and future research projects. In this respect, the development of low-cost instruments becomes a key factor, but optimal signal-processing techniques must be used to balance their measurements with those obtained from accurate but expensive instruments. In this chapter, in addition to the study presented in Chapter 2, a complete signal-processing chain to process the fluorescence spectra of marine organisms for taxonomic discrimination is also proposed. This time, the pre-processing methods presented have been designed to deal with noisy, narrow-band and low-resolution data obtained from low-cost sensors or instruments and to optimize its computational cost. It consists of four separated blocks that denoise, normalize, transform and classify the samples. For each block, several techniques are tested and compared to find the best combination that optimizes the classification of the samples. The main difference with the previous chapter is that the signal processing has been focused only on the Chlorophyll a (Chl a) fluorescence peak, since it presents the highest emission levels and it can be measured with sensors presenting poor sensitivity and signal-to-noise ratios. The whole methodology has been successfully validated by means of the fluorescence spectra emitted by five different cultures. The chapter has been organized similar to the previous chapter, and some of the techniques are shared, for this reason, or because are well-known techniques, some of them are just briefly introduced. First of all, 43
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 44 the chapter contains a short introduction explaining the importance of low-cost instrumentation and why the study has been centered in the Chlorophyll a peak. Following, a section introducing the methodologies used for this work is presented. The results are presented next, and finally the main conclusions extracted from the results are exposed. The results obtained in this chapter have been already published in ”Analysis of Discrimination Techniques for Low-Cost Narrow-Band Spectrofluorometers” [26] and ”Optimal processing algorithms for taxonomic discrimination with low-cost narrow-band spectrofluorometers” [27]. 3.1 Introduction Chlorophyll (Chl) fluorescence techniques have been widely used to assess the taxonomic composition of microscopic photosynthetic organisms (phytoplankton) in order to avoid the time constraints imposed by the microscopic analysis of water samples [15]. The basis of fluorometric taxonomic discrimination, as explained above, relies in the specific features of the excitation and emission spectra of each phytoplankton taxonomic group [15,86], and multiple approaches have been used to determine such differences. For instance, the spectral deconvolution analysis, used by Chekalyuk [87] to discriminate between two different organisms, or the self-organizing maps (SOM) technique presented in the previous chapter [24] to classify seven strains from different taxonomic groups of phytoplankton, among others. Nevertheless, these techniques have mostly been tested with accurate and precise data obtained with expensive instruments. This involves an important limitation, since the observational costs spent in infrastructure and instruments in order to obtain high volumes of accurate data in shallow or open water is extremely high, and consumes most part of the money budget available in a research project. In this regard, the concept of ”citizen science” has arisen as an effective methodology to mitigate the expenses while covering large areas with high temporal and spatial resolution measurements [88], but this concept only makes sense through the development of extreme low-cost sensors, as those presented in [89–95]. Reportedly, their accuracy (sensibility, resolution and signal-to-noise ratio (SNR)) is not comparable to the most precise (and consequently, expensive) alternatives, but they present a considerable potential if a correct pre-processing step is performed. Therefore, there is an increasing need for the development of signal-processing strategies able to suitably process the noisy and low-accurate data obtained from instruments based on low-cost sensors.
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 45 In this chapter, the analysis of the discrimination skills of a potential low-cost hyperspectral fluorescence instrument presenting a lower performance in terms of sensibility, SNR and processing capabilities is presented. To this end, three different techniques based on pattern recognition are tested, evaluated and compared to find which one presents the optimal performance considering two main constraints. First, a successful taxonomic discrimination must be obtained even when using as primary information only the highest fluorescence emission levels (if the SNR of the sensor is extremely low, only those levels would be reliable), which correspond to the Chl fluorescence peak (around the 680 nm). This consideration differs from [24,87] and the previous chapter, where the whole optical spectra bandwidth is analyzed, and it is actually feasible assuming that the fluorescence signal in this wavelength range is not only due to the Chl a emission peak, but also the Chl b, c and demission peaks along with additional complement pigments (such as the phycocyanin, whose fluorescence emission is located in the 630-to-660 nm band). Besides, this consideration relaxes the needed sensor’s spectra bandwidth performance. Second, the computational cost needed to develop the algorithms must be optimally reduced in order to decrease the electronic hardware requirements needed to implement the instrument (which will directly influence on its economic cost). In order to deal with these two requirements and considering high levels of noise in the measurement samples, three signal-processing blocks previous to the classification one have been established, accounting for denoising, normalization and transformation of the measured data. The denoising block reduces the noise introduced by the sensor; the normalization block equals the emission contribution measured at different growth states, which improves the discrimination outcomes; and the transformation block transforms and reduces the data dimension, improving the computational-cost efficiency. Thereby, the most convenient technique in each of these three blocks, which, in combination with the best classification algorithm, provides an optimal taxonomic discrimination even when dealing with the two measurement constraints described above, is sought. In order to test the performance of different algorithms in the presented signal-processing chain, the fluorescence spectra of five isolated cultures have been measured at different growth stages. Hyperspectral low-cost fluorescence instruments for in-situ or in-vivo measurements of phytoplankton responses have not been developed yet. Fluorescence sensors or instruments based on low-cost technology are presented in [89,93–95], but their measurements do not exhibit a hyperspectral performance. Therefore, measurements have been firstly obtained with an accurate fluorescence instrument and degraded afterwards in terms of resolution and SNR to emulate the potential lowcost sensor performance. Those measurements are then processed in each block, where well-known
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 46 methods such as moving average, wavelet or principal components, are put into practice along with other algorithms developed in this study specifically designed for this work. This new approach, mainly based on a reliable signal-processing chain, considerably reduces the sensor’s requirements (spectra bandwidth and computational cost) needed to perform a suitable classification. Besides, its conclusive results constitute an important stimulus to develop new and optimal low-cost fluorometers enhancing their discrimination capabilities and encouraging marine research groups to continue studying this field by considerably reducing the instrumentation costs. This chapter is then structured as follows. A brief introduction to the algorithms used in this study is presented in the next Section 3.2. In Section 3.3, measurements from five phytoplankton cultures from different taxonomic groups are used to perform a comparison of the different algorithms. The results presented in this section were processed first with the original data, and later with a degraded version of the measurements in order to simulate the performance of a low-cost sensor. Section 3.4 outlines the conclusions derived from this work. 3.2 Processing Techniques Figure 3.1 shows the block diagram of the four-step signal-processing chain. Three steps before addressing a classification method, where the taxonomic discrimination is performed, are proposed in order to optimize the processing efficiency. Any electro-optical sensor is a noisy source mainly due to the shot and thermal noise, and this is emphasized in low-cost sensors, which usually present a lower performance. Denoising techniques are firstly applied to mitigate the noise effect, considering that a careful attention must be paid in order to avoid the loss of information due to an excessive smoothing. The fluorescence intensity depends upon the cell concentration, the biological growth state, the temperature conditions and the incident light, among other factors, and measurements of the same culture may present significant range variations. Since the classification techniques are usually based on the Euclidean distance between the sample under test and a reference, their objective functions will not appropriately discriminate the samples if such variations are presented within the same culture. Therefore, all measurements must be normalized in a second step in order to make the contribution of their particular features equivalent. Finally, the transformation techniques that adapt the data to increase the discrimination capacity of the classification algorithms, and the reduction of dimension methods that increase the efficiency of the learning algorithms, are
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 47 included in the third step. In the latter, if the classification techniques have to deal only with those wavelengths that are more representative of the features that characterize the culture (obviating redundant information), the computational cost is considerably reduced. Figure 3.1: The four-step signal-processing chain proposed. The whole set of techniques used in each step are presented in Table 3.1, and described in the following subsections. Widely known methods such as moving average, principal components or k-neighbors are used along with other techniques developed and adapted to improve the taxonomic analysis proposed in this study. Moreover, the complete signal-processing chain has been centered in the Chl a fluorescence peak (around 680 nm), which largely simplifies the computational cost that the analysis of the whole hyperspectral data would need. Table 3.1: Algorithms of the four-step signal-processing chain. Denoising Normalization Transformation Classification WMA Min-Max Derivative k-neighbors Savitzky-Golay GSM Genetic-Algorithm* SOM Wavelet* SNV PCA GCS Modified SBN 3.2.1 Denoising In this chapter, a more exhaustive study of noise robustness is done. Optical detectors are subjected to several influences such as optical shot noise (which follows a Poisson distribution), thermal noise (Poisson distribution), read noise (approximately Gaussian), background light from blackbody radiation (Plank distribution), flicker noise (pink power distribution) and technical noise due to various imperfections (which do not follow a specific distribution). The noise-floor level in a measurement is determined by the thermal and the read noises, while the shot noise dominates at high signal values. In a low-performance sensor, it is expected to have significant levels of noise and, in consequence, a poor SNR. Therefore, a denoising block is needed as a first step for the proposed
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 48 processing chain. Three different techniques have been considered to smooth the measurements acquired for this study (see the first column of Table 3.1). These techniques are briefly described below. 3.2.1.1 The Weighted Moving Average Method The weighted moving average (WMA) [96] is the most widely used technique for denoising. In it, the output averaged data vector (y) can be computed as the weighted mean of the nearest 2 ·P wavelengths (Pwavelengths for each side) for each value of the noisy raw data (x), and can be expressed as: y(λ) = 1 P ρ=−Pw(ρ) P ρ=−P w(ρ)x(λ−ρ) (3.1) being wthe weighting factor vector and λthe wavelength. The particular case where all weighting factors are equal to one is usually known as the standard moving average. 3.2.1.2 The Savitzky Golay Method The Savitzky-Golay technique [96] computes a local polynomial regression to approximate the nearest noisy samples using the least squares method, as: y(λ) = P ρ=−P b0(ρ)x(λ−ρ) (3.2) being b0the steady-state Savitzky-Golay filter which coefficients are determined using the leastsquares fit. The main advantage of this approach is that it tends to preserve distribution features such as relative maxima, minima and width, usually flattened with the WMA technique at the expense of not removing as much noise as the WMA. 3.2.1.3 The Wavelet Method As mentioned in the previous chapter, the method for noise reduction proposed here is derived from a wavelet-thresholding algorithm based on Mallat’s scheme [42,97]. A fast wavelet algorithm
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 49 (FWT) that computes the DWT very efficiently was presented in [41]. The technique implemented in this thesis has been done following the steps explained in [44] and shown above in figure 2.2. The original signal is first decomposed into low and high-frequency components by convolutionsubsampling operations using low and high-pass filters directly on the discrete domain. The lowfrequency components (approximation coefficients) keep the global features of the signal, while the high-frequency components (detail coefficients) retain the local features. The decomposition process can be iterated recursively on the approximation coefficients. And, at the last iteration, both approximation and detail coefficients are kept to be used in the reconstruction step. Some of the resulting wavelet coefficients correspond to details in the data set. If the details are small, they might be omitted without substantially affecting the main features of the data set. The idea of thresholding, then, is to set to zero all coefficients that are less than a particular threshold because they are basically caused by noise. These coefficients are used in an inverse wavelet transformation to reconstruct the data set. The signal is transformed, thresholded and inverse-transformed. The technique is a significant step forward in handling noisy data because the denoising is carried out without smoothing out the sharp structures. The result is cleaned-up signal that still shows important details. Another important advantage of this method is that it not only optimizes the mean-square error but also ensures, with high probability, that the denoised signal is at least as smooth as the original [42]. This property is important because other techniques that optimize mean square error, in some cases, introduce artifacts that can affect signal characteristics. Figure 3.2: Growing Spectra Modeling (GSM) example: (a) Fluorescence measurements of a particular culture at different growth states; (b) Fluorescence at 660 nm and 680 nm related to each fluorescence maximum, and linear regression for the two cases.
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 50 3.2.2 Normalization Once the measurements have been denoised, the next step is the normalization. The second column of table 3.1 shows the normalization methods considered in this chapter and described below 3.2.2.1 The Min-Max Method The Min-Max is a simple method of fitting the fluorescence curve into a fixed range. Minimum and maximum values (aand b, respectively) become equal for all the samples, and the normalized curve is obtained with: y(λ) = a+(x(λ)−min(x))(b−a) max(x)−min(x)(3.3) 3.2.2.2 The Growing Spectra Modeling Method The growing spectra modeling (GSM) is a new method that exploits the simplicity of the Min-Max normalization but uses the information of all the values at each wavelength simultaneously in order to increase its robustness. In it, each wavelength fluorescence value of a particular culture and at a specific growth state is compared with its fluorescence maximum. Measurements on different cultures have shown that this relationship is linear at all wavelengths, which allows obtaining an accurate approximation using their linear regression coefficients, as: Flλ(m) = aλm+bλ(3.4) being aλand bλthe linear regression coefficients of the wavelength λ, and Flλits fluorescence evaluated at m. Figure 3.2a shows an example of the smoothed fluorescence measurements of a particular culture at different growth states. As can be observed, the fluorescence maximum presents an important level variability according to different growth states. When all the measurements on a single wavelength are plotted against its fluorescence maximum, a linear relationship, as shown in Figure 3.2b for two particular wavelengths (660 nm and 680 nm), is obtained. The linear regression has also been plotted for these two cases, showing a decrement on the slope when moving away from the maximum. In general, a unitary slope is obtained around 680 nm, and a
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 51 close-to-zero slope around the fluorescence minimum. The model finally uses the linear regression coefficients to compute the normalization factor for each measurement and wavelength. This is done by evaluating Equation 3.4 at two point values, the maximum fluorescence in that measurement (obtaining Flλ1) and the desired (or normalized) maximum fluorescence (obtaining Flλ2). The coefficient obtained from the relationship Flλ1/Flλ2is the normalization factor used to normalize the initial data value. 3.2.2.3 The Standard Normal Variate Method The standard normal variate (SNV) [98,99] is a robust method against noisy data. It is based on the mean and variance (µand σ2, respectively) matching of all the measured samples, as: y(λ) = (x(λ)−µ+µtot)σ2 tot √σ2(3.5) being µtot and σ2 tot the averaged mean and variance of the whole set of samples. A typical approximation is done considering µtot = 0 and σ2 tot = 1. 3.2.2.4 The Modified Scale Based Normalization Mehod The three previous methods significantly distort those signals that are more different from the general pattern, leading, in some cases, to a significant deformation of the small details that characterize the nature of the sample. A more general and flexible version of the SNV method is the scale-based normalization (SBN) introduced in [98]. When applying the wavelet decomposition to a signal, its variance is also faithfully decomposed, allowing a more precise scaling. Variance normalization to 1 (σ2 tot = 1) and mean to 0 (µtot = 0) is performed using only those wavelets that do not contain the high-frequency noise, as: y(λ) = Di(λ) + ... +Dj(λ) σ2 i+... +σ2 j (3.6) being Di(λ)+...+Dj(λ) the j−inoise-free detail functions of the wavelet decomposition (obtained with Equation 2.2), and σ2 i+... +σ2 jtheir variances. The SNV method constitutes a special case
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 58 Table 3.3: RMSE between the covariance of the original data and the covariance of the smoothed data at 684 nm. Algorithm Parameters RMSE WMA (Square window) n=3 0.012 n=7 0.047 n=11 0.102 WMA (Gaussian window) σ1=1.04 0.016 σ2=1.56 0.030 σ3=3.12 0.090 Savitzki-Golay n=13 0.015 n=17 0.028 n=23 0.062 Wavelet thr1 0.032 thr2 0.012 Finally, the wavelet method shows a smaller distance in the hard threshold case than in the soft one. Similar results have also been obtained using different cultures. Table 3.3 gives an idea about the smoothing rate introduced by each algorithm, but it is not decisive when selecting the most suitable one. Further results are shown in Subsection 3.3.4 when using them to classify the samples. Figure 3.6 shows the original and smoothed spectra for the Pl measurements using the Savitzky-Golay method and n=17 as an example. Figure 3.6: Example using Savitzky-Golay for denoising: (a) Original and (b) smoothed spectra of all the Pl measurements obtained with the Savitzky-Golay method and n=17. 3.3.2 Normalization Figure 3.7 shows the four normalization methods proposed in this chapter applied on the Thwi measurements after denoising using the wavelet method and soft threshold.
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 59 Figure 3.7: Different normalizations applied on the denoised Thwi measurements (with wavelet method and soft threshold) using: (a) the min-max method; (b) the GSM method; (c) the SNV method; (d) the modified SBN method. As can be observed, the four spectra plots are quite similar, but slight differences can be appreciated. The spectra obtained with the min-max method, Figure 3.7a, presents flat shapes around the 640 nm and 680 nm with minimum and maximum values, respectively, not seen in any other plot, due to the scaling method. Such distortion may affect the statistical properties present at those wavelengths. As expected, the GSM method improves the spectral shape, as shown in Figure 3.7b, and presents the best normalization below 660 nm. However, all the curves tend to concentrate around a single point in its maximum since it is taken as the reference and it can affect the classification step. The SVN and the modified SBN methods, Figure 3.7c and 3.7d, respectively, present similar curves and do not suffer from any distortion on their fluorescence values. The four algorithms are objectively compared in Subsection 3.3.4 against the denoising and classification methods to determine the most suitable one.
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 60 3.3.3 Transformation and Dimensionality Reduction 3.3.3.1 The Derivative Method The derivative of the denoised and normalized samples (using the wavelet and SBN methods, respectively) of the Thwi culture was obtained using different band separations, as shown in Figure 3.8. As the band separation increases, a smoother curve is obtained. In order to know if the derivatives of the original fluorescence signals contain hidden properties that may facilitate the discrimination process, they are used in Subsection 3.3.4, along with the three classification methods, to compare its taxonomic discrimination results with the ones obtained using the original measurements. Figure 3.8: Derivative of the Thwi samples for different band separation: (a) 5 sampling intervals; (b) 10 sampling intervals; (c) 20 sampling intervals. 3.3.3.2 The Genetic Algorithm Method The first step before using the genetic algorithm is to find the minimum data dimension that keeps a suitable classification efficiency. The maximum likelihood estimator (MLE) technique [105], which uses the principle of maximum likelihood on the distances between close neighbors to group them, was applied on the fluorescence data of the five cultures denoised and normalized with the wavelet (hard threshold) and SBN methods, respectively, obtaining a minimum dimension of 5 bands. Then, the genetic algorithm was used in combination with the k-neighbors classification method (with k=1) to find the value of these five wavelengths. Among the classification methods, the k-neighbors was selected since it does not need a training and thereby it is the fastest one. The results obtained after 20 generations over an initial vector of 100 solutions are 637 nm, 677 nm, 694 nm, 710 nm and 720 nm. Since these results are spaced along the whole bandwidth, it can be concluded that the particular features of each culture are not concentrated in a narrow band but widely distributed.
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 61 3.3.3.3 The PCA Method Figure 3.9 shows the results obtained with the PCA method applied on all the samples. As can be seen, a significant reduction of the data dimension can be applied since the first three components concentrate the 99% of the data variability. The other ones will not significantly contribute to obtain a better classification of the culture. In the next subsection, the first three components obtained with this algorithm are used in combination with the three classification methods to compare its results with previous methods. Figure 3.9: Example of PCA Analysis: Representation of (a) the first 20 eigenvectors, and (b) their percentage of variance. 3.3.4 Classification The comparison between the classification methods (and the different denoising and normalization techniques) has been performed using the confusion matrix and the Kappa index (K), introduced in the previous chapter [24]. As explained, the confusion matrix displays both the number of samples that were correctly and incorrectly classified, and, in the latter case, provides insight into which was the wrong chosen culture. The Kappa index is a measure of the global classification error whose calculation is made from elements of the confusion matrix. In order to maximize the performance of k-neighbors, a previous optimization of the variable k was done finding the value that presents the best classification results, obtaining k= 1. The training for the SOM and GCS classification algorithms, with 32 nodes each, was done using both sampling columns of Table 3.2 in a 5-fold cross-validation technique. Tables 3.4 to 3.6 resume the classification performance of the three classification methods, in combination with the three denoising methods and the four normalization techniques, through the Kappa index. As can be observed, all combinations present
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 62 accurate classifications achieving in some cases a perfect result. In general, the net growing concept seems to present a better performance than the static net of SOM. However, the three tables coincide in pointing out the k-neighbors as the best classification method. Besides, among the denoising techniques, the WMA denoising algorithm (Table 3.4) gives the higher Kappa indices, as the SNV does among the normalization ones. Therefore, an optimal solution for the signal-processing chain is obtained when using these two algorithms (WMA and SNV) in combination with the k-neighbors method. Table 3.4: Kappa indices obtained with the three classification techniques and the four normalization methods when denoising with the WMA method (Gaussian window with α2). Min-Max GSM SNV SBN k-neighbors 0.994 1 1 1 SOM 0.925 0.987 0.994 0.994 GCS 0.974 0.991 10.974 Table 3.5: Kappa indices obtained with the three classification techniques and the four normalization methods when denoising with the Savitzky-Golay method and n=13. Min-Max GSM SNV SBN k-neighbors 0.994 1 1 1 SOM 0.918 0.987 0.975 0.994 GCS 0.965 0.991 10.982 Table 3.6: Kappa indices obtained with the three classification techniques and the four normalization methods when denoising with the wavelet method (the letter denotes the use either soft or hard threshold). Min-Max GSM SNV SBN k-neighbors 0.984s/0.994h 0.994s/0.994h 1s/1h 1s/1h SOM 0.919s/0.931h 1s/0.975h 0.981s/0.987h 0.994s/0.994h GCS 0.965s/0.965h 1s/1h 0.991s/1h 0.982s/0.965h Table 3.7 shows the Kappa index obtained with the three classification methods and the three transformation methods when considering the WMA (Gaussian window with α2) as a denoising method and the modified SBN to normalize. The optimal performance is obtained again with the k-neighbors method using only the five bands given by the genetic algorithm (as expected since the genetic algorithm uses the k-neighbors method to find the suitable wavelengths) or in combination with the first three components given by the PCA method. SOM gives its best result using the five bands of the genetic algorithm and, in contrast, GCS gives its one in combination with the PCA.
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 63 Table 3.7: Kappa indices obtained with the three classification techniques and the three transformation methods, considering the WMA (Gaussian window with α2) as a denoising method and the modified SBN as a normalization method. Derivative Genetic PCA (First Order) Algorithm k-neighbors 0.994 1 1 SOM 0.962 0.994 0.987 GCS 0.974 0.982 0.991 All these results reinforce the idea that an important reduction of the data dimension can be applied without much loss of performance, which involves a reduction of the needed computational cost. In order to evaluate the degree of optimization, Table 3.8 shows the time employed on the execution of the previous example with and without a reduction of the data dimension using an Intel Pentium 4 at 3 GHz, with a 1 GB RAM and running a Windows 7. As can be seen, the time needed to complete the signal processing is reduced between 30%-33% in the SOM case, 8%-9% in the GCS case, and 20%-24% in the k-neighbors case. Even without the further computational-cost reduction, the k-neighbors algorithm is already much optimal than SOM and GCS algorithms, since training is not necessary, and therefore preferred from this point of view. On the other hand, the derivative of the original spectra does not seem to improve in a significant way the classification techniques, and its performance worsens proportionally to an increasing derivative order. Table 3.8: Computational cost expressed in terms of execution time (in seconds), considering the WMA (Gaussian window with α2) as a denoising method and the modified SBN as a normalization method. Standard Genetic PCA Algorithm k-neighbors 1.13 s 0.90 s 0.86 s SOM 19.91 s 14.10 s 13.40 s GCS 106.06 s 98.10 s 96.87 s Tables 3.9 and 3.10 show the averaged confusion matrices (due to the five-fold cross-validation) for the worst solutions obtained in Tables reftab:KappaResultsdenWMA to 3.8, that is, the one obtained with the SOM classification method when using the Savitzky-Golay and the Min-Max methods to denoise and normalize (Table 3.9), and the one obtained with the SOM algorithm when using the wavelet method with a soft threshold and the Min-Max method to denoise and normalize (Table 3.10). Thus, those samples that are more difficult to be suitable classified can be identified. In the Savitzky-Golay case (Table 3.9), some Pl samples are classified as Duna, but the major
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 64 error is produced with the Iso samples classified as Amin. In the wavelet case (Table 3.10), some Duna samples are classified as Pl, but, again, a considerable error is produced with the Iso samples classified as Amin. Both results show that small similarities exist between Duna and Pl samples and an important likeness between Iso and Amin. Table 3.9: Averaged confusion matrix obtained with the SOM classification method when using the Savitzky-Golay and the Min-Max methods to denoise and normalize. Results of the confusion matrix have been averaged due to the five-fold cross-validation. Thwi Duna Pl Amin Iso Thwi 8 0 0 0 0 Duna 0 8 0 0 0 Pl 0 0.2 7.8 0 0 Amin 0 0 0 8 0 Iso 0.2 0 0 2.2 5.6 Table 3.10: Averaged confusion matrix obtained with the SOM algorithm when using the wavelet (soft threshold) and the Min-Max methods to denoise and normalize. Results of the confusion matrix have been averaged due to the five-fold cross-validation. Thwi Duna Pl Amin Iso Thwi 8 0 0 0 0 Duna 0 7.8 0.2 0 0 Pl 0 0 8 0 0 Amin 0 0 0 8 0 Iso 0 0 0 2.4 5.6 3.3.5 Effect of noise in the Classification The fluorescence measurements used in the previous subsections were obtained with an accurate spectral resolution of 1 nm and using a slit width of 4 nm. In order to simulate the performance of a low-cost fluorometer, the signal quality was degraded by reducing its spectral resolution by a factor of 2 (since the monochromatic filter bandwidth was of 4 nm, no loss of information is produced) and its SNR by adding noise. Since the measurement’s noise-floor level is mainly described through a white Gaussian distribution, as stated in Subsection 3.2.1, the added noise followed a white Gaussian distribution with zero mean and a variance of 0.03. Figure 3.10 shows an example before and after this signal degradation.
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 65 Figure 3.10: Example of signal degradation:(a) Measured Thwi fluorescence, and (b) Thwi fluorescence after reducing the spectral resolution and adding noise, which follows a white Gaussian distribution with zero mean and a variance of 0.03, to emulate the measurements obtained with low-cost sensors. The algorithms of the signal-processing chain described in this chapter were applied on the degraded samplings to determine if a suitable classification performance could be obtained under such constraints. The WMA denoising method with a Gaussian window (and using α2) showed to be the most successful one, as seen in Tables 3.4 to 3.6. Therefore, this method was selected to compare the three classification methods against the four normalization techniques, as seen in Table 3.11. Again, the best result is obtained when using the SNV normalization method and the k-neighbors classification one, achieving almost a perfect classification. Table 3.11: Kappa indices obtained with the three classification techniques and the four normalization methods when denoising with the WMA method (Gaussian window with α2). Min-Max GSM SNV SBN k-neighbors 0.750 0.625 0.994 0.794 SOM 0.575 0.575 0.943 0.681 GCS 0.279 0.310 0.680 0.517 In order to determine if this new set of samples can be dimensionally reduced to decrease the computational cost but still obtaining accurate results, the genetic algorithm was applied again to obtain the five frequencies that mostly characterize them. By firstly using only these five frequencies along with the k-neighbors classification method, and secondly the PCA along with the k-neighbors classification method, the Kappa index obtained were of 0.890 and 0.981, respectively. This classification result shows that the best combination of algorithms, in the k-neighbors method case, include the PCA method as the optimal one to reduce the computational cost (the successful classification percentage only drops a 2% and thus an accurate classification result is still obtained,
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 66 whereas the execution time is reduced by a 28%). Additionally, in the PCA case, a study of the classification performance for different levels of noise was performed by modifying the variance of the Gaussian distribution, obtaining the results shown in Table 3.12. As expected, the Kappa index decreases as the variance increases. At a variance beyond 0.2 the classification cannot be considered successful (Kappa index drops far below 0.8). Table 3.12: Kappa indices obtained with the WMA method to denoise, the SNV method to normalize, the PCA method for a dimensional reduction and the k-neighbors to classify, by using the degraded samples with four different variances (σ) of the noise Gaussian distribution. σ=0.03 σ=0.06 σ=0.1 σ=0.2 Kappa index 0.981 0.937 0.850 0.787 The results presented above show that the best combination of algorithms in the signal-processing chain that showed an optimal classification with the accurate fluorescence measurements, is also the best combination when handling with low-accurate data (degraded in terms of resolution and SNR). As stated in Table reftab:KappadenoisingWMA, this combination includes the WMA method to denoise, the SNV method to normalize, the PCA method for a dimensional reduction and the kneighbors to classify the measurements (see Figure 3.11), reaching a discrimination performance with a Kappa index of 0.981 (when σ= 0.03). This is also the best solution in terms of computational cost, since the k-neighbors algorithm presents the lowest execution time (5% of the total SOM execution time and 1% of the whole GCS execution time), and, after adding the dimensional reduction block, the execution time is further reduced by a 24% (as seen in Table 3.8). Taking into account the results presented above, it has been proven the initial hypothesis of a feasible taxonomic classification using only the narrow 630-to-730 nm band of fluorescence emissions, which corresponds to the Chl a peak, since other cellular contents (which differs depending on the strain) also modify its spectral shape. With this, the two constraints given by the measurements obtained with low-performance sensors, i.e., high noise levels in a narrow bandwidth and an optimal computational cost, have been dealt. To conclude, despite the fact that a limited number of samples obtained from only five different cultures constituted the whole dataset, the results presented in this section are significantly accurate, and constitute an optimistic beginning to continue working in this direction, either designing more accurate algorithms or improving the current ones. Besides, they stimulate the investment in the development of a new hyperspectral low-cost sensor with discrimination capabilities cen-tred in the Chl a peak spectral range. It should be noted that, if a new low-cost sensor was actually developed, its measurement uncertainty (regarding to spectrometric
Chapter 3. Chlorophyll a Fluorescence Peak Analysis 67 and radiometric errors) should be perfectly characterized and evaluated to see if the performance of the signal processing presented in this chapter is severely degraded. Figure 3.11: Resulted optimal four-step processing chain: Optimal algorithms for the four-step signal-processing chain designed to deal with low-accurate narrowband fluorescence measurements. 3.4 Conclusions Current constraints in money budget for research projects have motivated the development of new low-cost technologies. That, in consequence, requires an effort to extract as much information as possible from instruments exhibiting a low performance. In this chapter, an optimal-computationalcost signal-processing chain designed to deal with fluorescence measurements featuring poor signalto-noise ratios, low resolution, narrow bandwidth and, therefore, suitable for low-cost sensors and instruments, has been presented. The main objective of this research was focused in finding that combination of algorithms that optimizes the instrument performance when discriminating between different taxonomic cultures of phytoplankton species present in marine environments considering two constraints. The first one was given by the potential low-performance of the sensor, which limits the measurable spectra (the lowest fluorescence emissions may be found under the noise floor level of the sensor). Thereby, only the highest values of fluorescence emissions, that is, those close to the Chlorophyll a peak placed around the 680 nm, were considered. The second constraint was given by the processing limits of a low-cost hardware. To minimize the signal-processing requirements, algorithms should exhibit an optimal performance in terms of computational cost. In order to fulfill with these two requirements, the signal-processing chain was established in four separated blocks, each one with a different processing function, which include the denoising, the normalization, the transformation and the classification of the input data. The denoising techniques were used to smooth the noisy data, the normalization methods to make the hyperspectral signature of cultures measured at different growth stages equivalent, the transformation block to modify and reduce the
Chapter 4. Detecting phytoplankton species in a mixed sample 74 Figure 4.2: Schematic representation of a Non-Linear Mixing example (Image extracted from [113]). hand, in a non-linear mixing scenario, the model for the scattered light is much more complex and typically application dependent. There exist particular situations in which a non-linear model can be approximated with a linear one. Although the application case faced in this chapter is probably a non-linear scenario, and there are some studies dealing with non-linear problems [119,120], it is much more complicated to solve and it would imply a complete thesis research. Thus, before facing this scenario, it is always suggested to prove if the problem can be approximately solved by a linear mixing model. To accomplish this goal, two different approaches are presented in this chapter: •Use of laboratory mixed samples: In this case, optical responses of phytoplankton cultures are acquired in the laboratory, as well as fluorescence responses of mixtures obtained merging these phytoplankton cultures. •Use of a radiative transfer model: For this study, a radiative transfer model implemented in a commercial software (Hydrolight [121]) is used to simulate several scenarios of bloom (single algae simulation) and mixture (two phytoplankton species mixed). Simulation outputs are used for testing unmixing techniques This chapter is then organized as follows: Next section introduces the processing techniques used for this study. Linear Mixing Model (LMM) is introduced to better understand how it works and
Chapter 4. Detecting phytoplankton species in a mixed sample 75 it is applied. After that, both approaches are presented. Generated data for each approach is explained and some conclusions are extracted from the results. Some of the results presented in this chapter were presented in [28]. 4.2 Linear Mixing Model (LMM) As it has been already shown during this thesis, each phytoplankton specie produces a different fluorescence spectra, since they are composed of different pigments that provide them their characteristic colors and different spectral responses, depending on the source that illuminates them (intensity and wavelength). This property allows to distinguish and classify the different species present in the water [24]. In marine environments more than one specie is present in the water, so the measurement equipment will not acquire just a pure fluorescent compound but the contribution of all the species present in the specific environment, as well as other fluorescent particles. The aim of this chapter is then to determine if it is possible to distinguish the proportion of each specie in a measurement using a linear unmixing method. As mentioned during the thesis, several techniques use phytoplankton fluorescence spectroscopy to discriminate between different phytoplankton groups [1,17]. Some of these techniques can achieve a high taxonomic discrimination but they are based on measurements that require excitation at different wavelengths. During this thesis, and the results presented in [24,26], the possibility to use the information contained in emission fluorescence spectra to discriminate between several phytoplankton species has been evaluated. Self-Organizing Maps (SOM), as well as other classification techniques such as P-SVM, k-neighbors, Growing Cell Structures, etc. were used and their performance were presented as feasible techniques to use in those studies, in which time acquisition is an important constraint, e.g. mobile platforms for high spatial resolution measurements. In this chapter, we go one step further. The fluorescence of different mixtures has been acquired and a first approach using the Linear Mixture Model (LMM) [114] is analyzed in order to evaluate how this model can be applied to this kind of data. LMM is herein briefly described and finally, the results of the unmixing using this linear approach are presented and discussed. There exist other techniques such as Independent component analysis (ICA) [122] or independent factor analysis (IFA) [123] which have been used by many authors to unmix hyperspectral data.
Chapter 4. Detecting phytoplankton species in a mixed sample 76 However, they are based in the assumption of source statistical independence and this is not satisfied in spectral applications [113], since the source are fractions and, thus, non-negative and sum one. That is why methods based in ICA or IFA have important limitations in this field. In [124] the impact of this source statistical dependence is addressed. In this study, evidence that the unmixing matrix minimizing the mutual information might be far from the true one is given. In this chapter, and bearing this in mind, linear spectral mixing model is used, and it is explained in the next section. 4.2.0.1 Linear Spectral Mixing Model The basic premise of mixture modeling is that within a given scene, the medium, in a general case, is dominated by a small number of distinct materials that have relatively constant spectral properties. Theses distinct substances are called endmembers and the fractions in which they appear in a mixture are called fractional abundances. In this sense, there should exist a linear relationship between the fractional abundance of the substances present in the medium being monitored and the fluorescence spectra measured by the sensor. This relation in a linear process is called Linear Spectral Mixing Model (LSMM), and it can be written as: x= M i=1 aiSi+w=Sa +w(4.1) where xis the measured spectrum vector (L×1 vector, being Lthe number of features), Sis aL×Mmatrix whose columns are the endmember signatures (L×1 vectors belonging to M endmembers), ais the M×1 fractional abundance vector whose entries are ai(being i= 1, ..., M) and w is the L×1 additive observation noise vector. Inversion algorithms consist on the estimation of the fractional abundances of the constituents present in a mixed sample, taking into account the spectrum acquired and the endmembers spectra. Minimizing the squared-error is one of the most widely used inversion algorithms. Least squares inversion methods are usually used as a first approach. They start from a simple approach and increase in complexity as further assumptions are imposed on the problem. Usually two assumptions or constraints are imposed: the sum of the abundances must be 1 (M i=1 ai= 1), and the abundances must be a positive value below 1 (0 ⩽ai⩽1, i= 1, ..., M). The results shown below present the
Chapter 4. Detecting phytoplankton species in a mixed sample 77 performance of the methodology with and without constraints. The approach from non-constraint least squares solutions to full-constraint followed in ths thesis is presented in [125]. Considering a Linear Spectral Mixing Model (LSMM) problem, described above (eq. 4.1), the assumption of no additional noise, and no constraints, the least squares estimation for the abundances (aLS) is: aLS = (STS)−1STx=S#x(4.2) where S#is the pseudo inverse matrix of S. In order to solve these equations, Singular Value Decomposition (SVD) can be used [126]. The SVD is a well-known eigenanalysis mathematical method of decomposing a matrix Ainto matrices U,Sand V, such that A=USV T, where UTU=VTV=I,Iis the identity matrix, and Sis a diagonal matrix whose diagonal elements are the singular values [127,128]. The previous approach imposed no constraints on the abundance solution. A partial constraint approach can be reach if the constraint of sum to one is applied (M i=1 ai= 1). In this case, the obtention of the abundances aSCLS can be obtained following equation 4.3: aSCLS =PMaLS + (STS)−11[1T(MTM)−1(4.3) where aLS is given by 4.1 and PM =I−(MTM)−11[1T(MTM)−11]−11Ta= (STS)−1STx=S#x(4.4) with 1 = (1,1, ..., 1)T. About the nonnegatively constrained least squares (NCLS) is based on the following optimization problem: Minimize LSE = (Sa −x)T(Sa −x)subject to ai⩾0where i = 1, ..., M (4.5)
Chapter 4. Detecting phytoplankton species in a mixed sample 78 where LSE is the least squares error used as the criterion for optimality and ai⩾0, i = 1, ..., M represents the nonnegativity constraint. The problem is now a set of inequalities that following the equation development shown in [129] the constraint optimization problem results in the next iterative equations given by: aNCLS = (MTM)−1MTx−(MTM)−1λ=aLS −(MTM)−1λ(4.6) λ=MT(x−MaNCLS) (4.7) These two equations 4.6 and 4.7 can be used to solve the optimal solution aNCLS and the Lagrange multiplier vector λ= (λ1, λ2, ..., λ3)T. Finally, following the method considered in [130,131], a full-constraint optimal solution (FCLS) can be obtained by making use of the previous NCLS algorithm, introducing the sum-to-one constraint. This constraint can be included modifying the signature matrix Mby a new one denoted by N, defined by: N= δM 1T (4.8) with 1 = (1,1, ..., 1)T, and a new vector x, denoted by s: s= δx 1 (4.9) The use of δin 4.8 and 4.9 controls the impact of the sum-to-one constraint, and can be used as a tunable parameter. The full-constraint algorithm is obtained replacing Mand xused in NCLS with Nand s.
Chapter 4. Detecting phytoplankton species in a mixed sample 79 4.3 Results and Discussion In this study, two different approximations can be clearly distinguished. First (case I), a study done with samples mixed in the laboratory using several phytoplankton cultures. The second approach (case II) is using a radiative transfer model to simulate different mixed scenarios. Each subsection will introduce the specific data generated for each approach and will show the results obtained. 4.3.1 CASE I: Unmixing phytoplankton fluorescence mixtures from laboratory samples This experiment shows the performance obtained unmixing emission fluorescence phytoplankton spectra acquired in the laboratory. First, the data acquired for this work is explained and after that, the results obtained are shown. Data preparation In order to create a data set with mixtures of different phytoplankton species, the fluorescence emission of approx. 80 mixtures were acquired. The measurements were done with the same spectrometer used during this thesis, an Aminco-Bowman Series 2 luminiscence spectrometer (configured with a 4 nm slit width and a scan speed of 20 nm/s). 59 mixtures of two components and 24 of three were acquired over several days, merging eight different phytoplankton species (Table 4.1), representing the major algae divisions. In this work, and having as a reference the goal of a fast discrimination method, emission fluorescence spectra excited at 490nm were acquired. These fluorescence spectra ranged between 520 - 735 nm, were used as an input data to work with. In order to select the best possible endmembers, they have been chosen as the emission fluorescence of the isolated cultures acquired the same day as the mixture. Figure 4.3 shows an example of the spectra acquired for each phytoplankton specie. These spectra are then used as endmembers in the unmixing process. As it can be seen, the spectra are very similar, unless Cryptomonas and Rhinomonas, both belonging to Cryptophyceae class, which present a tiny peak at 590 nm. From this image, the difficulty of the unmixing step can be appreciated, but it is even more obvious in the next figure. Figure 4.4 shows an example of a mixture spectra and the signature spectra of both phytoplankton species used in this mixture. Table 4.2 summarizes the available mixtures acquired for this study, and the abundances of each constituent.
Chapter 4. Detecting phytoplankton species in a mixed sample 80 Figure 4.3: This figure shows an example of endmember for each phytoplankton specie. Figure 4.4: This figure shows an example of a phytoplankton mixture and the constituent endmembers used. Iso and Amin were mixed in a 50% −50% proportion.
Chapter 4. Detecting phytoplankton species in a mixed sample 81 Table 4.1: Phytoplankton species merged for mixtures. The abbreviation is neccesary to follow the discussion and to understand table 4.2. Species Division Abbreviation Pleurochrysis elongata Prymnesiophyceae Pl (1) Isochrysis galbana Prymnesiophyceae Iso (2) Cryptomonas sp. Cryptophyceae Crypt (3) Rhinomonas reticulata Cryptophyceae Rhin (4) Dunaliella primolecta Chlorophyceae Duna (5) Nannochloropsis oculata Eustigmatophyceae Nano (6) Alexandrium minutum Dinophyceae Amin (7) Thalassiosira weissflogii Bacillariophyceae Thwi (8) Results After the effort of the laboratory work, the study presented in this section was the first approach to unmix phytoplankton mixtures. In this case, an intensive search of different combinations of two and three components is done. For instance, taking a two component mixture, all the possible combinations of two species from the list of eight (Table 4.1) are found (C8 2=8! 2!(8−2)! = 28 possible combinations). Then, LSMM with NCLS approximation was applied for each combination of two endmembers. From these abundances, the estimated mixtured is computed for each combination and the error between the original mixture and the reconstructed from the estimated abundances is computed as the norm of the difference vector. The combination of two endmembers with less error is the winner, and taken as the best fit combination for this specific mixture under study. The same procedure is followed for the rest of two and three mixtures left. In the case of mixtures containing three components, the number of possible combinations is 56. The performance is then computed as (eq. 4.10): performance =ncorrect predictions n total mixtures ∗100 (4.10) Taking into account all the two and three component mixtures measured for this study (Table 4.2) the performance of this approach methodology is about 65%. In order to improve the results, taking as a reference the results presented in the previous chapters, denoising and derivative analysis were perform. This time the results were not the expected, since applying wavelet denoising to the spectra the performance decrease to 61.44%. Derivative analysis does not make it better obtaining just a 60.24%.
Chapter 4. Detecting phytoplankton species in a mixed sample 82 Table 4.2: Summary of mixtures acquired in the laboratory. It is indicated for each mixture the phytoplankton species mixed and the abundance in volume percentage of each one. Day Mixed Species Abundance (%) Day Mixed Species Abundance (%) 1-2 50%-50%, 75%-25% 1-7 50%-50%, 75%-25% 3-4 50%-50%, 75%-25% 2-8 50%-50%, 75%-25% Day 1 5-6 50%-50%, 75%-25% Day 6 3-5 50%-50%, 75%-25% 7-8 50%-50%, 75%-25% 4-6 50%-50%, 75%-25% 1-2-3 33%-33%-33% 1-4-7 33%-33%-33% 1-3 50%-50%, 75%-25% 1-8 50%-50%, 75%-25% 2-4 50%-50%, 75%-25% 2-7 50%-50%, 75%-25% Day 2 5-7 50%-50%, 75%-25% Day 7 3-6 50%-50%, 75%-25% 6-7 50%-50%, 75%-25% 4-5 50%-50%, 75%-25% 4-5-6 50%-25%-25% 6-8-3 25%-50%-25% 3-4-7 33%-33%-33% 1-4 50%-50%, 75%-25% 3-5-8 25%-25%-50% 2-3 50%-50%, 75%-25% 2-6-8 50%-25%-25% Day 3 5-8 50%-50%, 75%-25% Day 8 1-5-7 25%-25%-50% 6-7 50%-50%, 75%-25% 4-7-8 33%-33%-33% 7-8-1 50%-25%-25% 2-6-7 50%-25%-25% 5-7-8 33%-33%-33% 3-4-2 25%-25%-50% 5-7-8 50%-25%-25% 1-5 50%-50% 3-5-7 33%-33%-33% 2-6 50%-50% 1-2-7 50%-25%-25% Day 4 3-8 50%-50% Day 9 2-5-8 25%-50%-25% 4-7 50%-50% 4-6-7 33%-33%-33% 2-4-8 33%-33%-33% 3-6-8 50%-25%-25% 3-4-6 25%-25%-50% 1-2-6 25%-50%-25% 1-2 50%-50% 1-6 50%-50%, 75%-25% 3-4 50%-50% 2-5 50%-50%, 75%-25% 5-6 50%-50% Day 5 3-7 50%-50%, 75%-25% Day 10 7-8 50%-50% 4-8 50%-50%, 75%-25% 2-4-7 33%-33%-33% 3-5-7 50%-25%-25% 2-4 50%-50% 5-7 50%-50% 3-7 50%-50%
Chapter 4. Detecting phytoplankton species in a mixed sample 83 At this point, there seems to exist a good behavior in terms of linearity, but the performance is still poor. For this reason, the next step was grouping the species and see if there exists a meaningful grouping criteria that helps increasing the performance and still offering enough useful information. In this direction, three grouping scenarios were studied: •Grouping species by class. The groups are shown in Table 4.3, where the species are just ordered by their taxonomic class. The results as expected increased the performance reaching the 71.08% (68.67% after denoising the spectra and 66.26% using derivative analysis). Table 4.3: List of species used in this chapter experiment grouped by Class. Class Species Bacillariophyceae Thwi Eustigmatophyceae Nano Chlorophyceae Duna Dinophyceae Amin Cryptophyceae Crypt, Rhin Prymnesiophyceae Iso, Pl •Grouping species by functional algal grups (following Beutler’s proposal [1]). In [1] Beutler groups the phytoplankton species into major funtional algal groups (Table 4.4, which are mainly named by their color. The results as expected increased the performance reaching the 74.70% (72.29% after denoising the spectra and 74.70% using derivative analysis). Table 4.4: List of species grouped by major algal functional groups following the proposal stated by Beutler in [1]. Class Species Green (Chlorophyceae) Duna Blue (Cyanobacteria) X Brown (Dinophyceae, Prymnesiophyceae, Amin, Iso, Pl, Nano, Thwi Eustigmatophyceae,Bacillariophyceae) Mixed group (Cryptophyceae) Crypt, Rhin •Grouping species doing a study of distances among spectra. This time, instead of fixing the groups in advance following certain criteria like it was done in both previous approaches, the relation among the species spectra was studied. A clustering algorithm based on spectral similarity was applied. Hierarchical clustering Analysis (HCA)
Chapter 4. Detecting phytoplankton species in a mixed sample 90 Results This section is divided in two parts. The first one is devoted to show how both parameters (Irradiance Reflectance (R) and Diffuse Attenuation Coefficient (Kd)) change with depth or concentration, in order to check if they can be used as endmembers for phytoplankton detection in the whole water column. The second part presents the results using these parameters to detect the most abundant algae present in the mixed profile. Study of invariance of Rand Kd In order to use Kdor Rto discriminate phytoplankton species in the water column, it is necessary to use reference spectra, or endmembers, representing the spectral signature of each specie. Furthermore, these spectra have to be depth invariant, or one spectrum per depth would be needed, and also invariant to algae concentration. Figure 4.8 shows how these two parameters change their spectral response with depth. All five phytoplankton groups and 300 depth samples for each one are represented. It can be observed that both parameters present an expected slight attenuation with depth, while their shapes are fairly constant. It can be also seen that they present differentiable shapes, an important characteristic in order to discriminate them. Figure 4.9 shows an example of the spectra from just one group for three different concentration levels simulated with HE. It can be seen that Ris less sensitive to changes in concentration than Kd. However, although Kdchanges in value, it seems to maintain its shape, which makes it suitable also to shape analysis techniques. Most abundant algae detection As mentioned previously, in this section a simulation of a triangular phytoplankton distribution in the water column (figure 4.7) is used. This simulation presents parts with only one algae present and another where they are mixed. In this case, similarity indices were used to detect the most abundant phytoplankton specie. The reference spectra were taken from pure culture simulations. Specifically, reference spectra were picked at 7.5m depth and 12 mg·Chl/m3. Similarity indices between reference spectra and mixed samples are computed, and the phytoplankton specie with higher similarity index is selected as the most abundant at a specific depth. Table 4.7 summarizes the results where performance values are between 0 and 1. It shows the percentatge of correct matching of the correct solution in the water column for different combinations of concentration levels between both species.
Chapter 4. Detecting phytoplankton species in a mixed sample 91 Figure 4.8: Example of variations of Rand Kdwith depth. Resulting Rand Kdfrom homogeneous simulations (concentration 6 mg·Chl/m3). All 300 samples of each group are represented. Figure 4.9: Example of variations of Rand Kdwith concentration and depth. Resulting R and Kdfrom homogeneous simulations (concentration 6, 12 and 18 mg·Chl/m3). All 300 samples represented.
Chapter 4. Detecting phytoplankton species in a mixed sample 92 Table 4.7: Performance results of similarity for different mixed profiles. Combinations of concentrations between the first and second peak (bold for Kd). Concentration level (First peak-Second peak) Similarity method 12-12 12-6 18-12 18-18 6-12 6-6 Euclidean 0.37 0.41 0.46 0.41 0.19 0.24 0.62 0.67 0.42 0.3 0.55 0.5 Cosine 0.64 0.52 0.66 0.75 0.61 0.5 0.92 0.83 0.93 0.95 0.93 0.85 Correlation 0.83 0.73 0.82 0.85 0.83 0.71 0.91 0.84 0.93 0.93 0.84 0.81 The results show how Kdobtain better results than Rwhen detecting the most abundant phytoplankton specie present in the mixed simulation. In order to enhance the differences among spectra, derivative analysis has been also applied, as well as wavelet denoising to reduce noise before computing the derivative. The results are slightly better as can be appreciated in table 4.8. Table 4.8: Performance results of similarity for different mixed profiles after derivative analysis. Combinations of concentrations between the first and second peak (bold for Kd). Concentration level (First peak-Second peak) Similarity method 12-12 12-6 18-12 18-18 6-12 6-6 Euclidean 0.74 0.67 0.79 0.84 0.57 0.47 0.85 0.85 0.88 0.85 0.63 0.58 Cosine 0.93 0.93 0.94 0.96 0.91 0.87 0.96 0.94 0.97 0.96 0.96 0.95 Correlation 0.91 0.91 0.93 0.94 0.89 0.85 0.97 0.95 0.97 0.97 0.97 0.97 From these results, it can be concluded that Diffuse Attenuation Coefficient (Kd) presents slightly better results than Irradiance Reflectance (R), and it might be a better variable to use in order to detect species present in a mixed environment. Next experiment is then centered in this coefficient applying Linear Spectral Unmixing to phytoplankton mixture simulations. 4.3.2.2 b) Diffuse Attenuation Coefficient (Kd) simulations Here, the performance of Diffuse Attenuation Coefficient used to unmix phytoplankton spectra is shown. Simulations using Hidrolight-Ecolight were performed and the results, using the Linear Spectral Mixing Mode explained above, presented.
Chapter 4. Detecting phytoplankton species in a mixed sample 93 Data preparation This time the attention has been centered at Diffuse Attenuation Coefficient (Kd). In these simulations, again using Hydrolight-Ecolight software for its Radiative Transfer Model, three different phytoplankton species belonging to three different phytoplankton groups (Table 4.9) are used to obtain nine simulations of mixture vertical profiles (Table 4.10) with distinct abundances. The simulations are 10 m deep, taking a sample each meter, with a wavelength resolution ranged between 352-748 nm every 4 nm (total 100 bands). The concentration this time was set to 10 mg·Chl/m3. Following the same procedure used in the above simulations, optical properties obtained from simulations of the water column using just one compound were used as endmembers. Table 4.9: Phytoplankton groups using Diffuse Attenuation Coefficient (Kd) for unmixing using Linear Spectral Mixing Model (LSMM) Group Bacillariophyceae Dinophyceae Prymnesiophyceae Table 4.10: Summary of mixtures simulated in Hydrolight-Ecolight to be unmixed. It is listed the abundance of each constituent. %spA Bacillariophyceae %spB Dinophyceae %spC Prymnesiophyceae Mixture 1 50% 50% 0% Mixture 2 50% 0% 50% Mixture 3 0% 50% 50% Mixture 4 25% 25% 50% Mixture 5 50% 25% 25% Mixture 6 25% 50% 25% Mixture 7 30% 30% 40% Mixture 8 40% 30% 30% Mixture 9 30% 40% 30% Results In this case, as explained in the data preparation section, a part from phytoplankton water column simulations with pure cultures and a constant profile concentration (used to extract endmembers spectra), there have been also simulated water column mixtures. In contrast with the previous study, the profile concentration of the mixtures was constant as well, following table 4.10. Figure 4.10 shows the spectral signatures, or endmembers, for each phytoplankton group, and their variations with depth.
Chapter 4. Detecting phytoplankton species in a mixed sample 94 Figure 4.10: Kdphytoplankton endmembers at different depths: SpA) Bacillariophyceae, spB) Dinophyceae, and spC) Prymnesiophyceae. The results of the Linear Spectral Mixing Model (LSMM) are presented in figures 4.11,4.12 and 4.13. In these figures the abundances and species predicted are shown. As it can be seen, the prediction is particularly good at the surface and the near surface meters. It almost have a perfect prediction the first 2 and 3 meters, but it is getting worse while depth is increasing. In order to illustrate this, two different error indices are computed. Figure 4.14a presents the error between the real abundance vector and the predicted at each depth. Cosine distance is used for this result, and it can be visually appreciated, in a logarithmic representation, how the error increases with depth. On the other hand, figure 4.14b represents the error between the real mixture under study and the reconstructed mixture from the predicted abundances. In this case, from the abundances predicted with the LSMM, the predicted mixture is reconstructed and the cosine distance between both is computed and represented in this figure. The result corroborate the conclusion that the linear prediction decreases its performance with depth. 4.4 Conclusions This chapter was an exploratory work towards the detection of phytoplankton species present in mixed samples. As explained, in natural water, phytoplankton will not appear isolated. There exist other particles able to contribute in the spectral response, as well as other phytoplankton species. When light interacts with these particles before reaching the sensor, it can affected by different processes and the acquired signal is a combination of all of them. As explained in this chapter, basically two different scenarios can be pictured. Depending on the interaction of the light with the materials present in the scene or sample, the response signal acquired by the sensor can be linear o
Chapter 4. Detecting phytoplankton species in a mixed sample 95 Figure 4.11: Resulted abundances using LSMM: a) Mixtures 1, b) Mixture 2, and c) Mixture 3. On the left, the mixture spectra for each depth on the water column. The figure on the right shows the abundances of each constituent at each depth.
Chapter 4. Detecting phytoplankton species in a mixed sample 96 Figure 4.12: Resulted abundances using LSMM: a) Mixtures 4, b) Mixture 5, and c) Mixture 6. On the left, the mixture spectra for each depth on the water column. The figure on the right shows the abundances of each constituent at each depth.
Chapter 4. Detecting phytoplankton species in a mixed sample 97 Figure 4.13: Resulted abundances using LSMM: a) Mixtures 7, b) Mixture 8, and c) Mixture 9. On the left, the mixture spectra for each depth on the water column. The figure on the right shows the abundances of each constituent at each depth.
Chapter 4. Detecting phytoplankton species in a mixed sample 98 Figure 4.14: Performance error from retrieved abundances: a) Similarity between abundances vectors and predicted abundances, and b) Similarity between predicted mixture and real mixture under study. non-linear, and the study of these data depends on this characteristic. In our case, light interacts with different particles before being acquired, and it is presumably a non-linear scenario. However, since non-linear solutions are usually more complex and time consuming, it is always preferably to explore if there exists any possibility where the problem can be solved with linear solutions. This is the case studied in this chapter. Two different scenarios were studied depending on the data origin. In the first one, an intensive laboratory work was done in order to acquire phytoplankton fluorescence responses of cultures and mixtures. The results using Linear Spectral Mixing Model were not conclusive, although they presented a good performance detecting the presence of a major phytoplankton group. This might be interesting when the target is included in one of these groups. However, due to the important laboratory work necessary to carry out more experiments and better characterize the data acquired, it was decided to continue the research using Radiative Transfer Models, which simulates the interaction of light with particles present in the water column. The use of numerical modeling of the ocean optical properties can help to understand how they react to changes in the water column composition. For this reason, Hydrolight-Ecolight (HE), a well-known software used in the whole community, was used in order to simulate different scenarios to study how two Apparent Optical Properties (AOPs) such as Irradiance Reflectance (R) and Diffuse Attenuation Coefficient (Kd) could be used to detect the presence of an specific group in the water column.
Chapter 4. Detecting phytoplankton species in a mixed sample 99 Rand Kdwere shown fairly invariant with depth and constant in shape when several concentrations were simulated. A simple preliminary approach using similarity indices was used to evaluate if these parameters could be used as reference spectra for unmixing problems in the water column. The results showed that there is a good correlation between the reference spectra and the simulated mixed profile when the most abundant algae is the target. Kdseemed to get better results, and for this reason this parameter was used in the last part of this chapter applying LSMM to unmix simulated data. The results of the last experiment showed how linear mixing models obtained good unmixing performance in the first couple of meters, while the error increased with depth. Although the results are still preliminary, it is a very promising starting point for studying phytoplankton mixes in the ocean because of the difficulty to have references to detect the different components present in natural waters. The big variability of environmental conditions (light attenuation with depth, physiological state of the cell, changing meteorological conditions, etc...) makes the task of obtaining automated methods for discrimination of phytoplankton in situ a very difficult one.
Bibliography [1] Beutler, M.; Wiltshire, K.H.; Meyer, B.; Moldaenke, C.; L¨uring, C.; Meyerh¨ofer, M.; Hansen, U.P.; Dau, H. A fluorometric method for the differentiation of algal populations in vivo and in situ. Photosynthesis Research 2002,72, 39–53. [2] Council, N.; Studies, D.; Board, O.; Plan, C. A Review of the Ocean Research Priorities Plan and Implementation Strategy; National Academies Press, 2007. [3] Dickey, T.D.; Bidigare, R.R. Interdisciplinary oceanographic observations: the wave of the future. Scientia Marina 2005,69, 23–42. [4] Dickey, T.D. The role of new technology in advancing ocean biogeochemical research. OCEANOGRAPHY-WASHINGTON DC-OCEANOGRAPHY SOCIETY2001,14, 108– 120. [5] Dickey, T.D.; Chang, G.C. Recent advances and future visions: temporal variability of optical and bio-optical properties of the ocean. OCEANOGRAPHY-WASHINGTON DCOCEANOGRAPHY SOCIETY2002,14, 15–29. [6] Chang, G.; Mahoney, K.; Briggs-Whitmire, A.; Kohler, D.; Mobley, C.; Lewis, M.; Moline, M.; Boss, E.; Kim, M.; Philpot, W.; others. The new age of hyperspectral oceanography 2004. [7] Dickey, T.D. Emerging ocean observations for interdisciplinary data assimilation systems. Journal of Marine Systems 2003,40, 5–48. [8] McClain, C.R. A decade of satellite ocean color observations*. Annual Review of Marine Science 2009,1, 19–42. 107
Bibliography 108 [9] Fallon, M.F.; Papadopoulos, G.; Leonard, J.J.; Patrikalakis, N.M. Cooperative AUV navigation using a single maneuvering surface craft. The International Journal of Robotics Research 2010, p. 0278364910380760. [10] Fiorelli, E.; Leonard, N.E.; Bhatta, P.; Paley, D.; Bachmayer, R.; Fratantoni, D.M.; others. Multi-AUV control and adaptive sampling in Monterey Bay. Oceanic Engineering, IEEE Journal of 2006,31, 935–948. [11] Gong, G.C.; Wen, Y.H.; Wang, B.W.; Liu, G.J. Seasonal variation of chlorophyll a concentration, primary production and environmental conditions in the subtropical East China Sea. Deep Sea Research Part II: Topical Studies in Oceanography 2003,50, 1219–1236. [12] Desortov´a, B. Relationship between Chlorophyll-αConcentration and Phytoplankton Biomass in Several Reservoirs in Czechoslovakia. Internationale Revue der gesamten Hydrobiologie und Hydrographie 1981,66, 153–169. [13] Gitelson, A.A.; Gurlin, D.; Moses, W.J.; Barrow, T. A bio-optical algorithm for the remote estimation of the chlorophyll-a concentration in case 2 waters. Environmental Research Letters 2009,4, 045003. [14] Yentsch, C.; Phinney, D. Spectral fluorescence: an ataxonomic tool for studying the structure of phytoplankton populations. Journal of Plankton Research 1985,7, 617–632. [15] Cowles, T.; Desiderio, R.; Neuer, S. In situ characterization of phytoplankton from vertical profiles of fluorescence emission spectra. Marine Biology 1993,115, 217–222. [16] Kolbowski, J.; Schreiber, U. Computer-controlled phytoplankton analyzer based on 4wavelengths PAM chlorophyll fluorometer. Photosynthesis: from light to biosphere 1995, 5, 825–828. [17] Zhang, Q.Q.; Lei, S.H.; Wang, X.L.; Wang, L.; Zhu, C.J. Discrimination of phytoplankton classes using characteristic spectra of 3D fluorescence spectra. Spectrochimica Acta Part A: Molecular and Biomolecular Spectroscopy 2006,63, 361–369. [18] Moore, C.; Da Cuhna, J.; Rhoades, B.; Twardowski, M.; Zaneveld, J.; Dombroski, J. A new in-situ measurement and analysis system for excitation-emission fluorescence in natural waters. Ocean Optics XVII, Freemantle, Australia 2004.
Bibliography 109 [19] Cowles, T.; Desiderio, R. Resolution of biological microstructure through in situ fluorescence emission spectra. Oceanography 1993,6, 105–111. [20] Cowles, T.; Desiderio, R.; Carr, M.E. Small-scale planktonic structure: persistence and trophic consequences. OCEANOGRAPHY-WASHINGTON DC-OCEANOGRAPHY SOCIETY1998,11, 4–9. [21] George, R.A.; Gee, L.A.; Hill, A.W.; Thomson, J.A.; Jeanjean, P.; others. High-resolution AUV surveys of the eastern Sigsbee Escarpment. Offshore Technology Conference. Offshore Technology Conference, 2002. [22] Margalef, R. Perspectives in ecological theory. University of Chicago Press 1968. [23] Das, J.; Rajan, K.; Frolov, S.; Pyy, F.; Ryan, J.; Caron, D.; Sukhatme, G.S.; others. Towards marine bloom trajectory prediction for AUV mission planning. Robotics and Automation (ICRA), 2010 IEEE International Conference on. IEEE, 2010, pp. 4784–4790. [24] Aymerich, I.F.; Piera, J.; Soria-Frisch, A.; Cros, L. A rapid technique for classifying phytoplankton fluorescence spectra based on self-organizing maps. Applied spectroscopy 2009,63, 716–726. [25] Aymerich, I.F.; Piera, J.; Soria-Frisch, A. Potential support vector machines and SelfOrganizing Maps for phytoplankton discrimination. Neural Networks (IJCNN), The 2010 International Joint Conference on. IEEE, 2010, pp. 1–5. [26] Aymerich, I.F.; S´anchez, A.M.; P´erez, S.; Piera, J. Analysis of Discrimination Techniques for Low-Cost Narrow-Band Spectrofluorometers. Sensors 2014,15, 611–634. [27] Shez, A.M.; Aymerich, I.F.; Prez, S.; Piera, J. Optimal processing algorithms for taxonomic discrimination with low-cost narrow-band spectrofluorometers. 2015. [28] Aymerich, I.F.; Pons, S.; Piera, J.; Torrecilla, E.; Ross, O.N. Comparing the use of hyperspectral Irradiance Reflectance and Diffuse Attenuation Coeficient as indicators for algal presence in the water column. Hyperspectral Image and Signal Processing: Evolution in Remote Sensing (WHISPERS), 2010 2nd Workshop on. IEEE, 2010, pp. 1–4. [29] Richardson, A.; Risien, C.; Shillington, F. Using self-organizing maps to identify patterns in satellite imagery. Progress in Oceanography 2003,59, 223–239.
Bibliography 110 [30] Ainsworth, E.J.; Jones, I.S. Radiance spectra classification from the ocean color and temperature scanner on ADEOS. Geoscience and Remote Sensing, IEEE Transactions on 1999, 37, 1645–1656. [31] Richardson, A.; Pfaff, M.; Field, J.; Silulwane, N.; Shillington, F. Identifying characteristic chlorophyll a profiles in the coastal domain using an artificial neural network. Journal of Plankton Research 2002,24, 1289–1303. [32] Silulwane, N.; Richardson, A.; Shillington, F.; Mitchell-Innes, B. Identification and classification of vertical chlorophyll patterns in the Benguela upwelling system and AngolaBenguela Front using an artificial neural network. South African Journal of Marine Science 2001,23, 37–51. [33] Talsky, G. Derivative spectrophotometry; Wiley-VCH Verlag GmbH, 1994. [34] Curran, P.J.; Dungan, J.L.; Macler, B.A.; Plummer, S.E.; Peterson, D.L. Reflectance spectroscopy of fresh whole leaves for the estimation of chemical concentration. Remote Sensing of Environment 1992,39, 153–166. [35] Demetriades-Shah, T.H.; Steven, M.D.; Clark, J.A. High resolution derivative spectra in remote sensing. Remote Sensing of Environment 1990,33, 55–64. [36] Pe˜nuelas, J.; Gamon, J.; Fredeen, A.; Merino, J.; Field, C. Reflectance indices associated with physiological changes in nitrogen-and water-limited sunflower leaves. Remote Sensing of Environment 1994,48, 135–146. [37] Philpot, W.D. The derivative ratio algorithm: avoiding atmospheric effects in remote sensing. Geoscience and Remote Sensing, IEEE Transactions on 1991,29, 350–357. [38] BUTLER, W.t.; Hopkins, D. Higher derivative analysis of complex absorption spectra. Photochemistry and Photobiology 1970,12, 439–450. [39] Fell, A.; Smith, G. Higher derivative methods in ultraviolet-visible and infrared spectrophotometry. Analytical Proceedings, 1982, Vol. 54, pp. 28–32. [40] Millie, D.; Schofield, O.; Kirkpatrick, G.; Johnsen, G.; Tester, P.; Vinyard, B. Phytoplankton pigments and absorption spectra as potential biomarkers for harmful algal blooms: A case study of the Florida red-tide dinoflagellate, Gymnodinium breve. Limnol. Oceanogr 1997,42, 1240–1251.
Bibliography 111 [41] Mallat, S.G. A theory for multiresolution signal decomposition: the wavelet representation. Pattern Analysis and Machine Intelligence, IEEE Transactions on 1989,11, 674–693. [42] Donoho, D.L. De-noising by soft-thresholding. Information Theory, IEEE Transactions on 1995,41, 613–627. [43] Donoho, D.L.; Johnstone, I.M. Adapting to unknown smoothness via wavelet shrinkage. Journal of the american statistical association 1995,90, 1200–1224. [44] Piera, J.; Roget, E.; Catalan, J. Turbulent patch identification in microstructure profiles: A method based on wavelet denoising and Thorpe displacement analysis. Journal of Atmospheric and Oceanic Technology 2002,19, 1390–1402. [45] Michie, D.; Spiegelhalter, D.J.; Taylor, C.C. Machine learning, neural and statistical classification 1994. [46] Hughes, G. On the mean accuracy of statistical pattern recognizers. Information Theory, IEEE Transactions on 1968,14, 55–63. [47] Watanachaturaporn, P.; Varshney, P.K.; Arora, M.K. Evaluation of factors affecting support vector machines for hyperspectral classification. the American Society for Photogrammetry & Remote Sensing (ASPRS) 2004 Annual Conference, Denver, CO, 2004. [48] Kohonen, T. Self-organizing maps; Vol. 30, Springer Science & Business Media, 2001. [49] Liu, Y.; Weisberg, R.H. A review of self-organizing map applications in meteorology and oceanography; INTECH Open Access Publisher, 2011. [50] Cavazos, T. Using self-organizing maps to investigate extreme climate events: An application to wintertime precipitation in the Balkans. Journal of climate 2000,13, 1718–1732. [51] Raju, K.S.; Kumar, D.N. Classification of Indian meteorological stations using cluster and fuzzy cluster analysis, and Kohonen artificial neural networks. Nordic Hydrology 2007, 38, 303–314. [52] Tadross, M.; Hewitson, B.; Usman, M. The interannual variability of the onset of the maize growing season over South Africa and Zimbabwe. Journal of climate 2005,18, 3356–3372. [53] Guti´errez, J.; Cano, R.; Cofi˜no, A.; Sordo, C. Analysis and downscaling multi-model seasonal forecasts in Peru using self-organizing maps. Tellus A 2005,57, 435–447.
Bibliography 112 [54] Chang, F.J.; Chang, L.C.; Kao, H.S.; Wu, G.R. Assessing the effort of meteorological variables for evaporation estimation by self-organizing map neural network. Journal of Hydrology 2010,384, 118–129. [55] Risien, C.M.; Reason, C.; Shillington, F.; Chelton, D.B. Variability in satellite winds over the Benguela upwelling system during 1999–2000. Journal of Geophysical Research: Oceans (1978–2012) 2004,109. [56] Liu, Y.; Weisberg, R.H. Ocean currents and sea surface heights estimated across the West Florida Shelf. Journal of Physical Oceanography 2007,37, 1697–1713. [57] Iskandar, I. Variability of satellite-observed sea surface height in the tropical Indian Ocean: comparison of EOF and SOM Analysis. MAKARA SAINS 2009,13, 173–179. [58] Yacoub, M.; Badran, F.; Thiria, S. A topological hierarchical clustering: Application to ocean color classification. In Artificial Neural NetworksICANN 2001; Springer, 2001; pp. 492–499. [59] Telszewski, M.; Pad´ın, X.; R´ıos, A.F.; others. Estimating the monthly pCO2 distribution in the North Atlantic using a self-organizing neural network. Biogeosciences 2009,6, 1405– 1421. [60] Liu, Y.; Weisberg, R.H. Patterns of ocean current variability on the West Florida Shelf using the self-organizing map. Journal of Geophysical Research: Oceans (1978–2012) 2005, 110. [61] Jin, B.; Wang, G.; Liu, Y.; Zhang, R. Interaction between the East China Sea Kuroshio and the Ryukyu Current as revealed by the self-organizing map. Journal of Geophysical Research: Oceans (1978–2012) 2010,115. [62] Soria-Frisch, A. Unsupervised construction of fuzzy measures through self-organizing feature maps and its application in color image segmentation. International journal of approximate reasoning 2006,41, 23–42. [63] Vesanto, J.; Alhoniemi, E. Clustering of the self-organizing map. Neural Networks, IEEE Transactions on 2000,11, 586–600. [64] Vesanto, J.; Himberg, J.; Alhoniemi, E.; Parhankangas, J. SOM toolbox for Matlab 5; Citeseer, 2000.
Bibliography 113 [65] Cover, T.M. Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition. Electronic Computers, IEEE Transactions on 1965, pp. 326–334. [66] Boser, B.E.; Guyon, I.M.; Vapnik, V.N. A training algorithm for optimal margin classifiers. Proceedings of the fifth annual workshop on Computational learning theory. ACM, 1992, pp. 144–152. [67] Cortes, C.; Vapnik, V. Support-vector networks. Machine learning 1995,20, 273–297. [68] Vapnik, V.N.; Vapnik, V. Statistical learning theory; Vol. 1, Wiley New York, 1998. [69] Vapnik, V. The nature of statistical learning theory; Springer Science & Business Media, 2000. [70] Shah, C.; Watanachaturaporn, P.; Varshney, P.; Arora, M. Some recent results on hyperspectral image classification. Advances in Techniques for Analysis of Remotely Sensed Data, 2003 IEEE Workshop on. IEEE, 2003, pp. 346–353. [71] Hochreiter, S.; Obermayer, K. Support vector machines for dyadic data. Neural Computation 2006,18, 1472–1510. [72] Hochreiter, S.; Obermayer, K. Classification, regression, and feature selection on matrix data; Citeseer, 2004. [73] Hochreiter, S.; Obermayer, K. Nonlinear feature selection with the potential support vector machine. In Feature Extraction; Springer, 2006; pp. 419–438. [74] Duan, K.B.; Keerthi, S.S. Which is the best multiclass SVM method? An empirical study. In Multiple Classifier Systems; Springer, 2005; pp. 278–285. [75] Hsu, C.W.; Lin, C.J. A comparison of methods for multiclass support vector machines. Neural Networks, IEEE Transactions on 2002,13, 415–425. [76] Melgani, F.; Bruzzone, L. Classification of hyperspectral remote sensing images with support vector machines. Geoscience and Remote Sensing, IEEE Transactions on 2004, 42, 1778–1790.
Bibliography 114 [77] Zhang, K.L.; Liu, C.M.; Huang, F.Q.; Zheng, C.; Wang, W.D. Study of the electronic structure and photocatalytic activity of the BiOCl photocatalyst. Applied Catalysis B: Environmental 2006,68, 125–129. [78] Tsai, F.; Philpot, W. Derivative analysis of hyperspectral data. Remote Sensing of Environment 1998,66, 41–51. [79] Vaiphasa, C. Consideration of smoothing techniques for hyperspectral remote sensing. ISPRS journal of photogrammetry and remote sensing 2006,60, 91–99. [80] Torrecilla, E.; Aymerich, I.F.; Pons, S.; Piera, J. Effect of spectral resolution in hyperspectral data analysis. Geoscience and Remote Sensing Symposium, 2007. IGARSS 2007. IEEE International. IEEE, 2007, pp. 910–913. [81] S´aez, A.G.; Probert, I.; Young, J.R.; Edvardsen, B.; Eikrem, W.; Medlin, L.K. A review of the phylogeny of the Haptophyta. In Coccolithophores; Springer, 2004; pp. 251–269. [82] S´aez, A.G.; Zaldivar-River´on, A.; Medlin, L.K. Molecular systematics of the Pleurochrysidaceae, a family of coastal coccolithophores (Haptophyta). Journal of plankton research 2008,30, 559–566. [83] Balech, E. Redescription of Alexandrium minutum Halim (Dinophyceae) type species of the genus Alexandrium. Phycologia 1989,28, 206–211. [84] Delgado, M.; Estrada, M.; Camp, J.; Fern´andez, J.V.; Santmart´ı, M.; Llet´ı, C.; others. Development of a toxic Alexandrium minutum Halim (Dinophyceae) bloom in the harbour of Sant Carles de la Rapita (Ebro Delta, northwestern Mediterranean). Scientia marina 1990,54, 1–7. [85] Van Lenning, K.; Probert, I.; Latasa, M.; Estrada, M.; Young, J.R. Pigment diversity of coccolithophores in relation to taxonomy, phylogeny and ecological preferences. In Coccolithophores; Springer, 2004; pp. 51–73. [86] MacIntyre, H.L.; Lawrenz, E.; Richardson, T.L. Taxonomic discrimination of phytoplankton by spectral fluorescence. In Chlorophyll a fluorescence in aquatic sciences: methods and applications; Springer, 2010; pp. 129–169. [87] Chekalyuk, A.; Hafez, M. Advanced laser fluorometry of natural aquatic environments. Limnology and Oceanography: Methods 2008,6, 591–609.
Bibliography 115 [88] Cohn, J.P. Citizen science: Can volunteers do real research? BioScience 2008,58, 192–197. [89] Leeuw, T.; Boss, E.S.; Wright, D.L. In situ measurements of phytoplankton fluorescence using low cost electronics. Sensors 2013,13, 7872–7883. [90] Bardaj´ı, R.; Zafra, E.; Simon, C.; Piera, J. Comparing water transparency measurements obtained with low-cost citizens science instruments and high-quality oceanographic instruments. OCEANS 2014-TAIPEI. IEEE, 2014, pp. 1–4. [91] Kantzas, E.P.; McGonigle, A.J.; Bryant, R.G. Comparison of low cost miniature spectrometers for volcanic SO2 emission measurements. Sensors 2009,9, 3256–3268. [92] Yeh, T.S.; Tseng, S.S. A low cost LED based spectrometer. Journal of the Chinese Chemical Society 2006,53, 1067–1072. [93] Attivissimo, F.; Guarnieri Calo Carducci, C.; Lanzolla, A.M.L.; Massaro, A.; Vadrucci, M.R. A portable optical sensor for sea quality monitoring. Sensors Journal, IEEE 2015. [94] Kissinger, J.; Wilson, D. Portable fluorescence lifetime detection for chlorophyll analysis in marine environments. Sensors Journal, IEEE 2011,11, 288–295. [95] Blockstein, L.; Yadid-Pecht, O. Lensless miniature portable fluorometer for measurement of chlorophyll and CDOM in water using fluorescence contact imaging. Photonics Journal, IEEE 2014,6, 1–16. [96] Orfanidis, S. Introduction to Signal Procesing; Beijing: Tsinghua University Publishing House, 1998. [97] Donoho, D.L.; Johnstone, J.M. Ideal spatial adaptation by wavelet shrinkage. Biometrika 1994,81, 425–455. [98] Barnes, R.; Dhanoa, M.; Lister, S.J. Standard normal variate transformation and detrending of near-infrared diffuse reflectance spectra. Applied spectroscopy 1989,43, 772– 777. [99] Randolph, T.W. Scale-based normalization of spectral data. Cancer Biomarkers 2006, 2, 135–144.