Full text
mathematics Article Data Classification Methodology for Electronic Noses Using Uniform Manifold Approximation and Projection and Extreme Learning Machine Jersson X. Leon-Medina 1,2 , Núria Parés 3,* , Maribel Anaya 4, Diego A. Tibaduiza 5and Francesc Pozo 1,6,* Citation: Leon-Medina, J.X.; Parés, N.; Anaya, M.; Tibaduiza, D.A.; Pozo, F. Data Classification Methodology for Electronic Noses Using Uniform Manifold Approximation and Projection and Extreme Learning Machine. Mathematics 2022,10, 29. https://doi.org/10.3390/ math10010029 Academic Editors: Bo-Hao Chen and Cornelio Yánez Márquez Received: 4 November 2021 Accepted: 17 December 2021 Published: 22 December 2021 Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. Copyright: © 2021 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). 1Control, Modeling, Identification and Applications (CoDAlab), Department of Mathematics, Escola d’Enginyeria de Barcelona Est (EEBE), Campus Diagonal-Besòs (CDB), Universitat Politècnica de Catalunya (UPC), Eduard Maristany 16, 08019 Barcelona, Spain; jersson.xavier[email protected] 2Ingeniería Mecánica Investigación-USTA (IMECI-USTA), Faculty of Mechanical Engineering, Universidad Santo Tomás-Seccional Tunja, Calle 19 # 11-64, Tunja 150001, Colombia 3Laboratori de Càlcul Numèric (LaCàN), Department of Mathematics, Escola d’Enginyeria de Barcelona Est (EEBE), Campus Diagonal-Besòs (CDB), Universitat Politècnica de Catalunya (UPC), Eduard Maristany 16, 08019 Barcelona, Spain 4MEM (Modelling-Electronics and Monitoring Research Group), Faculty of Electronics Engineering, Universidad Santo Tomás, Bogota 110231, Colombia; [email protected] 5Departamento de Ingeniería Eléctrica y Electrónica, Universidad Nacional de Colombia, Cra 45 No. 26-85, Bogota 111321, Colombia; [email protected] 6Institute of Mathematics (IMTech), Universitat Politècnica de Catalunya (UPC), Pau Gargallo 14, 08028 Barcelona, Spain *Correspondence: [email protected] (N.P.); [email protected] (F.P.) Abstract: The classification and use of robust methodologies in sensor array applications of electronic noses (ENs) remain an open problem. Among the several steps used in the developed methodologies, data preprocessing improves the classification accuracy of this type of sensor. Data preprocessing methods, such as data transformation and data reduction, enable the treatment of data with anomalies, such as outliers and features, that do not provide quality information; in addition, they reduce the dimensionality of the data, thereby facilitating the tasks of a machine learning classifier. To help solve this problem, in this study, a machine learning methodology is introduced to improve signal processing and develop methodologies for classification when an EN is used. The proposed methodology involves a normalization stage to scale the data from the sensors, using both the well-known min −max approach and the more recent mean-centered unitary group scaling (MCUGS). Next, a manifold learning algorithm for data reduction is applied using uniform manifold approximation and projection (UMAP). The dimensionality of the data at the input of the classification machine is reduced, and an extreme learning machine (ELM) is used as a machine learning classifier algorithm. To validate the EN classification methodology, three datasets of ENs were used. The first dataset was composed of 3600 measurements of 6 volatile organic compounds performed by employing 16 metal-oxide gas sensors. The second dataset was composed of 235 measurements of 3 different qualities of wine, namely, high, average, and low, as evaluated by using an EN sensor array composed of 6 different sensors. The third dataset was composed of 309 measurements of 3 different gases obtained by using an EN sensor array of 2 sensors. A 5-fold cross-validation approach was used to evaluate the proposed methodology. A test set consisting of 25% of the data was used to validate the methodology with unseen data. The results showed a fully correct average classification accuracy of 1 when the MCUGS, UMAP, and ELM methods were used. Finally, the effect of changing the number of target dimensions on the reduction of the number of data was determined based on the highest average classification accuracy. Keywords: electronic nose (EN); data transformation; data reduction; manifold learning; meancentered unitary group-scaling (MCUGS); uniform manifold approximation and projection (UMAP); extreme learning machine (ELM); odor recognition Mathematics 2022,10, 29. https://doi.org/10.3390/math10010029 https://www.mdpi.com/journal/mathematics
Mathematics 2022,10, 29 2 of 38 1. Introduction An electronic nose (EN), or e-nose, is an electronic device that is used as an artificial olfactory system. In general terms, it is used to identify gases having a wide spectrum of odor patterns by utilizing the interaction with a sensor array. As a system, it takes advantage of its various properties for classification tasks and is composed of an array of sensors, a data acquisition system, and a pattern recognition approach [ 1 ]. These components, including both software and hardware, require the application of odor capturing by a sensor array, signal conditioning, signal processing, data gathering, data preprocessing, and pattern recognition techniques to perform odor classification [2]. Different approaches to model the response of the chemical sensors in electronic noses have been developed [ 3 , 4 ]. Firstly, deterministic models have been developed to reproduce the sensing mechanism in the metal oxide gas sensors based on the adsorption and desorption process carried out in the sensing material [ 5 – 7 ]. Some physicochemical parameters have been used to improve the selectivity of a metal-oxide gas sensor. For example, a time-domain characterization technique is developed in [ 8 ], searching to improve the discrimination ability of an electronic nose fitting a fully analytical model of the electric resistance transient sensor response. Secondly, numerical approaches have been developed to simulate the chemical reaction in the sensing material upon contact with the analyte studied [ 9 , 10 ]. Finally, stochastic models [ 11 , 12 ] have also been developed to reproduce the sensor response and have been linked in the process of classification of substances using electronic nose. Pattern recognition in an EN enables the treatment of signals based on the identification of features that can be used in classification tasks. The signal processing methodology of an EN is composed of different stages, including the application of dimensionality reduction techniques, such as feature extraction or feature selection, to identify the best information, suppress outliers and noises in the collected signals, and decrease the false detection/classification rates [ 13 ]. Although a dimensionality reduction technique can filter redundant information by denoising, some useful information in the original data may be lost, thus affecting the final results of the classification. In addition, the classification method can automatically remove the useless components during the learning process of the sample [14]. Although several studies have been conducted with this aim, as described in the following section, this remains an open research topic because of the option to include new techniques to improve the classification process. To solve this problem, a classification methodology for classifying data from an EN-type sensor is proposed in this study. This methodology consists of several stages developed for EN-type sensor arrays using machine learning and signal processing techniques. Uniform manifold approximation and projection (UMAP) is used as a type of manifold learning to achieve dimensionality reduction, and an extreme learning machine (ELM) is used as a classifier. The methodology is evaluated by using three different EN datasets, and it achieves excellent results. The overall contributions of this study can be summarized as follows: 1. A methodology for classifying data obtained by using a sensor array was developed for ENs. A crucial part of this methodology is data preprocessing, including a data arrangement using an unfolding procedure, min −max or mean-centered unitary group-scaling (MCUGS) for dealing with the different magnitudes of the sensors (data normalization), a reduction of the manifold learning data by applying the supervised UMAP algorithm. 2. The data classification process in the methodology is performed through an ELM classifier, which is a fast and accurate method for discriminating the different classes. A two-stage validation/verification evaluation process is carried out. First, in the training stage, 5-fold cross-validation (CV) is used to prevent an overfitting of 75% of the data, and second, a testing set is formed with the remaining 25% of the data and is used to verify the performance of the methodology with unseen data.
Mathematics 2022,10, 29 3 of 38 3. The methodology was validated using three different datasets for classification: (i) Six distinct gases (dataset #1), (ii) three different qualities of wine (dataset #2), and (iii) three classes of gases (dataset #3). The average classification accuracy was used as a performance measure of the proposed process. The high average accuracy achieved in both training and testing sets indicated the effectiveness of the proposed methodology. 4. The methodology can be used in imbalanced multi-class classification problems, as evidenced on datasets #2 and #3, which exhibited an imbalanced behavior in the number of samples per class. The remainder of this paper is organized as follows. Section 2presents a brief review of the related studies on the various uses of ENs and the available methods established in the literature for data analysis and processing of this type of sensor. Section 3describes the methodology needed to improve the classification of data derived from EN sensors, along with the classification methodology applied and its step-by-step processing. Next, in Section 4, three EN datasets used to validate the classification methodology are described and a brief overview of the data acquisition and transformation methods used in the proposed EN data classification methodology is presented. Next, the data reduction stage using UMAP is described in Section 5. The data classification stage using the ELM method is then detailed in Section 6. Section 7outlines the validation/verification procedure and describes the training/test data split. The classification performance measures are then defined. In Section 8, the experiment results and a discussion are provided. In addition, the main results after the application of the developed methodology for the three datasets are presented, including the normalization, dimensionality reduction, confusion matrix, average classification performance metrics, and tuning parameters for each method. Finally, some concluding remarks are provided in Section 9. 2. Related Work Data preprocessing can improve the classification of different analytes in data from several sensors, including the sensors used by ENs. Different linear and nonlinear methods can be used in the data reduction stage. Although principal component analysis (PCA) [ 15 ] is used as a linear method, in most cases, data have nonlinear characteristics that cannot be identified through a linear method. For this reason, different nonlinear data reduction methods, which are members of the group of manifold learning algorithms [ 16 , 17 ], can be used to deal with the data reduction stage. In recent years, different studies related to the use of manifold learning algorithms in ENs have been conducted. For instance, the modified unsupervised discriminant projection [ 18 ] and Laplacian eigenmaps (LEs) [ 19 ] were used as data reduction methods for the rapid determination of the freshness of Chinese mitten crab, similarly the supervised locality-preserving projection was used to process the feature matrix before applying it to the classifier to improve the performance of an EN [20]. Different perspectives in the literature have been developed to deal with classification problems in ENs [ 21 ]. One such perspective is the machine learning approach, which can be divided in supervised, semi-supervised, or unsupervised methods. It delivers qualitative answers by which different types of analytes are identified [ 13 ]. Different supervised learning algorithms have been used to solve classification tasks in ENs, including support vector machines (SVMs) [ 22 ], artificial neural networks [ 23 ], and a new kernel discriminant analysis [ 14 ]. In addition, the k -nearest neighbor (KNN) algorithm with one neighbor and the Euclidean distance has proven to be efficient [ 24 ] solving multi-class problems, and obtaining high classification rates. There are a significant number of datasets available in the literature related to the classification of EN signals. One such dataset was used in the study by Vergara et al. (2012) [25] , who measured six volatile organic compounds during a 3-year period under strongly controlled operating conditions and using a series of 16 metal-oxide gas sensors [ 26 ]. In
Mathematics 2022,10, 29 4 of 38 that study, the authors used a classifier ensemble as a supervised learning method, to treat the sensor drift problem in ENs. Leon-Medina et al. [ 21 ] developed an EN machine learning classification methodology. In their study, four nonlinear data reduction algorithms were compared in an EN classification task: A kernel PCA, LEs, locally linear embedding, and Isomap. A KNN algorithm was used as the classification model. The results when applying these algorithms reached a classification accuracy of 98.33% after performing a holdout CV. In 2019, Vallejo et al. [ 27 ] reviewed some models for classification and regression tasks usually determined by human perception, such as smell, and they called these measurement techniques soft metrology. In their study, the classification process in soft metrology is composed of several stages such as a database construction and preprocessing, construction of an effective representation space, model choice training and validation, and model maintenance. Different preprocessing techniques for data reduction in big data have been developed; for example, wavelet packet decomposition (WPD) is used in [ 28 ] for selecting features that have maximal separability according to the Fisher distance criterion. The WPD method was compared with fast Fourier transform and autoregressive (AR) models for data transformation. The best classification accuracy was obtained using the WPD method, reaching a value of 90.8%. Several studies on the use of machine learning in sensor arrays such as ENs have recently been conducted. Some drift problems in ENs are normally found owing to the deterioration of the sensors over time. A cross-domain discriminative subspace learning (CDSL) method was developed in [ 29 ]. This domain adaptation-based odor recognition method allows solving recognition tasks in systems of master and slave ENs. In this way, two different configurations of the source and target domains were solved. The maximum average accuracies reached by the CDSL method for the two configurations of the source and target domains were 78.96% and 80.17%, respectively. In 2014, Zhang and Zhang [ 30 ] developed the domain adaptation ELM (DAELM) method to solve the drift problem. This method learns a robust classifier by leveraging a limited number of labeled data from a target domain for drift compensation in ENs. The developed DAELM method was validated using a dataset of an EN sensor array captured during a 36-month period, and two different settings were validated using this method, reaching average accuracies of 91.8% under both settings. In 2020, Kumar and Ghosh [ 31 ] developed a system identification approach for data transformation and reduction in sensor arrays. The equivalent circuit parameter method was compared against a discrete wavelet transform and the neighborhood component analysis, which found a significant reduction in the number of features in the machine learning tasks. The best average prediction accuracy was reached by the neural network regression method, with a value of 98%. A novel bio-inspired neural network [ 32 ] processes the raw data of an EN without any signal preprocessing, feature selection, or reduction. This significantly simplifies the data processing procedure in ENs. The results showed a classification accuracy of 100% in the task of recognizing seven classes of Chinese liquors. It is worth mentioning that although a classification accuracy of 100% is reached, this was only tested for the signals acquired with the electronic nose developed in [ 32 ], in addition, the proposed olfactory neural network has a problem with respect to the parameter tuning. A portable EN instrument was developed in [ 33 ] for qualitative discrimination among different gas samples. A multivariate data analysis of this instrument was conducted using a PCA. Good results were obtained in the qualitative and quantitative tests conducted using this instrument for three industrial gases: Acetone, chloroform, and methanol. Krutzler et al. [ 34 ] showed that the determination of gas concentration is improved through the implementation of sensors with smaller resistance variations when the reference library from one sensor is also used for other sensor elements. In 2012, Brudzewski et al. [ 35 ] used a nonlinear classifier in the form of a Gaussian kernel SVM to classify EN data. A total of 11 classes of coffee brand mixtures were correctly classified with an average error of 0.21%
Mathematics 2022,10, 29 5 of 38 using a differential nose system containing two arrays of semiconductor sensors composed of pairs of the same sensors. An incremental-learning fuzzy model for the classification of black tea using an EN is described in Tudu et al. [ 36 ]. With the use of the approach they developed, a universal computational model can evolve incrementally to automatically include the newly presented patterns in the training dataset without affecting the class integrity of the previously trained system. After a 10-fold CV procedure, an average classification rate of 80% with a standard deviation of 4.969% was obtained. In 2014, Zhang and Tian [ 37 ] developed a multilayer perceptron (MLP) based on a multiple-input, single-output approach for an estimation of the simultaneous concentration of multiple types of chemicals in an EN. The lowest average mean square error of prediction was 3.33%, obtained by a particle swarm optimization method based on the bacterial chemotaxis backpropagation (BP) method. A study describing the extension of neuromorphic methods for artificial olfaction is presented in [ 38 ]. The neuromorphic engineering aims to overcome data processing challenges and reduce the output latency. The neuromorphic approach was applied in [ 39 ] for gas recognition of three target gases, namely, ethanol, methane, and carbon monoxide, with a 100% classification accuracy. A hybrid approach using both convolutional and recurrent neural networks (CRNNs), based on the long short-term memory module, was proposed in [ 40 ]. A deep learning method is well-suited for extracting the valuable transient feature contained at the very beginning of the response curve. The reported accuracy dramatically outperformed the previous algorithms, including gradient tree boosting, random forest (RF), SVM, KNN, and linear discriminant analysis. The CRNN approach reached an average classification accuracy of 98.28% when 4-s signals were used. In 2019, Liu et al. [ 41 ] proposed a multi-task model based on the BP neural network (MBPNN) for EN classification problems. The approach they developed was compared against RF and SVM. Their study showed that the EN is effective for the classification and evaluation of organic green tea. The MBPNN method was satisfactorily used in two datasets of ENs for classification tasks, reaching accuracies of 99.83% and 99.67%. A multi-task learning-long-short-term memory (MLSTM) recurrent network was proposed in [ 42 ] for gas classification and concentration estimation (regression) in ENs. The best classification accuracies obtained with the MLSTM method for each of the three datasets presented were 99.99%, 95.17%, and 100%. Cheng et al. [ 43 ] proposed a solution to the problem of dynamically growing odor datasets while both the training samples and number of classes increase over time. The solution uses a deep nearest class mean (DNCM) model based on a deep learning framework and the nearest class mean method. The results showed that the DNCM model is extremely efficient for incremental odor classification, particularly for new classes with only a few training examples. The best average recognition accuracy was 78.09%, obtained using the DNCM method. In 2020, an AR process associated with an observation of zero-mean Gaussian noise was used to solve drift issues in EN-type sensor arrays [44]. In this approach, each sensor response passes through a separate Kalman filter and a regression technique is used to predict the sensor response. In 2018, Zhang et al. [ 45 ] developed a subspace learning methodology for classifying data originating from a sensor array. A sliding windowbased smooth filter was employed for denoising and feature selection, a local discriminant preservation projection approach was applied for dimensionality reduction, and a kernel ELM was used as a classifier. The best results showed a classification accuracy of 98.22% on a training set using 5-fold CV. In summary, some limitations of the existing methods in the literature that lead to the development of the methodology proposed in this study are: The low flexibility of the models, the need to make the dimensionality reduction model each time a new data is obtained, and the inability to perform a quick classification. To treat these limitations, the methodology for electronic nose signal classification developed in this study includes: (i) The correct verification in three different electronic nose datasets, (ii) the use of the supervised variant of the UMAP method for data reduction allowing for mapping new
Mathematics 2022,10, 29 6 of 38 data to a low dimensional space using the model saved in the training, and (iii) the use of the ELM classifier algorithm with all its capabilities related with an extremely fast learning speed and good generalization performance. 3. Methodology Description This study presents a methodology for classifying signals acquired by EN systems that can yield high classification rates. The classifier is constructed using four stages (see Figure 1). DATA ACQUISITION DATA TRANSFORMATION DATA REDUCTION Step 1.1 Feature extraction Step 1.2 Min-max UMAP Dataset #1 Datasets #2 & #3 Step 2. Unfolding Step 1. MCUGS DATA CLASSIFICATION ELM VALIDATION / HYPERPARAMETER TUNING Figure 1. The data classification methodology for ENs(Electronic Noses) is divided into four stages: Data acquisition, data transformation, data reduction, and data classification. MCUGS, meancentered unitary group-scaling; UMAP, uniform manifold approximation and projection; ELM, extreme learning machine. The first stage is data acquisition, where the EN (as a multi-sensor system) collects data in a three-dimensional matrix: The rows and columns of the matrix contain the data acquired during each experiment/test per time instant for a specific sensor, whereas the third dimension stores the information of each sensor. In the second stage, these data are first properly transformed to consider whether the data captured by each sensor or their associated extracted features can present significant differences in their magnitudes. Owing to the different nature of the considered data, two different normalization approaches have been considered: The min −max normalization and MCUGS. Later, with the use of an unfolding process, the three-dimensional matrix containing the data is stored as a two-dimensional matrix . The third step concerns data reduction, whereas the UMAP method is applied to the high-dimensional EN-transformed data for discarding features that are irrelevant to the classification problem. After data reduction, the low-dimensional data serve as an input for training a machine learning classifier. In this case, the ELM algorithm is used owing to its capabilities, such as better generalization ability, robustness, controllability, and fast learning. Once the classifier is set, given a new experimental sample, the classification algorithm can predict its class. However, before the classifier is used to predict new class samples, it is crucial to evaluate its performance. This evaluation is a challenging task because the available data must be used to both define the classifier and estimate its performance. Here, a two-stage validation/verification evaluation is considered (see Figure 2). During the validation step, a standard 5-fold CV is performed by using a training set containing 75% of the total data (see [ 15 ]). This allows for the suitable tuning of the number of extracted features of the UMAP method and the hyperparameters of the ELM classifier using a GridSearch algorithm, while providing standard estimates of the overall performance of the model. However, although validation performance
Mathematics 2022,10, 29 7 of 38 measures are usually used in the literature to assess the quality of a classification strategy, an extra verification step is added to avoid an overfitting. Once the UMAP model and the ELM classifier are defined, the strategy is tested using the remaining 25% of the data. The key point here is that these test data are completely independent of the training step because they are not involved in the supervised feature extraction step nor in the tuning of the classifier parameters. Local train UMAP supervised with all train data Local train Local train Local train Local test Local train Local train Local train Local train Local train Local train Local train Local train Local train Local train Local train Local train Local train Local train Local train Local train Local test Local test Local test Local test Estimation 1 Estimation 2 Estimation 3 Estimation 4 Estimation 5 Validation performance measure CV GridSearch UMAP projection ELM classification Verification performance measure 1st fold 2nd fold 3rd fold 4th fold 5th fold test (25%) train full data set Figure 2. Model validation/verification using 75% of the data in the 5-fold CV used to tune the hyperparameters and 25% of the data for the final verification. Consequently, two different performance measures are obtained, one for the training set (associated with the final hyperparameters after applying the CV GridSearch) and one for the test set. 4. Datasets Description, Data Acquisition, and Transformation The contributions of this study include the approaches to data preprocessing: How the data are arranged, scaled, transformed, and reduced. In this section, data preprocessing is presented in adequate detail, along with the three different datasets used for the validation of the proposed approach. More precisely, the preprocessing is divided into three phases: Data acquisition, data transformation (feature extraction and min −max normalization for dataset #1, and MCUGS for datasets #2 and #3), and data reduction (Appendix A). Figure 1 illustrates these three phases. 4.1. Dataset #1 The first dataset was created by Vergara et al. [ 25 ] and is composed of 3600 measurements of six volatile organic compounds (acetaldehyde, ethanol, toluene, ammonia, ethylene, and acetone) under strongly controlled operating conditions recorded from 16 metal-oxide semiconductor gas sensors manufactured by Figaro, Inc. (figarosensor.com, accessed on 10 October 2021). 4.1.1. Data Acquisition The EN employed to collect dataset #1 used a test chamber linked to a computercontrolled continuous flow system [ 21 ]. The total flow rate through the detection chamber was set to 200 mL/min. Synthetic dry air was used as the background for all measurements to keep the humidity level constant at 10% relative humidity (measured at 25 ± 1 ◦ C ). The operating temperature of the sensors was set to 400 ° C. The dataset was obtained under laboratory conditions with a controlled atmosphere, which favors those materials of the sensors that interact with the gas in a reversible manner [25].
Mathematics 2022,10, 29 8 of 38 The experiments were carried out using an EN for measuring the behavior of 6 analytes during a 36-month period. The current study considered the 10th batch of the measurements, corresponding to the 36th month. More precisely, batch number 10 corresponded to the last batch of recordings in the study by Vergara et al. [ 25 ]. This batch contained 3600 measurements from the analytes captured during month 36 of the author’s study and was collected 5 months after batch 9. During this 5-month period between the collection of measurements from batch 9 to batch 10, the sensors were turned off and their response capacity was affected owing to a lack of operating temperature [ 46 ], making the data collected during the 10th batch extremely relevant for validating the classification strategies. As previously stated, there were a total of 3600 experiments, divided into six classes. Table 1describes the number of samples per class (gas) in this dataset. The duration, T , of the experiments was not constant, spanning at least 300 s (comprising three phases of measurement: Baseline measurements injecting only pure air, test gas measurements injecting the gas, and a recovery phase) under a sampling frequency of 100 Hz. Therefore, during each experiment, at least 300 s× 100 Hz = 30,000 resistance measurements per sensor were acquired. These collected data were arranged into a three-dimensional matrix with a size of n×M×N , where n= 3600 indicates the total number of experiments, M> 30,000 represents the number of time measures, and N= 16 shows the number of sensors. For the data transformation, the sensor measures were grouped into n×M matrices, X1 raw,X2 raw, . . . , XN raw (see Figure 3). experiments (n) time (M) sensor 1 X1 raw sensor 2 X2 raw sensor 3 X3 raw sensor N−1 XN−1 raw sensor N XN raw sensors (N) Figure 3. Initial arrangement of the collected data. Table 1. Number of measures in batch #10 of the EN dataset #1. Index Class Name Number of Measures 1 Acetaldehyde 600 2 Ethanol 600 3 Toluene 600 4 Ammonia 600 5 Ethylene 600 6 Acetone 600 Total 3600 4.1.2. Data Transformation Feature Extraction: Feature extraction usually plays an important role in data preprocessing for chemo-sensory applications, transforming the raw sensor responses while preserving the most meaningful portion of the information contained in the original sensor signal. For this dataset, two distinct types of features were considered:
Mathematics 2022,10, 29 9 of 38 (a) Two steady-state responses of the sensor element; and (b) Three rising and three decaying transient portions of the sensor response. Therefore, from the M> 30,000 resistance measurements per sensor stored in the columns of matrix Xk raw,k=1, . . . , N, only F=8 features were extracted. Specifically, for a given experiment i= 1, . . . , n and sensor k= 1, . . . , N , let rk i=rowi(Xk raw)∈RM be the i th row of matrix Xk raw collecting all the time profiles of the resistance measures for the given sensor associated with the given experiment. Note that the j th component of this vector rk ij , j= 1, . . . , M , corresponds to the resistance acquired at time t=j/ 100 ∈[ 0, T] . Therefore, Xk raw = (rk ij) is the matrix having a raw resistance value of rk ij as the (i,j)th entry. The two features corresponding to the steady-state response of the sensor are the difference between the maximum and minimum/baseline resistance measurements, ∆Rik =max(rk i)−min(rk i), and its normalized version, ∆nRik =∆Rik min(rk i), where the min(·) and max(·) functions return the minimum and maximum values of a vector, respectively. The second set of six features reflecting the sensor dynamics of the increasing/decaying transient portions of the sensor response are computed by first introducing the exponential moving average (EMA) of the signal. Given a smoothing factor α∈[ 0,1 ] , the EMA of the sensor resistance with respect to the smoothing factor α is a new vector emaα(rk i)∈RM defined as emaα(rk i)1=0 and for l=2, . . . , Mis defined as: emaα(rk i)l= (1−α)emaα(rk i)l−1+α(rk il −rk i,l−1), where emaα(rk i)l is the l st component of the EMA signal. The maximum and minimum values of the EMA signal, Mik α=max(emaα(ri)) and mik α=min(emaα(rk i)) , respectively, are characteristic features of the rising and decaying portion of the sensor response. Therefore, the second set of six features is obtained by computing the minimum and maximum values of the EMA signals for α=0.001, 0.01, and 0.1. Thus, each vector rk i∈RMis mapped into vector ˜ rk i∈R8: ˜ rik = (∆Rik,∆nRik, Mik 0.001,mik 0.001,Mik 0.01,mik 0.01,Mik 0.1,mik 0.1). By computing the abovementioned F= 8 features for each of the experiments and each of the N= 16 sensors in the prerecorded time series, we can map the n×M matrices X1 raw,X2 raw, . . . , XN raw to the n×Fmatrices X1,X2, . . . , XN, where rowi(Xk) = ˜ rk i. Figure 4illustrates a reference time profile of the sensor resistance where the three phases (baseline, test gas, and recovery) can be clearly distinguished along with their associated EMA signals and extracted features.
Mathematics 2022,10, 29 16 of 38 This allows for introducing the following average performance measures: average accuracy =1 l l ∑ i=1 tpi+tni tpi+tni+fni+fpi , (2) average precision =1 l l ∑ i=1 tpi tpi+fpi , (3) average recall =1 l l ∑ i=1 tpi tpi+fni , (4) average specificity =1 l l ∑ i=1 tni tni+fpi , (5) F1-score =2precision ×recall precision +recall, (6) which consider the imbalance problem presented in the EN dataset #2 and the existing differences between the samples of the wine classes (see [61] for more details). As shown in Figure 2, the 5-fold CV produces five different confusion matrices providing five performance measures that are averaged to obtain the final validation performance measure. However, despite the performance measures being nonlinear, the common practice of computing the total performance measure by first adding the five confusion matrices associated with each fold and then computing the performance of the total added confusion matrix is used. 8. Experimental Results and Discussion The results of the developed manifold learning classification methodology for ENs are shown in this section for the three described datasets. These results were obtained using the available online implementations of the UMAP [55] and ELM [62] methods. 8.1. Data Transformation and Unfolding 8.1.1. Dataset #1 The EN used to acquire dataset #1 is composed of 16 sensors. Specifically, four Figaro sensors of each of the following types were used: TGS2600, TGS2602, TGS2610, and TGS2620. After conducting the experiments, acquiring the signals by each sensor, and selecting the corresponding eight features per sensor, we applied the min-max normalization. The data were then unfolded to obtain the two-dimensional data matrix X∈R3600×128 , where each row contains the information regarding all 16 sensors of a particular experiment (see Section 4.4). Figure 6shows the 128 feature vectors for a particular experiment before and after applying the min-max normalization. As shown in the figure, before the normalization, the non-normalized steady-state feature, ∆Rik (first feature per sensor), stands out among the other features, i.e., both the normalized steady-state feature (second feature per sensor) and the dynamic features (features 3–8). After the values are scaled to lie between − 1 and 1, the data are comparable and ready to be used for classification.
Mathematics 2022,10, 29 17 of 38 20 40 60 80 100 120 Feature index 0 0.5 1 1.5 2 2.5 Feature value 105 20 40 60 80 100 120 Feature index −1 −0.8 −0.6 −0.4 −0.2 0 0.2 0.4 0.6 0.8 1 Normalized feature value Figure 6. Two unfolded samples in dataset #1 (ethanol compound in blue and ethylene compound in magenta). Initial feature values, i.e., 8 features per 16 sensors ( left ) and min-max normalized values (right). The steady-state features are highlighted by solid circles. 8.1.2. Dataset #2 As mentioned in Section 4.2, in this study, all of the sensor values of the full time profile were maintained. Figures 7and 8show the data collected in matrix X∈R235×19,980 before and after applying MCUGS normalization. Recall that the non-preprocessed data contain the responses in terms of resistivities over time for each sensor (the values per sensor are placed one after another, with a total of 6 blocks × 3300 values per sensor). As can be seen in the figure, there are clear differences in the magnitudes associated with each of the six sensors; specifically, the values of the fourth sensor are significantly larger. These differences can be eliminated by applying the MCUGS method. The resulting MCUGS normalized data have zero mean (per feature) and both a unitary block and a global deviation. Figure 7. Unfolded data in dataset #2. Initial values ( left ) and MCUGS signal ( right ). All experiments shown are superposed (3330 values per 6 sensor signals for 235 experiments).
Mathematics 2022,10, 29 18 of 38 Figure 8. Three-dimensional plot of the unfolded data in dataset #2. Initial values ( left ) and MCUGS signal ( right ). All experiments are shown superposed (3330 values per 6 sensor signals for 235 experiments). 8.1.3. Dataset #3 Figure 9shows the data collected in matrix X∈R309×10,000 before and after applying MCUGS normalization. The non-preprocessed data contain the measures for each sensor (5000 values for 2 sensors) for all 309 experiments. It can be seen in the figure that, in the MCUGS signals, the differences between the three classes are accentuated and, therefore, are better suited for classification purposes. Figure 9. Unfolded data in dataset #3. Initial values ( left ) and MCUGS signal ( right ). All experiments are shown superposed (5000 values per 2 sensor signals for 309 experiments). 8.2. Validation Step: Tuning the UMAP/ELM Parameters As mentioned in Section 3, a 5-fold CV using a training set containing 75% of the total data was applied to properly tune the number of extracted features of the UMAP method and the hyperparameters of the ELM classifier using a GridSearch algorithm (see Figure 2). A standard method used to find the optimal parameters applies a simple grid search combined with a CV to evaluate the accuracy of the model for each set of candidate hyperparameters. However, a simultaneous exhaustive grid search within the entire grid of the joined UMAP and ELM hyperparameters is prohibitively expensive because we have numerous parameters to be tuned (see Tables 4and 5). Therefore, a simplified strategy is used. First, the set of hyperparameters to be tuned is decreased by fixing some of the hyperparameters to their default values (the UMAP metric (d) is set to Euclidean, n_epochs
Mathematics 2022,10, 29 19 of 38 is set to none, and the spread is set to 1), and therefore, only five parameters have to be tuned: The UMAP n_neighbors k , the UMAP min_dist, the UMAP n_components d , the number of hidden layer nodes of the ELM ˜ d , and the ELM activation function g . Then, the five-dimensional space is separated into five separate one-dimensional spaces, and sequential one-dimensional grid searches are carried out, while keeping the values of the other hyperparameters fixed. Although this approach is generally not optimal, because the hyperparameters are typically interdependent, in the current scenario it provides excellent classification results and the use of more involved techniques is not considered (see [ 63 ]). The optimal one-dimensional searches are conducted using the GridSearchCV function of scikit learn [ 64 ], where the accuracy of the model for each set of parameters is measured using the average accuracy (see Equation (2)). Table 7presents the sequential GridSearch conducted for datasets #1 and #3 along with the final selected parameters. Since the results obtained are shown for the three selected activation functions (tanh, tribas, and hardlim), the GridSearch strategy is not applied to determine the optimal ELM activation function. To determine the other parameters, the ELM activation function is set to its default (tanh). Table 7. One-dimensional sequential GridSearch of the UMAP and ELM hyperparameters (the other parameters are set to their default: Metric ( d ) is set to the Euclidean, n_epochs is set to none, and the spread is set to 1) for datasets #1 and #3. The range for the one-dimensional searches is highlighted in bold. Dataset #1 Dataset #3 kmin_dist d˜ dg k min_dist d˜ dg UMAP n_neighbors k[6,55]0.5 8 100 tanh [6,55]0.5 8 100 tanh UMAP min_dist 16 [0.1,0.9]8 100 tanh 6 [0.1,0.9]8 100 tanh UMAP n_components d16 0.1 [2,20]100 tanh 6 0.5 [2,20]100 tanh ELM hidden layer nodes ˜ d16 0.1 8 [10,200]tanh 6 0.5 8 [10,200]tanh Final Values 16 0.1 8 60 tanh 6 0.5 8 50 tanh The results obtained for dataset #2 are not detailed because a perfect classification was obtained for all the tested parameters. The final parameters for dataset #2 were set to n_neighbors =k=16, min_dist =0.1, n_components =d=8, and ˜ d=100. As can be seen in Table 7, the first parameter that has been tuned is the number of neighbors k of the UMAP neighboring graph when the remaining parameters are fixed. Figure 10 shows the effectivity index associated to the average accuracy, defined as ρ=1−average accuracy, obtained when varying kfrom 6 to 55 for datasets #1 and #3. Since a perfect classification corresponds to an average unitary accuracy, the best results are obtained when ρ is close to zero. As can be seen in the figure, a jump in the quality of the classifier is obtained when using 16 neighbors for dataset #1 and 6, 34, or 35 neighbors for dataset #3, where the average accuracies obtained are 0.99901 and 0.99711, respectively. It can also be seen that, for parameter values larger than 40, an increase in the accuracy can be obtained (even yielding perfect classification results); however, considering the trade-off between computational cost and accuracy, and because the accuracy is also improved by tuning the remaining parameters, settings of k= 16 and 6 are the best option. The results for dataset #2 are not shown because all values of k yielded a perfect classification ( ρ= 0). In this case, k=16 is considered. After the optimal number of neighbors k is fixed, GridSearch is applied with respect to the min_dist parameter that controls the minimum distance between the points in the low-dimensional representation (for the same fixed remaining parameters). Figure 11 shows the effectivity index obtained when varying min_dist within the interval [ 0.1,0.9 ] for datasets #1 and #3.
Mathematics 2022,10, 29 20 of 38 10 15 20 25 30 35 40 45 50 55 UMAP n_neighbours k 10-3 10-2 Effectivity index ( 1 − average accuracy ) Dataset #1 Dataset #3 Perfect classification 99% 99.9% 99.5% Figure 10. Average accuracy vs. the number of neighbors in the UMAP graph k for datasets #1 and #3 (semi-logarithmic plot in the vertical axis). The vertical axis represents the effectivity index, i.e., one minus the average accuracy. Perfect classifications ( ρ= 0) are marked with solid black circles. The optimal selected parameters are highlighted in red. The accuracy range for dataset #1 is from 99.864% to 99.925% and that for dataset #3 is from 96.248% to 100%. 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 UMAP min_dist 10-3 10-2 Effectivity index ( 1 − average accuracy ) Dataset #1 Dataset #3 99% 99.9% 99.5% Figure 11. Average accuracy versus the minimum distance between points in the low-dimensional UMAP representation min_dist for datasets #1 and #3 (semi-logarithmic plot in the vertical axis). The vertical axis represents the effectivity index. The optimal selected parameters are highlighted in red. The accuracy range for dataset #1 is from 99.839% to 99.901% and for dataset #3 is from 96.536% to 99.711%. Again, the results for dataset #2 are not shown because a perfect classification was obtained for all values, taking min_dist = 0.1. Regarding dataset #1, the highest average accuracy of 0.99901 is obtained for values of min_dist of up to 0.6, and the default value of 0.1 is maintained. It is worth noting that the results in this case are not overly sensitive to
Mathematics 2022,10, 29 21 of 38 changes in the parameter value. Finally, for dataset #3, changing the initially prescribed value of min_dist =0.5 provides worse results, and therefore, the value is maintained. The final hyperparameter of the UMAP method to be tuned is n_components, i.e., the dimension of the low-dimensional space, d (number of retained features used for classification). It is worth noting that this parameter, which is the dimension of the feature vector inserted into the ELM classifier, is key to the accuracy of the final classification results (see [ 65 ]). Figure 12 shows the values of the effectivity index when varying the n_components within the interval [ 2,20 ] for datasets #1 and #3. For dataset #3, a jump in quality is obtained for n_components, = 8. Regarding dataset #1, excellent results are obtained for nearly all parameter values and for a value of d= 8. As can be seen in the figure, the supervised version of the UMAP method obtains excellent accuracies for the training dataset even for an extremely low number of retained features (an accuracy of 0.99901 is obtained even for d= 2 in dataset #1, and an accuracy of 0.98845 is obtained for d= 4 in dataset #3, instead of 0.99711 in the case of d= 8). However, because the classifier is needed to classify new unseen data, the value of d= 8 is retained in both cases as a tradeoff between computation cost and accuracy. 2468101214161820 UMAP n_components 10-3 10-2 Effectivity index ( 1 − average accuracy ) Dataset #1 Dataset #3 99% 99.9% 99.5% Figure 12. Average accuracy vs. the dimension of the UMAP low-dimensional space d n_components for datasets #1 and #3 (semi-logarithmic plot in the vertical axis). The vertical axis represents the effectivity index. The optimal selected parameters are highlighted in red. The accuracy range for dataset #1 is from 99.888% to 99.901% and that for dataset #3 is from 95.093% to 99.711%. Finally, a GridSearch was conducted to set the number of hidden layer nodes ˜ d of the ELM method. It is well known that the number of hidden nodes is a key factor for achieving a good performance of the ELM method. The optimal value of the parameter usually depends on the size of the training set, n , the number of input features, d , and the number of output classes, L . Figure 13 shows the sensitivity analysis of the accuracy of the combined UMAP-ELM strategy with respect to the number of hidden nodes of the ELM classifier. The selected parameters are ˜ d= 60 and 50 for datasets #1 and #3, respectively, with associated final accuracies of 99.901% and 99.711%, respectively. As depicted in the figure, the combined approach allows excellent accuracies to be obtained even for an extremely low number of hidden nodes as opposed to the standard large number of hidden nodes required in other existing approaches (see for instance [ 45 ]). It is also worth noting that, for both datasets, all reported accuracies are above 90% (excluding the value ˜ d= 1 for dataset #1, where an accuracy of 88% was reached).
Mathematics 2022,10, 29 22 of 38 ñ 20 40 60 80 100 120 140 160 180 200 ELM d 10-3 10-2 10-1 Effectivity index ( 1 − average accuracy ) Dataset #1: n=2700, d=8, L=6 Dataset #3: n=231, d=8, L=3 90% 95% 99% 99.5% 99.9% Figure 13. Average accuracy vs. the number of hidden layer nodes ELM ˜ d for datasets #1 and #3 in the training set (semi-logarithmic plot in the vertical axis). The vertical axis represents the effectivity index. The optimal selected parameters are highlighted in red. The accuracy range for dataset #1 is from 88.333% to 99.901% and that for dataset #3 is from 91.053% to 99.711%. Although it has been reported that the selection of the UMAP and ELM methods has a significant impact on the performance, it is worth noting that the combined approach presented here is extremely robust. Figures 10–13 show the results of the sensitivity analysis of the combined UMAP-ELM strategy proposed, showing the accuracies with respect to the variation of the parameters. It can be seen in the figures that, for the first three parameters studied, the accuracies are all above 95%, and regarding the number of hidden nodes, all values ranging from 20 to 150 also ensure an accuracy of 95%. 8.3. Classification Results This section shows the validation and verification performance measures obtained for the three considered datasets (see Figure 2). 8.3.1. Dataset #1 This dataset contains 3600 samples from six different analytes, where each sample has 128 initial features. After normalizing and unfolding all of the data, we applied a 3:1 partition of the total samples: 2700 samples were selected for the initial training (450 for each gas) and 900 samples were kept for the final testing (150 for each gas). The training samples were used in the supervised UMAP manifold learning algorithm to reduce the number of features to be used for classification to only d= 8 (see Table 7for the details of the UMAP hyperparameters used). Figure 14 illustrates the first three features extracted using the UMAP method in dataset #1 for the training samples and the test samples, separately. As can be seen in the figure, the supervised UMAP allows the successful clustering of most samples of the training set, even though a clear visual separation cannot be achieved. A careful look at the eight features provided by the UMAP method provides similar results to those shown in Figure 14. Most of the samples seem to be clustered however, some remaining overlapping occurs with no clear visual separation. Therefore, the use of a machine learning classifier algorithm is crucial to be able to properly classify all of the samples. When the learned UMAP mapper is used to extract the features of the test dataset, it can be seen that most new points also follow the same pattern, even though the samples do not cluster as clearly.
Mathematics 2022,10, 29 23 of 38 Acetaldehyde Ethanol Toluene Ammonia Ethylene Acetone −4 −2 0 2 4 6 8 10 12 3rd 15 10 2nd 15 510 1st 5 00 −5 −4 −2 0 2 4 6 8 10 12 3rd 15 10 2nd 15 510 1st 5 00 −5 Figure 14. Scatter plot of the first three UMAP features of the data in dataset #1 (the training set is shown on the left and the test set is shown on the right ). The six different analytes are shown in different colors. The classification results obtained by using a 5-fold CV on the training dataset with the hyperparameters reported in Table 7are summarized in the total added confusion matrix shown in Table 8. This table also shows the classification results on the test dataset for the same value of the parameters. As can be seen in the table, all the training samples from acetaldehyde, ammonia, and acetone are perfectly recognized and only eight samples from among the 2700 are misclassified (two ethylene samples, three ethanol samples, and three toluene samples), yielding an accuracy of 99.90%. Regarding the classification of the 900 samples of the test data, in this case, the toluene and ethylene samples are perfectly recognized and only 17 samples are misclassified, yielding an accuracy of 99.37%. Tables 9and 10 show all the performance metrics associated with both the validation and verification steps, and the effect of the choice of the activation function of the ELM classifier. As expected, because the parameters of the methods were trained using these data, excellent performance metrics were obtained when the combined UMAP-ELM classifier was used in the training set for the hyperbolic tangent (tanh) activation function. However, it can be seen that the same parameters also provided good results for the other two activation functions (tribas and hardlim). In addition, the accuracies obtained in the validation step with the unseen test dataset were excellent. Table 8. Confusion matrix for the UMAP-ELM training and test sets for dataset #1 (for the ELM activation function tanh). Predicted Class (Train Set) Predicted Class (Test Set) Acet. Etha. Tolu. Ammo. Ethy. Acet. Acet. Etha. Tolu. Ammo. Ethy. Acet. Actual Class Acetaldehyde 450 00000148 0 0 20 0 Ethanol 0 447 10200144 301 2 Toluene 0 3447 00000150 0 0 0 Ammonia 0 0 0 450 00030147 0 0 Ethylene 0 101448 00000150 0 Acetone 0 0 0 0 0 450 14010144
Mathematics 2022,10, 29 24 of 38 Table 9. Validation (train) classification performance measures of the ELM classifier for three different transfer functions of the ELM algorithm in dataset #1. Validation Performance Measures (Training Set) Accuracy Precision Specificity Recall F1 Score ELM tanh 0.99901 0.99704 0.99941 0.99704 0.99704 ELM tribas 0.95333 0.88147 0.97200 0.86000 0.87060 ELM hardlim 0.99864 0.99593 0.99919 0.99593 0.99593 Table 10. Verification (test) classification performance measures of the ELM classifier for three different transfer functions of the ELM algorithm in dataset #1. Verification Performance Measures (Test Set) Accuracy Precision Specificity Recall F1 Score ELM tanh 0.99370 0.98117 0.99622 0.98111 0.98114 ELM tribas 0.95741 0.87711 0.97444 0.87222 0.87466 ELM hardlim 0.99333 0.98029 0.99600 0.98000 0.98015 8.3.2. Dataset #2 This dataset contains 235 samples of wine with three different qualities, where each initial sample has 19,980 features. After the MCUGS was applied and all of the data were unfolded, a 3:1 partition of the total samples was achieved: 176 samples were selected for the initial training (32,38, and 106 for the high-, average-, and low-quality samples, respectively) and 59 samples were kept for the final testing (11,13, and 35 for the high- , average-, and low-quality samples, respectively). As mentioned in Section 8.2, the combined UMAP-ELM strategy provides a perfect classification on the training set for all the selected parameters when using the tanh activation function. Note that, although some imbalance in the data is present, this does not indicate a poor predictive performance of the model for the minority classes. Figure 15 illustrates the first three features extracted using the UMAP method in dataset #2 for both the training and test samples. As can be seen in the figure, the capabilities of the UMAP method for clustering the data are evident owing to the perfect separation of the wine-quality classes as evidenced for both the training and test datasets. This clustering explains the perfect classification results obtained using the ELM classifier. 10 2nd 5 1 2 3 4 3rd 5 6 0 7 1st 50 10 15 10 2nd 5 1 2 3 4 3rd 5 6 0 7 1st 50 10 15 High Quality Wine Average Quality Wine Low Quality Wine Figure 15. Scatter plot of the first three UMAP features of the data in dataset #2 (the training set is on the left and the test set is on the right). The three qualities of the wine are shown in different colors. The performance measures obtained for dataset #2 using three different activation functions on the ELM method are given in Tables 11 and 12. As mentioned, the tanh activation function allows to perfectly classify both the training and test datasets. The
Mathematics 2022,10, 29 25 of 38 results for the tribas and hardlim functions were also excellent, yielding only some minor misclassifications (see the associated confusion matrices shown in Table 13). Table 11. Validation (train) classification performance measures of the ELM classifier for three different transfer functions of the ELM algorithm in dataset #2. Validation Performance Measures (Training Set) Accuracy Precision Specificity Recall F1 Score ELM tanh 1 1 1 1 1 ELM tribas 0.96970 0.93333 0.98148 0.96922 0.95094 ELM hardlim 1 1 1 1 1 Table 12. Verification (test) classification performance measures of the ELM classifier for three different transfer functions of the ELM algorithm in dataset #2. Verification Performance Measures (Test Set) Accuracy Precision Specificity Recall F1 Score ELM tanh 1 1 1 1 1 ELM tribas 1 1 1 1 1 ELM hardlim 0.98870 0.97222 0.99306 0.97436 0.97329 Table 13. Confusion matrices for the UMAP-ELM for dataset #2. Training set using the tribas activation function (left) and test set using the hardlim activation function (right). Predicted Class Predicted Class (Train Dataset & Tribas) (Test Dataset & Hardlim) HQ AQ LQ HQ AQ LQ Actual Class High-Quality Wine 32 0 0 11 0 0 Average-Quality Wine 137 0112 0 Low-Quality Wine 7099 0 0 35 The current combined UMAP-ELM approach improves the classification results presented in the original study [ 47 ]. This study also involves a feature extraction method along with the use of two different classifiers (an SVM and a deep MLP neural network) and presents the results of a 5-fold validation in the training and verification test stages. The reported average accuracies for the test stages were 97.34% for the conventional SVM approach, where 69 features were selected for the classification, and 97.68% for the improved approach, where 300 features were selected for the classification. In this case, the use of the combined UMAP-ELM methods allows perfect classifications to be obtained, therefore improving the aforementioned results. 8.3.3. Dataset #3 This dataset consists of 309 nearly balanced samples (92, 100, and 112 samples of carbon monoxide ( CO ), formaldehyde ( HCHO ), and nitrogen oxide ( NO2 ), respectively) split into 231 training samples (72, 75, and 84, respectively) and 78 test samples (24, 25, and 29, respectively). Each sample contained 10,000 initial features that were reduced to 8 using the supervised UMAP method. Figure 16 illustrates the first three features extracted using the UMAP method in dataset #3 for both the training and test samples. As can be seen in the figure, the supervised UMAP seems to properly cluster nearly all the samples of the training dataset however, when the test data were projected, a slight overlapping occurred. The small ambiguities present in the separation of the training dataset were amplified when projecting the test dataset. It is worth noting that, in this case, the characteristics of the features did not allow such an easy clustering as in dataset #2 and that the supervised UMAP method only had 231 samples to train however, in dataset #1, where the global structure was quite complex, 2700 samples were used for training.
Mathematics 2022,10, 29 32 of 38 x1 x2 x3 x4 x5 x6 x7 p12 p13 p14 p15 p16 p17 p24 p25 p27 p34 p37 p56 p67 y1 y2 y3 y4 y5 y6 y7 q12 q13 q14 q15 q16 q17 q24 q25 q27 q34 q37 q56 q67 Figure A3. Illustration of a weighted undirected graph for points in xi∈RD and a low-order representation of the graph where yi∈Rd . The weights pij(X) in the original space are computed from distances measured in the RD d-metric of the k -neighbors of each point, whereas the weights qij(Y) = (1+akyi−yjk2b 2)−1are computed from pairwise RdEuclidean distances. Note that the values of pij , which are computed from X , are fixed and that the weights qij are computed from the d features of the samples stored in Y . Therefore, after the constant terms in the objective function are removed, the problem at hand translates to finding Y∈Rn×dby minimizing: min ˜ Y∈Rn×d− n ∑ i,j=1pij log(qij(˜ Y)) + (1−pij)log(1−qij(˜ Y)). (A4) To use the stochastic gradient descent (SGD) method in the previous optimization problem, we consider a smooth function describing the similarities between the pairs of points in the low-dimensional space Rd: qij = (1+akyi−yjk2b 2)−1, where a and b can be user-defined positive parameters or automatically set by solving a nonlinear least squares fitting, once the minimum distance between the points in the low-dimensional representation (min_dist hyperparameter) and the effective scale of the embedded points (spread hyperparameter) is fixed. Specifically, a and b are found, if not previously given, fitting the real function (1+ax2b)−1to the exponential decay function: ψ(x) = 1, x<min_dist exp(−(x−min_dist)/spread), otherwise , for x∈[ 0,3 ·spread] . For instance, for the reference values min_dist = 0.1 and spread = 1, we have a=1.58 and b=0.90 (see [66] for details). Therefore, after an initial guess Yini is computed, an iterative procedure is applied to find the minimum value in Equation (A4) . The initial guess for the SGD method is the largest eigenvectors of the symmetric normalized graph Laplacian associated with P ( see [67] ). Specifically, we denote by D∈Rn×n the degree matrix of the undirected graph P ( D=diag(Dii) , where Dii =∑n j=1pij ), and by L=D1/2(D−P)D1/2 the normalized graph Laplacian; then, Yini is the matrix containing only the first d eigenvectors with the largest eigenvalues. Appendix B. Practical Computational Description of Extreme Learning Machine Let Y∈Rn×d be a matrix collecting the n samples given in Equation (A3) to be classified, where each specific sample is given by yi=rowi(Y)T= (yi1 , yi2 , . . . , yid)T∈Rd
Mathematics 2022,10, 29 33 of 38 and d indicates the relevant features selected by the UMAP algorithm. In addition, let {C1 , C2 , . . . , CL} be the classes/labels into which the samples must be classified, where L denotes the total number of classes. The ELM classifier takes as input a particular sample, yi∈Rd , and returns the output vector, oi= (oi1 , oi2 , . . . , oiL)T∈RL , containing the raw score of the class membership to each of the L classes. By default, the ELM method uses binarized {− 1,1 } class labels, and thus, if sample yi belongs to class Cci , we expect that (oi)j=oij ≈ − 1 + 2 δj,ci , i.e., we expect the ci th component of the vector to be close to 1 and all the other components to be close to − 1. The class prediction is then computed by selecting the maximum component of oi: cpi=argmax j=1,...,L oij ∈ {1, 2, . . . , L}. The output vector, oi , is computed from yi by introducing a hidden intermediate layer in the neural network represented by vector hi∈R˜ d , where ˜ d denotes the number of nodes in the hidden layer. In general, every layer of a neural network that takes as input a feature vector, xin ∈Rnin, and returns an output vector, xout ∈Rnout, is characterized by: (1) An activation function g:R→R; (2) An nout threshold or biases bj∈R,j=1, . . . , nout, collected in vector b∈Rnout ; and (3) nout weighting vectors wj∈Rnin , j= 1, . . . , nout , collected in the weighting matrix W= (w1,w2, . . . , wnout )T∈Rnout×nin . Thus, the output vector is computed as follows: xout =g◦(Wxin +b) = g(w1·xin +b) . . . g(wnout ·xin +b) , where g◦ is the element-wise activation function that applies the activation function, g , to each component of the vector. Based on this definition, it is easy to recover the usual alternative expression for the jth component of the output vector: (xout)j=g(wj·xin +b) = nin ∑ s=1 g(wjs(xout)s+bs). In particular, the ELM neural network consists of a hidden layer that converts the initial samples yi∈Rd into the hidden feature vector hi∈R˜ d and an output layer that converts the hidden vector hi∈R˜ d into the output vector oi∈RL (see Figure A4). The hidden layer consists of a weight matrix W= (w1 , w2 , . . . , w˜ d)T∈R˜ d×d , a bias vector b∈R˜ d , and a user-defined activation function g , whereas the output layer consists of a weight matrix B= (β1 , β2 , . . . , βL)T∈RLט d , a zero-bias vector, and an identity activation function. Therefore, after composing the two layers, we obtain an output vector of: oi=Bhi=Bg◦(Wyi+b). Once the neural network parameters (g , W , b , B) are set, it is straightforward to provide the class prediction for any given sample. As mentioned before, the training step of the ELM neural network is comparable to that of a traditional neural network, because the weights and biases of the hidden layer are randomly assigned at the beginning of the learning process and remain unchanged during the entire training process. In particular, once the hidden number of features ˜ d
Mathematics 2022,10, 29 34 of 38 is set, the hidden weights and biases are randomly computed using the standardized normal distribution: W=randN(0,1)(˜ d,d)∈R˜ d×d, b=randN(0,1)(˜ d,1)∈R˜ d. Y∈Rn×dH∈Rnט dO∈Rn×L yi∈Rdhi∈R˜ doi∈RL hij =g(wj·yi+bj)oik =βk·hi hi=g◦(Wyi+b)oi=Bhi yi1 yi2 . . . yid hi1 hi2 . . . hij . . . hi˜ d oi1 oi2 . . . oik . . . oiL w11 w21 wj1 w˜ d1 w12 w22 wj2 w˜ d2 w1d w2d wjd w˜ dd β11 β21 βk1 βL1 β12 β22 βk2 βL2 β1˜ d β2˜ d βk˜ d βL˜ d g w1,b1 w2,b2 wj,bj w˜ d,b˜ d β1 β2 βk βL W∈R˜ d×d b∈R˜ d,g B∈RLט d Figure A4. The ELM neural network consists of a hidden layer that converts the initial samples yi∈Rd into the hidden feature vector hi∈R˜ d and an output layer that converts the hidden vector hi∈R˜ dinto the output vector oi∈RL. Therefore, only the weights of the output layer B are computed using the training dataset. To illustrate the computation of B , for ease of presentation and without loss of generality, let us assume that all data Y∈Rn×d are used to train the neural network (in practice, only some of the total data are used for training the ELM classifier, and Y∈Rn×d should be replaced by Y∈Rntrain×d). Because the class of the training samples is known, the raw score of the class membership for each sample is stored in the true output matrix: T=(t1,t2, . . . , tn)T= tT 1 . . . tT n ∈Rn×L, (A5) where ti=rowi(T)T∈RL corresponds to the raw scores of the i th sample (containing − 1 in the non-actual class and 1 in the actual class). Specifically, if the i th sample belongs to
Mathematics 2022,10, 29 35 of 38 class Cci , then (ti)j= 1 if j=ci and − 1 otherwise. Moreover, because the hidden layer is already known and is described by (g , W , b) , we can compute all the hidden features of the input samples hi=g◦(Wyi+b)∈R˜ dand collect them in the matrix: H=(h1,h2, . . . , hn)T= hT 1 . . . hT n ∈Rnט d, (A6) where the vectors hi=rowi(H)T are now placed in the rows of matrix H . Then, the output of the ELM neural network is: O=(o1,o2, . . . , on)T= oT 1 . . . oT n =HBT∈Rn×L, (A7) where recall that oi=Bhi. The output weights are then computed by minimizing the distances between the true outputs T and the predictions O , namely, by finding a least-squares solution ˆ B of the linear system HBT=T: ˆ B=argmin B∈RLט d kHBT−TkF, (A8) where the Frobenius norm of a matrix is kAkF=ptrace(ATA) . Note that the system of equations, HBT=T , represents the n×L restrictions to be met, matrix BT has Lט d degrees of freedom, and, in general, the ELM can only approximate the training samples with zero error if the number of hidden nodes coincides with the total number of samples, ˜ d=n , in which case, BT=H−1T . In the general case in which ˜ d6=n (note that the number of hidden nodes is usually much smaller than the number of samples), the output weights are computed by solving the least-squares problem in Equation (A8), yielding: ˆ BT=H†T, (A9) where H† is the Moore-–Penrose generalized inverse of H (see [ 68 ]). To further describe this key computation of the training step, despite the Moore–Penrose pseudo-inverse usually being computed using a singular decomposition of the matrix, in the usual case in which ˜ d<n, we have the following: ˆ BT= (HTH)−1HTT. The training steps of the ELM are summarized in Algorithm A1. Algorithm A1. shows that the hidden layer parameters are randomly generated, independently of the dataset, and that the weights associated with the output layer are the only parameters that need to be trained. Equation (A9) reveals that the weights are computed explicitly without the use of an iterative procedure, and therefore, the ELM has a higher training speed than the BP learning algorithm.
Mathematics 2022,10, 29 36 of 38 Algorithm A1. ELM neural network training for classification. Input: A training dataset, Y∈Rn×d , of n samples containing d features; the class labels of the samples, {ci}i=1,...,n∈ { 1, . . . , L} ; the number of nodes in hidden layer, ˜ d ; the activation function, g, of the hidden layer. Output: ELM classifier 1: Create the true raw score class membership matrix, T=(t1,t2, . . . , tn)T∈Rn×L , using binarized {−1,1}class labels, where tij =1 if j=cior −1, otherwise. 2: Randomly generate weighting matrix W∈R˜ d×d and bias vector b∈R˜ d of the hidden layer, W=randN(0,1)(˜ d,d), where b=randN(0,1)(˜ d,1). 3: Compute the output matrix of the hidden layer, H∈Rnט d , associated with Y , namely, H=(h1,h2, . . . , hn)T, where hi=g◦(Wyi+b)∈R˜ d. 4: Compute the weights of the output layer, B= (H†T)T∈RLט d. References 1. Gardner, J.W.; Bartlett, P.N. A brief history of electronic noses. Sens. Actuators B Chem. 1994,18, 210–211. [CrossRef] 2. Karakaya, D.; Ulucan, O.; Turkan, M. Electronic nose and its applications: A survey. Int. J. Autom. Comput. 2020 ,17, 179–209. [CrossRef] 3. Marco, S.; Gutierrez-Galvez, A. Signal and data processing for machine olfaction and chemical sensing: A review. IEEE Sens. J. 2012,12, 3189–3214. [CrossRef] 4. Yan, J.; Guo, X.; Duan, S.; Jia, P.; Wang, L.; Peng, C.; Zhang, S. Electronic nose feature extraction methods: A review. Sensors 2015 , 15, 27804–27831. [CrossRef] [PubMed] 5. Lundström, I. Approaches and mechanisms to solid state based sensing. Sens. Actuators B Chem. 1996,35, 11–19. [CrossRef] 6. Gutierrez-Osuna, R.; Nagle, H.T.; Schiffman, S.S. Transient response analysis of an electronic nose using multi-exponential models. Sens. Actuators B Chem. 1999,61, 170–182. [CrossRef] 7. Gutierrez-Osuna, R.; Gutierrez-Galvez, A.; Powar, N. Transient response analysis for temperature-modulated chemoresistors. Sens. Actuators B Chem. 2003,93, 57–66. [CrossRef] 8. Varpula, A.; Novikov, S.; Haarahiltunen, A.; Kuivalainen, P. Transient characterization techniques for resistive metal-oxide gas sensors. Sens. Actuators B Chem. 2011,159, 12–26. [CrossRef] 9. Ducéré, J.M.; Hemeryck, A.; Estève, A.; Rouhani, M.D.; Landa, G.; Ménini, P.; Tropis, C.; Maisonnat, A.; Fau, P.; Chaudret, B. A computational chemist approach to gas sensors: Modeling the response of SnO2 to CO, O2, and H2O Gases. J. Comput. Chem. 2012,33, 247–258. [CrossRef] [PubMed] 10. Kamarudin, K.; Mamduh, S.M.; Md Shakaff, A.Y.; Mad Saad, S.; Zakaria, A.; Abdullah, A.H.; Kamarudin, L.M. Flexible and autonomous integrated system for characterizing metal oxide gas sensor response in dynamic environment. Instrum. Sci. Technol. 2015,43, 74–88. [CrossRef] 11. Siqueira, A.; Melo, M.; Giordani, D.; Galhardi, D.; Santos, B.; Batista, P.; Ferreira, A. Stochastic modeling of the transient regime of an electronic nose for waste cooking oil classification. J. Food Eng. 2018,221, 114–123. [CrossRef] 12. Siqueira, A.F.; Vidigal, I.G.; Melo, M.P.; Giordani, D.S.; Batista, P.S.; Ferreira, A.L. Assessing waste cooking oils for the production of quality biodiesel using an electronic nose and a stochastic model. Energy Fuels 2019,33, 3221–3226. [CrossRef] 13. Scott, S.M.; James, D.; Ali, Z. Data Analysis for Electronic Nose Systems; Springer: Berlin/Heidelberg, Germany, 2006. [CrossRef] 14. Zhang, L.; Tian, F.C. A new kernel discriminant analysis framework for electronic nose recognition. Anal. Chim. Acta 2014 , 816, 8–17. [CrossRef] [PubMed] 15. Leon-Medina, J.X.; Anaya, M.; Parés, N.; Tibaduiza, D.A.; Pozo, F. Structural damage classification in a Jacket-type wind-turbine foundation using principal component analysis and extreme gradient boosting. Sensors 2021,21, 2748. [CrossRef] [PubMed] 16. Leon-Medina, J.X.; Anaya, M.; Tibaduiza, D.A.; Pozo, F. Manifold Learning Algorithms Applied to Structural Damage Classification. J. Appl. Comput. Mech. 2021,7, 1158–1166. [CrossRef] 17. Leon-Medina, J.X.; Anaya, M.; Pozo, F.; Tibaduiza, D. Nonlinear Feature Extraction Through Manifold Learning in an Electronic Tongue Classification Task. Sensors 2020,20, 4834. [CrossRef] [PubMed] 18. Zhu, P.; Du, J.; Xu, B.; Lu, M. Modified unsupervised discriminant projection with an electronic nose for the rapid determination of Chinese mitten crab freshness. Anal. Methods 2017,9, 1806–1815. [CrossRef] 19. Ding, L.; Guo, Z.; Pan, S.; Zhu, P. Manifold learning for dimension reduction of electronic nose data. In Proceedings of the 2017 International Conference on Control, Automation and Information Sciences, ICCAIS 2017, Xi’an, China, 31 October–1 November 2017; pp. 169–174. [CrossRef] 20. Jia, P.; Huang, T.; Wang, L.; Duan, S.; Yan, J.; Wang, L. A novel pre-processing technique for original feature matrix of electronic nose based on supervised locality preserving projections. Sensors 2016,16, 1019. [CrossRef] [PubMed]
Mathematics 2022,10, 29 37 of 38 21. Leon-Medina, J.X.; Anaya, M.; Pozo, F.; Tibaduiza, D.A. Application of manifold learning algorithms to improve the classification performance of an electronic nose. In Proceedings of the 2020 IEEE International Instrumentation and Measurement Technology Conference (I2MTC), Dubrovnik, Croatia, 25–28 May 2020; pp. 1–6. [CrossRef] 22. Pardo, M.; Sberveglieri, G. Classification of electronic nose data with support vector machines. Sens. Actuators B Chem. 2005 , 107, 730–737. [CrossRef] 23. Tan, J.; Kerr, W.L. Determining degree of roasting in cocoa beans by artificial neural network (ANN)-based electronic nose system and gas chromatography/mass spectrometry (GC/MS). J. Sci. Food Agric. 2018,98, 3851–3859. [CrossRef] 24. Leon-Medina, J.X.; Cardenas-Flechas, L.J.; Tibaduiza, D.A. A data-driven methodology for the classification of different liquids in artificial taste recognition applications with a pulse voltammetric electronic tongue. Int. J. Distrib. Sens. Netw. 2019 ,15, 1550147719881601. [CrossRef] 25. Vergara, A.; Vembu, S.; Ayhan, T.; Ryan, M.A.; Homer, M.L.; Huerta, R. Chemical gas sensor drift compensation using classifier ensembles. Sens. Actuators B Chem. 2012,166, 320–329. [CrossRef] 26. Fonollosa, J.; Rodríguez-Luján, I.; Huerta, R. Chemical gas sensor array dataset. Data Brief 2015,3, 85–89. [CrossRef] 27. Vallejo, M.; de la Espriella, C.; Gómez-Santamaría, J.; Ramírez-Barrera, A.F.; Delgado-Trejos, E. Soft metrology based on machine learning: A review. Meas. Sci. Technol. 2019,31, 032001. [CrossRef] 28. Ting, W.; Guo-Zheng, Y.; Bang-Hua, Y.; Hong, S. EEG feature extraction based on wavelet packet decomposition for brain computer interface. Measurement 2008,41, 618–625. [CrossRef] 29. Zhang, L.; Liu, Y.; Deng, P. Odor recognition in multiple E-nose systems with cross-domain discriminative subspace learning. IEEE Trans. Instrum. Meas. 2017,66, 1679–1692. [CrossRef] 30. Zhang, L.; Zhang, D. Domain adaptation extreme learning machines for drift compensation in E-nose systems. IEEE Trans. Instrum. Meas. 2014,64, 1790–1801. [CrossRef] 31. Kumar, S.; Ghosh, A. A Feature Extraction Method Using Linear Model Identification of Voltammetric Electronic Tongue. IEEE Trans. Instrum. Meas. 2020,69, 9243–9250. [CrossRef] 32. Jing, Y.Q.; Meng, Q.H.; Qi, P.F.; Cao, M.L.; Zeng, M.; Ma, S.G. A bioinspired neural network for data processing in an electronic nose. IEEE Trans. Instrum. Meas. 2016,65, 2369–2380. [CrossRef] 33. Ozmen, A.; Dogan, E. Design of a portable E-nose instrument for gas classifications. IEEE Trans. Instrum. Meas. 2009 , 58, 3609–3618. [CrossRef] 34. Krutzler, C.; Unger, A.; Marhold, H.; Fricke, T.; Conrad, T.; Schütze, A. Influence of MOS gas-sensor production tolerances on pattern recognition techniques in electronic noses. IEEE Trans. Instrum. Meas. 2011,61, 276–283. [CrossRef] 35. Brudzewski, K.; Osowski, S.; Dwulit, A. Recognition of coffee using differential electronic nose. IEEE Trans. Instrum. Meas. 2012 , 61, 1803–1810. [CrossRef] 36. Tudu, B.; Metla, A.; Das, B.; Bhattacharyya, N.; Jana, A.; Ghosh, D.; Bandyopadhyay, R. Towards versatile electronic nose pattern classifier for black tea quality evaluation: An incremental fuzzy approach. IEEE Trans. Instrum. Meas. 2009 ,58, 3069–3078. [CrossRef] 37. Zhang, L.; Tian, F. Performance study of multilayer perceptrons in a low-cost electronic nose. IEEE Trans. Instrum. Meas. 2014 , 63, 1670–1679. [CrossRef] 38. Vanarse, A.; Osseiran, A.; Rassau, A. Neuromorphic engineering—A paradigm shift for future im technologies. IEEE Instrum. Meas. Mag. 2019,22, 4–9. [CrossRef] 39. Al Yamani, J.H.J.; Boussaid, F.; Bermak, A.; Martinez, D. Bio-inspired gas recognition based on the organization of the olfactory pathway. In Proceedings of the 2012 IEEE International Symposium on Circuits and Systems (ISCAS), Seoul, Korea, 20–23 May 2012; pp. 1391–1394. 40. Pan, X.; Zhang, H.; Ye, W.; Bermak, A.; Zhao, X. A fast and robust gas recognition algorithm based on hybrid convolutional and recurrent neural network. IEEE Access 2019,7, 100954–100963. [CrossRef] 41. Liu, H.; Yu, D.; Gu, Y. Classification and evaluation of quality grades of organic green teas using an electronic nose based on machine learning algorithms. IEEE Access 2019,7, 172965–172973. [CrossRef] 42. Liu, H.; Li, Q.; Gu, Y. A multi-task learning framework for gas detection and concentration estimation. Neurocomputing 2020 ,416, 28–37. [CrossRef] 43. Cheng, Y.; Wong, K.Y.; Hung, K.; Li, W.; Li, Z.; Zhang, J. Deep Nearest Class Mean Model for Incremental Odor Classification. IEEE Trans. Instrum. Meas. 2018,68, 952–962. [CrossRef] 44. Grover, A.; Lall, B. A Novel Method For Removing Baseline Drifts in Multivariate Chemical Sensor. IEEE Trans. Instrum. Meas. 2020,69, 7306–7316. [CrossRef] 45. Zhang, L.; Wang, X.; Huang, G.B.; Liu, T.; Tan, X. Taste recognition in e-tongue using local discriminant preservation projection. IEEE Trans. Cybern. 2018,49, 947–960. [CrossRef] [PubMed] 46. Leon-Medina, J.X.; Pineda-Muñoz, W.A.; Burgos, D.A.T. Joint Distribution Adaptation for Drift Correction in Electronic Nose Type Sensor Arrays. IEEE Access 2020,8, 134413–134421. [CrossRef] 47. Gamboa, J.C.R.; da Silva, A.J.; de Andrade Lima, L.L.; Ferreira, T.A. Wine quality rapid detection using a compact electronic nose system: Application focused on spoilage thresholds by acetic acid. LWT 2019,108, 377–384. [CrossRef] 48. Gamboa, J.C.R.; da Silva, A.J.; Ferreira, T.A. Electronic nose dataset for detection of wine spoilage thresholds. Data Brief 2019 , 25, 104202. [CrossRef] [PubMed]
Mathematics 2022,10, 29 38 of 38 49. Yin, X.; Zhang, L.; Tian, F.; Zhang, D. Temperature modulated gas sensing E-nose system for low-cost and fast detection. IEEE Sens. J. 2015,16, 464–474. [CrossRef] 50. Plastria, F.; De Bruyne, S.; Carrizosa, E. Dimensionality Reduction for Classification. In Proceedings of the 4th International Conference on Advanced Data Mining and Applications; Springer: Berlin/Heidelberg, Germany, 2008; pp. 411–418. [CrossRef] 51. Saul, L.; Weinberger, K.; Ham, J.; Sha, F. Spectral methods for dimensionality reduction. Semisupervised Learn. 2006 , 293–306. [CrossRef] 52. Schölkopf, B.; Smola, A.J.; Bach, F. Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond; MIT Press: Cambridge, MA, USA, 2002. 53. Scholkopf, B.; Mika, S.; Burges, C.J.; Knirsch, P.; Muller, K.R.; Ratsch, G.; Smola, A.J. Input space versus feature space in kernel-based methods. IEEE Trans. Neural Netw. 1999,10, 1000–1017. [CrossRef] [PubMed] 54. McInnes, L.; Healy, J.; Melville, J. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv 2018 , arXiv:1802.03426. 55. McInnes, L. UMAP for Supervised Dimension Reduction and Metric Learning. 2018. Available online: https://umap-learn. readthedocs.io/en/latest/supervised.html (accessed on 19 July 2021). 56. Sainburg, T.; McInnes, L.; Gentner, T.Q. Parametric UMAP: Learning embeddings with deep neural networks for representation and semi-supervised learning. arXiv 2020, arXiv:2009.12981. 57. McInnes, L. Transforming New Data with UMAP. 2018. Available online: https://umap-learn.readthedocs.io/en/latest/ transform.html (accessed on 10 October 2021). 58. Huang, G.B.; Zhu, Q.Y.; Siew, C.K. Extreme learning machine: Theory and applications. Neurocomputing 2006 ,70, 489–501. [CrossRef] 59. Xu, S.; Liu, S.; Feng, L. Fuzzy granularity neighborhood extreme clustering. Neurocomputing 2020,379, 236–249. [CrossRef] 60. Huang, G.B.; Zhou, H.; Ding, X.; Zhang, R. Extreme learning machine for regression and multiclass classification. IEEE Trans. Syst. Man Cybern. 2011,42, 513–529. [CrossRef] [PubMed] 61. Sokolova, M.; Lapalme, G. A systematic analysis of performance measures for classification tasks. Inf. Process. Manag. 2009 , 45, 427–437. [CrossRef] 62. McGinnis, W. Extreme Learning Machines, Sklearn-Extensions. 2015. Available online: http://wdm0006.github.io/sklearnextensions/extreme_learning_machines.html (accessed on 19 July 2021). 63. Jiménez, Á.B.; Lázaro, J.L.; Dorronsoro, J.R. Finding optimal model parameters by deterministic and annealed focused grid search. Neurocomputing 2009,72, 2824–2832. [CrossRef] 64. Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res. 2011,12, 2825–2830. 65. León, J.X.; Pineda Muñoz, W.A.; Anaya, M.; Vitola, J.; Tibaduiza, D.A. Structural Damage classification using machine learning algorithms and performance measures. In Proceedings of the 12th International Workshop On Structural Health MonitoringIWSHM 2019, Stanford, CA, USA, 10–12 September 2019. 66. Melville, J. Fine-Tuning UMAP Visualizations. 2020. Available online: https://jlmelville.github.io/uwot/abparams.html (accessed on 10 October 2021). 67. Shen, T. The Mathematics Behind Spectral Clustering And The Equivalence To PCA. arXiv 2021, arXiv:2103.00733. 68. Serre, D. Matrices Theory and Applications, 2nd ed.; Springer: New York, NY, USA, 2010.