scieee AI-readable full text Open interactive document viewer

Criteria for optimizing kernel methods in fault monitoring process: A survey

Llanes Santiago, Orestes,Bernal de Lázaro, José M.,Silva Neto, Antônio J.,Cruz Corona, Carlos Alberto

Abstract

Coordenação de Aperfeiçoamento de Pessoal de Nível Superior, Brazil (Finance Code 001 and CAPES-PRINT Process No. 88881.311758/2018-01)

Full text

JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 1 Criteria to optimizing kernel methods in fault monitoring process: A survey JM. Bernal-de L´ azaro, O. Llanes-Santiago, A. Silva-Neto, and C. Cruz-Corona Abstract—Over two last decade, diverse fault diagnosis strategies based on kernel methods have been studied intensively to ensure the safety and reliability of large-industrial processes. However, how choice a proper kernel function and its parameters are still two open research issues. This paper provides an overview of kernel-based preprocessing methods in fault monitoring tasks, with emphasis on helpful procedures to select the optimal kernel parameters in fault diagnosis applications. The discussion is also towards evaluating some kernel functions available from the machine learning literature but less used in fault diagnosis applications. With this aim, a KPCA-based monitoring scheme is employing to validate the effectiveness of sixteen kernel functions in two case studies. Index Terms—Kernel methods, kernel function, fault monitoring, preprocessing methods, tuned parameters, kernel PCA F 1 INTRODUCTION KErnel learning has become an important research topic for the data-based monitoring field because it provides potential solutions to nonlinear problems, including feature extraction, classification, and clustering tasks. Accordingly, numerous kernel approaches have been addressed to deal with the increasing complexity of modern industrial plants [1–5]. For example, the Support Vector Machines (SVM) paradigms have widely used to build effective discriminant models that are applied in fault isolation tasks [5–8]. Kernelized regression procedures have been also addressed to enhance the fault identification and the failure time prediction in complex production processes [9–12]. On the other hand, adaptive fault diagnosis schemes using kernel clustering strategies have been proposed for handling imbalanced classes and measurements affected by noise and outliers [13–16]. Many other kernelized algorithms have also been successfully employed as data pre-treatments methods in fault diagnosis applications, including algorithms as the Principal Component Analysis (KPCA) [17–19], Partial Least Squares (KPLS) [20–23], Fisher Discriminate Analysis (KFDA) [24–26], Entropy Component Analysis (KECA) [27–32], Independent Component Analysis (KICA) [33–35], as well as, Canonical Component Analysis (KCCA) [36–38]. Among all these data-pretreatments methods, kernel PCA is the more popular and extensively studied for fault diagnosis applications. The main advantage of the above kernel-based solutions stems from the inherent mapping of input data into a higher-dimensional space through the kernel trick, which allows us to handle more effectively their nonlinear •A. Silva-Neto is with Department of Mechanical Engineering, Universidade do Estado do Rio de Janeiro, IPRJ-UERJ, RJ, Brazil. •C. Cruz-Corona is with Departament of Computer Science and Artificial Intelligence, University of Granada, Spain. •JM. Bernal-de L´azaro and O. Llanes-Santiago are with Department of Computer Systems and Automation Engineering, Universidad Tecnol´ogica de La Habana ”Jos´e Antonio Echeverr´ıa”, CUJAE, Cuba. E-mail: [email protected] Manuscript received August xx, 2020; revised August xx, 2020. relationships. However, two bottlenecks are limiting the application of these approaches: (a) the computation complexity usually increases with sample size and dimension size [4]; (b) the proper choice of the kernel function and its parameters must be ensured to achieve high performances. Most authors agree that the selection and tuning of kernel functions are two open research issues [22, 39, 40]. However, the number of published papers that explicitly address the adjusting of the kernel-based diagnosis schemes is still very limited. Likewise, a systematic review of kernel parameter optimization criteria for the data-based process monitoring has not been recently reported. Since the lack of rigorous theoretical framework and appropriate guidelines, many times the kernel parameters are arbitrary selected without there being any guarantee of achieving acceptable performances [39, 41, 42]. These type of intuitive solutions was referred to here as heuristic criteria. Other strategies for the kernel parameter selection will be classified into two categories: (a) direct evaluation criteria that are more closely related to the kernel matrix; (b) indirect evaluation criteria where the fault diagnosis performance is evaluated to tuning the kernel parameters. These categories differ from those proposed by Xiao et al. [43], and they are not necessarily focused on the type of searching algorithm that is utilized, but rather the objective function employed to find the optimal parameters of the candidate kernels. The first aim of this paper is to provide readers an overview of the kernel-based preprocessing methods in fault monitoring tasks, with emphasis on helpful procedures to achieve the proper tuning of the kernel functions selected. From the above view of point, the present work provides additional remarks that complement the comprehensive literature review on fault diagnosis applications of data-driven methods that have been supplied by several papers recently published [3–6, 35, 44–49], wherein these topics are not particularly detailed. The second motivation of this paper is to identify kernel functions potentially useful, that are available from JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 2 the machine learning literature [50–54] but doesn’t be usually evaluated in fault diagnosis applications. The main contribution of this paper stems from the proposal, analysis, and justification of several kernel optimization criteria that facilitates finding the optimal kernel parameters in fault diagnosis applications. In order to further compare their performance from a practical perspective, sixteen kernel functions are evaluated using two case studies. Even though the effectiveness of the proposal is validated through a KPCA-based monitoring scheme, all optimization criteria herein discussed can be easily extensible to other kernelized algorithms. This paper is organized as follows. Section 2 is addressed to review the theoretical concepts and related works. Section 3 presents the kernel optimization criteria and the framework proposed for their application. Section 4 validates the procedures for the adjusting of the kernel functions through two benchmark process. Finally, the conclusions are given. 2 RELATED WORKS This section outlines and provides further references for kernel methods applied in fault diagnosis tasks. 2.1 Methodology The literature reviewed in the paper preparation includes publications since 2004 that have been obtained from ScienceDirect, IOPscience, IEEE Xplore, Thomson Reuters, and Taylor &Francis citation databases. Aiming provided an overview of the kernel-based monitoring trends, some papers published in world-renowned congresses were also analyzed. Through the information associated with titles, abstracts, and keywords, were identify hundreds of works related to kernel-based monitoring applications. The reviewed articles were also categorized into two relevant areas of research: electro-mechanical systems and complex chemical processes. Finally, the works more closely to the following scopes were considered. −Research papers developing comparatives studies about kernel methods in fault monitoring issues. −Research papers using several kernelized algorithms to improve the fault diagnosis procedures. −Research papers evaluating several kernel functions. −Research papers considering the choice of kernel parameters part of the fault diagnosis procedure. It is important to highlight that a detailed mathematical description for all reviewed kernel methods is beyond that the scope of this paper. For interested readers, it is suggested consider to original works. 2.2 Data-pretreatment through kernel methods In the data-driven process monitoring the measurements routinely collecting from the industrial processes are represented in a finite-dimensional space through vectors (i.e., so-called observations). A kernel function represents the inner product of these observations, which are implicitly mapped in a high dimensional space to handle more effectively the non-linear relations between the variables of the process. Let {Xi}m i=1 ∈Rpdenote the historical records with mobservations, a kernel function can be defined by: k(xi,xj) = hΦ(xi),Φ(xj)i(1) where Φindicates the implicit mapping of input data into the Reproducing Kernel Hilbert Space (RKHS). If a symmetric positive definite function that satisfies Mercer’s conditions is employed, then the kernel represents the inner product in the feature space F, which amounts to the angle between the observations. The kernel methods assume that similarities between input data can be completely represented by the kernel matrix Kij;∀i, j ={1, . . . , m}, resulting to evaluate the kernel functions. Kij =   ΦT 1Φ1· · · ΦT 1Φm . . ..... . . ΦT mΦ1· · · ΦT mΦm   =   k(x1, x1)· · · k(x1, xm) . . ..... . . k(xm, x1)· · · k(xm, xm)   (2) Therefore, the spatial structures of the input data and their similarity values are both determined by the tuning parameters of the kernel function. In order to validate this idea from the fault diagnosis perspective, the Table 1 illustrates different kernel functions that were considered for the present study. According to Genton (2001) [50], these kernels may be classified as anisotropic stationary, isotropic stationary, compactly supported, locally stationary, nonstationary, and separable nonstationary functions. For Interested practitioners, it is also suggested review the works provides by Canu and Smola (2006) [55], Al-Daoud and Turabieh (2013) [52], Zhu et al. [53], and Tian and Wang (2017) [54], where more kernel functions are analyzed for applications that go beyond the fault diagnosis field. From the literature review, it is remarkable the number of published papers wherein the kernel-based preprocessing methods are employed to provide enhanced diagnosis schemes. The data-pretreatment through kernel methods allows to reduce the dimension of the original data and the overload of computational operations but also increases the robustness against uncertainties by preserving the relevant information of observations [56]. Hence, the kernelized algorithms require that the embedding data sets to be transformed into a reduced subspace (dp) utilizing a transformation matrix P∈Rm×d, such that the dot products of images can be substituted by a kernel [57]. Then, considering that the projection vector is a linear combination of images in F, the matrix Pcan be expressed as Pj= m X i=1 βj,iΦ(xi)=Φβj(3) where Pjis the jth-transformation vector of P= [P1,...,Pd], and βj= [βj,1, . . . , βj,m]Tis the jth-vector from the pseudo transformation matrix B= [β1, . . . , βd]∈Rm×d. Through the kernel trick, the Eq. (3) can be modify as BTk(·, x) = BTKα(4) where k(·, x) = [k(x1, x), . . . , k(xm, x)]T= ΦTΦ(x), and K= ΦTΦ∈Rm×mis the kernel Gram matrix which is symmetric and positive semidefinite. According to Zhang et al. [57], the problem of the dimensionality reduction in the JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 3 TABLE 1 Kernel functions. No. Kernel Comments 1k1(xi, xj) = exp −kxi−xjk2/2θ2Gaussian kernel 2k2(xi, xj)=1/1 + kxi−xjk2/θCauchy kernel 3k3(xi, xj)=1− kxi−xjk2/kxi−xjk2+θRational Quadratic kernel 4k4(xi, xj) = 1 + kxi−xjkθ−1Generalized T-Student Kernel 5k5(xi, xj) = exp (−kxi−xjk/θ)Laplacian kernel 6k6(xi, xj) = exp −kxi−xjk/2θ2Exponential kernel 7k7(xi, xj) = hxT ixji+cθPolynomial kernel 8k8(xi, xj) = pkxi−xjk2+θ2Multiquadratic kernel 9k9(xi, xj)=1/pkxi−xjk2+θ2Inverse Multiquadratic kernel 10 k10(xi, xj) = −kxi−xjkθPower kernel 11 k11(xi, xj) = −log 1 + kxi−xjkθLog kernel 12 k12(xi, xj)=1−sin (πkxi−xjk/2θ)Nonparametric kernel (NNk13) 13 k13(xi, xj) = ωαhxT ixji+cθ+ (1 −ω)exp −kxi−xjk2/2θ2RBFpoly kernel 14 k14(xi, xj)=1− 3 X i=1 (−kxi−xjk/θ)i/2i−1Nonparametric kernel (NNk10) 15 k15(xi, xj)=1−"1 + 3 X i=1 (−1)i/1+(kxi−xjk/θ)i#Nonparametric kernel (NNk12) 16 k16(xi, xj)=1− 3 X i=1 (−kxi−xjk/θ)i/i−1Nonparametric kernel (NNk11) More information of nonparametric kernels in Al-Daoud and Turabieh (2013) [52] kernel feature space can be therefore formulating through the pseudo-transformation matrix in RKHS. In general, by applying eigenvalue decomposition to the kernel matrix, the pseudo-transformation vectors in the higher-dimensional space may be obtained. In this way, new descriptors through different data projections can found. For example, KPCA is a dimensionality reduction technique wherein the projections preserve the meaningful variability information in the original data set. The KICA algorithm finds a non-linear representation of the original set of variables for which the statistical dependence of the components result minimized [58]. In this context, the KECA reveals the angular structure of the transformed data set using projections onto those KPCA axes that contribute to the entropy estimate but not the variances of the input space data set [28]. In KPLS, the dimensions of two matrix blocks (i.e., input variables X and response variables Y) result simultaneously reducing to find the latent variables through projections that maximizing the covariance between input and output scores. A similar concept is addressed by the KCCA algorithm that maximizing the multi-dimensional correlation between X and Y, which performs better in terms of prediction power. Likewise, KCVA finds relations between two sets of variables considering serial correlations between the time instant to ppast and ffuture measurements. On the other hand, KFDA is a supervised data-pretreatment algorithm that finds the transformation vectors maximizing the separability between the different sets of variables. 2.3 Application Examples Next, several applications of kernel methods for feature extraction in electro-mechanical systems and complex chemical processes are analyzed. To facilitate the analyses on the subsequent sections, those studies where the kernel parameter selection forms part of the fault diagnosis procedure were summarized in Table 2. 2.3.1 Applications in electro-mechanical systems. The rotating machinery is essential to transmit power and motion in the modern manufacturing industries [89]. There are several kernel approaches reported in the literature for the health condition monitoring of such devices. For example, Zhang et al. [90] propose a KPCA-SVM scheme to predict the bearing degradation process. Dong et al. [91] applied the RBF-KPCA to reduce the feature mutual correlation, while a Morlet-SVM classifier [92] is used to identify the bearing running state. On the other hand, Shao et al. [93] utilize the RBF-KPCA for the fault feature extraction on gear systems. Otherwise, Liu et al. [94] integrate KPCA and TWSVM to analyze the health conditions of bearing and bevel gear faults. Likewise, Cheng et al. [69] investigate the joint working of KPCA and Learning Vector Quantization (LVQ) Neural Network for the planetary gear diagnosis. At around the same time, Wan et al. [95] analyzes the performance of the PCA, KPCA, and Isomap algorithms considering five different gear crack. Sakthivel et al. [96] discusses the benefits of JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 4 TABLE 2 Summary of papers: Kernel method, case studies, kernel functions, and fitness criterion used. Kernelized Benchmark Kernel Fitness number of No. Reference. Year Methods evaluated function criteria Scientific Journal citations 1 Lee et al.[59] 2004 PCA NE, WWTP RBF Heuristic Chemical Engineering Science 915 2 Choi et al.[60] 2005 PCA NE, CSTR RBF Heuristic Chemom. Intell. Lab. Syst. 419 3 Cho et al.[61] 2005 PCA NE, CSTR RBF Heuristic Chemical Engineering Science 294 4 Yoo et al.[62] 2005 PCA NE, WWTP RBF Heuristic Process Biochemistry 39 5 Jia et al.[63] 2010 PCA NE, PenSim RBF Heuristic Chemom. Intell. Lab. Syst. 101 6 Khediri et al.[64] 2011 PCA NE, TEP RBF Heuristic Computers and Industrial Engineering 63 7 Ni et al.[65] 2011 PCA PS RBF Direct IEEE Transactions on Power Delivery 95 8 Deng et al.[66] 2013 PCA NE, TEP RBF Heuristic Neurocomputing 57 9 Wang et al.[67] 2015 PCA SEP RBF Heuristic Chemom. Intell. Lab. Syst. 24 10 Yao et al.[68] 2015 PCA PenSim RBF Direct Journal of Process Control 33 11 Cheng et al.[69] 2016 PCA RM RBF Direct Measurement 57 12 Bernal et al.[34] 2016 PCA TEP RBF Indirect Chemical Engineering Science 30 13 Zhang et al.[70] 2017 PCA NE, PenSim RBF Indirect IEEE Access 23 14 Fu et al.[23] 2017 PCA DP RBF Indirect Chemom. Intell. Lab. Syst. 8 15 Deng et al.[71] 2017 PCA NE, TEP RBF Heuristic ISA Transactions 17 16 Mansouri et al.[17] 2018 PCA BP RBF Heuristic IEEE Transactions on Nanobioscience 5 17 He et al.[72] 2018 PCA RM RBF Direct Journal of Vibroengineering 1 18 Harkat et al.[18] 2019 PCA NE, TEP RBF Heuristic Chemical Engineering Science 9 19 Qian et al.[73] 2020 PCA NE, TEP RBF Heuristic Journal of Automatica Sinica – 20 Zhou et al.[74] 2020 PCA TEP RBF Heuristic Neurocomputing 1 21 Hamrouni et al.[19] 2020 PCA TEP RBF Direct Int. J. Adv. Manuf. Technol. – 22 Zhang et al.[75] 2007 ICA NPP RBF Heuristic Ind. Eng. Chem. Res 111 23 Fan et al. [33] 2014 ICA TEP RBF Heuristic Information Sciences 103 24 Bernal et al.[34] 2016 ICA TEP RBF Indirect Chemical Engineering Science 30 25 Zhang et al.[70] 2017 ICA NE, PenSim RBF Indirect IEEE Access 23 26 Samuel et al.[76] 2015 CVA TEP RBF Heuristic IFAC-PapersOnLine 18 27 Bai et al.[31] 2020 ECA RM RBF Indirect Applied Soft Computing – 28 He et al. [77] 2008 FDA TEP RBF Direct Chemom. Intell. Lab. Syst. 34 29 Zhu et al. [78] 2010 FDA TEP RBF Heuristic Chem Eng. Res. Des. 68 30 Zhu et al.[79] 2011 FDA TEP RBF Heuristic Expert Systems with Applications 82 31 Rong et al. [80] 2013 FDA WWTP, TEP RBF Indirect Computers and Chemical Eng. 12 32 Liu et al. [81] 2013 FDA RM RBF Direct Int. J. Adv. Manuf. Technol 52 33 Bernal et al. [82] 2015 FDA TEP RBF Direct Computers and Industrial Eng. 41 34 Feng et al. [24] 2016 FDA TEP RBF Heuristic IEEE Trans. Autom. Sci. Eng. 31 35 Deng et al. [26] 2016 FDA NE, CSTR RBF Heuristic Chemom. Intell. Lab. Syst. 30 36 Peng et al. [83] 2013 PLS HSMP RBF Indirect Control Engineering Practice 74 37 Zhang et al. [84] 2013 PLS NE, TEP, HSMP RBF Heuristic Math. Probl. Eng. 50 38 Godoy et al. [40] 2014 PLS NE RBF Heuristic Chemom. Intell. Lab. Syst. 25 39 Sheng et al.[20] 2015 PLS TEP RBF Heuristic IEEE Trans. Autom. Sci. Eng. 22 40 Jiao et al. [22] 2017 PLS NE, TEP RBF Indirect ISA Transactions 31 41 Fu et al. [23] 2017 PLS DP RBF Indirect Chemom. Intell. Lab. Syst. 8 42 Shi et al.[85] 2009 PCA RM RBFpoly Direct IFAC Proceedings Volumes 4 43 Jia et al.[86] 2012 PCA NE, PenSim RBF,Poly, Indirect Computers & Chemical Engineering 48 Sigmoid 44 Van et al. [87] 2015 FDA RM RBF, Morlet Indirect IEEE Trans Instrum Meas. 37 45 Shi et al. [88] 2016 FDA TEP RBFpoly Indirect Int. J. Syst. Sci. 11 46 Pilario et al.[36] 2019 CVA NE, CSTR RBFpoly Indirect Computers & Chemical Engineering 11 NPP, Nosiheptide production process; CSTR, Continuous stirred-tank reactor; WWTP, Wastewater treatment plant; PenSim, Penicillin fermentation process; NE, Numerical example; TEP, Tennessee Eastman process; HSMP, Hot strip mill process; RM, Rotating machinery; DP, Distillation process; SEP, Industrial semiconductor etch process; PS, Power systems the KPCA to reveal the health conditions of several pump components, including the bearing and the impeller of the centrifugal pump. Besides, Wong et al. [97] evaluate the KPCA-SVM (i.e., linear, RBF, and polynomial kernel) to detect gearbox faults for gas turbine generator systems. From a different perspective, Jiang et al. [98] studies the bearing degradation process using KICA and several kernel functions (i.e., Hermite, RBF, and polynomial kernel). The advantages of KICA and LS-SVM are further combining by Ma et al. [99] to recognize several faults in bearing elements. To reveal the operation state of the gearbox, Li et al. [100] utilized a Fuzzy k-Nearest Neighbor (FKNN) having as inputs the fault features extracted by a KICA processing stage. After this, Li et al. [101] suggest the faults detection of gears though KICA and BP Neural Network. Further, Liu et al. [81] utilize a hybrid kernel dimension reduction through RBF-KFDA for the fault level diagnosis of planetary gearboxes. Meanwhile, Jiang et al. [102] propose the Semi-supervised Kernel Marginal Fisher Analysis (SSKMFA) to extracts the low-dimensional characteristics from the raw high dimensional vibration signals in the gearbox. At the same work, the classifiers k-NN and RBF-SVM are also utilizing to evaluate the performances of five feature extraction methods (i.e., JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 5 SAEMD, KPCA, KFDA, KMFA, and SSKMFA). Somewhat later, He et al. [72] investigates the feature extraction and fault gearbox recognition when the KPCA is regularizing through the WPSO-FDA strategy. More recently, Bai et al. [31] employed a KECA-SVM scheme to recognize the state information of rolling bearing and gear, while a whale optimization algorithm is using to find the kernel parameters. 2.3.2 Applications in complex chemical processes The kernel methods have been widely investigated for fault diagnosis in complex chemical plants, which including continuous and batch processes. For example, dynamic extensions of kernel PCA [103–105], kernel PLS [106–108], and kernel ICA [33, 109] have been proposed to minimize the computation time and enhance fault detection performances. Kernelized algorithms have also exhibited a tremendous potential to handle the high nonlinearity [20, 22, 23, 40, 59, 60, 62, 82], the non-gaussianity [34, 75, 110–112], the multiple operation modes [113–115], and the inherently time-varying dynamics [70, 116, 117] that are typical on the chemical processes. Likewise, the kernel learning approaches have proven to be useful to detect incipient faults [34, 36, 76, 118–121], even at dynamically varying process operating conditions. Nevertheless, the related literature suggests integrating the kernel solutions to ensure more efficient operations on the critical systems. In this context, Vitale et al. [122] deal with the non-linear data structures of the chemical plants by using a kernel PLS with a pseudo-sample projection strategy. Meanwhile to aid the monitoring process, Zhang et al. [123] develop the kernel concurrent projection to latent structure (KCPLS) method aiming to reveal the more fault-relevant directions. On the other hand, He et al [77] and Li et al [124] propose two modified versions of KFDA to extracting discriminative features in the fault monitoring tasks. For more reliable quality monitoring, Jia et al. [63] integrate the KPCA and ARMAX models to characterize the batch processes. Somewhat later, Jiao et al. [22] discuss the nonlinear quality-related faults detection when KPLS is used to build the linear relationship between kernel and output matrices. Moreover, Zhu et al. [79] evaluate two discriminant approaches using KFDA-GMM and KFDA-kNN to perform the fault classification process. Beyond that, Khediri et al. [64] introduce an adaptive KPCA scheme to facilitate the recursive calculation of chart control limits between different batches. Soon later, Deng et al. [66] propose a novel similarity factor using KPCA to reveal the statistics information hidden in original measured variables. Alternatively, Yao et al. [68] explores the concept of generalized additive KPCA for the batch monitoring processes. On the other hand, Wang et al. [67] utilize the functional KPCA to handle nonlinear correlations between monitoring variables and/or sampling times. From a data-driven perspective, Feng et al. [24] take advantage of KFDA with local and global manifold to removes outliers caused by disturbances. With the above idea in mind, Hu et al. [125] suggests the spherical KPLS in order to handle data sets contaminated with outliers, while Deng et al. [71] uses KPCA with double-weighted local outlier factor (LOF). Furthermore, Zhu et al. [78] addressed the problem of imbalanced data through a KFDA-based fault diagnosis scheme. Besides, Mansouri et al. [17] propose a multiscale kernel generalized likelihood ratio test (MS-KGLRT) to enhance the fault monitoring abilities of nonlinear biological plants. Given the above analysis, Harkat et al. [18] and Hamrouni et al. [19] investigate the KPCA approach using a nonlinear interval of data values to address the problem of uncertainties in these systems. To enhance the fault detection performance, Deng et al. [26] considers a Bayesian perspective to built the fault monitoring statistics from the kernel feature space. More recently, Zhou et al. [74] make full use of low-order and high-order statistics of the process to preserves both local structure and global structure in the kernel feature space improving the fault monitoring of chemical processes. 3 CRITERIA TO CHOICE KERNEL PARAMETERS As already mentioned, the present work grouped the procedures for tuning the kernel parameters in three categories. In the first group are the heuristical selection procedures. A criterion widely used in this context is θ=cmσ2, where mand σ2are related to the training data set, and cis an empirical constant. Table 2 shows several published papers wherein different variants of the above rule utilized. In second place were considered the direct evaluation criteria, which take advantage of the data spatial structure to find the optimal kernel parameters. As illustrated in Figure 1, the direct evaluation criteria are independent of the fault diagnosis scheme used. Therefore, they have low computational complexity. Besides, from a theoretical view of point, the kernel functions tuned by the direct evaluation criteria could be successfully reusable for other fault monitoring schemes without repeat the parameter optimization procedure. Discriminate diagnosis algorithms Kernel methods Evaluations of the fault diagnosis tasks Kernel methods Data Indirect kernel parameter optimization Direct kernel parameter optimization 1 2 Fig. 1. Kernel evaluation approaches using for the selection of the optimal parameters. (1) A direct conception to built the objective function, (2) An indirect conception to built the objective function. In contrast to this, the indirect evaluation criteria choose the kernel parameters considering the performance of fault diagnosis tasks. In general, higher fault diagnosis performances are achieving through this approach, but with more effort and computational costly during the training phase of algorithms. Furthermore, both direct and indirect evaluation approaches can be indistinctly employed with either optimization algorithms available in the literature. JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 6 3.1 Direct evaluation criteria The direct kernel optimization approaches implement the diagnosis tasks as supervised learning problems, which depend on the learned spatial structure by the kernel gram matrix computed over the training data. Since the observations of each class can consider as one cluster, a kernel induced transformation will be appropriate if the separability of the data-classes in the inducing feature space is maximum. From this view of point, different inter-cluster distances have been using to choose the kernel parameters, but mostly restricted to gaussian kernels. For example, Teixeira et al. [126] choose the tuned parameter for the RBF kernel by the maximum distance of the samples from the mean vector of the class. Meanwhile, Ni et al. [65] consider the sum of the difference between the maximum and minimum for each variable in the data matrix. In the work reported by Zhang et al. [127], the median value of the reciprocal distances between each sample and the sample mean is utilizing as parameter selection criteria. More recently, Cheng et al. [69] employe the total distance among the sample centers of different fault gears in the feature space. On the other hand, Hamrouni et al. [19] suggest the average minimum distance between two points in the training data. According to Wu et al. [128], the following inter-cluster distance measures can be also used to find the kernel parameters. 1The minimum inter-cluster distance: min Jf(˜ θf) = kφ(x+)−φ(x−)k(5) 2The maximum inter-cluster distance: max Jf(˜ θf) = kφ(x+)−φ(x−)k(6) 3The average inter-cluster distance: max Jf(˜ θf) = 1 N+N−X x+∈X+X x−∈X− kφ(x+)−φ(x−)k(7) 4The distance between means of the two clusters: max Jf(˜ θf) = kµ+−µ−k(8) 5The distance between the points of one cluster from the mean of the other cluster: max Jf(˜ θf) = X x+∈X+ kφ(x+)−µ−k+X x−∈X− kφ(x−)−µ+k N++N− (9) where N+and N−represents the number of data points in the positive class X+and the negative class X+respectively. The values x+∈X+and x−∈X−are the observations with means µ+and µ−in the kernel induced feature space. Perhaps the inter-cluster distance more widely recognized in fault diagnosis applications [72, 82, 85, 129] is the Fisher measure based on the Rayleigh quotient [130]. Through the implicit information on the kernel matrix, the Fisher measure finds the projection ωthat maximizes the separability of different classes in the empirical feature space, while it minimizes their spatial dispersion. min Jf(˜ θf) = ωTSbω ωTSwω(10) Sb= (µ+−µ−)(µ+−µ−)T(11) Sw=X x+∈X+ (x+−µ+)(x+−µ+)T+X x−∈X− (x−−µ−)(x−−µ−)T(12) where Sband Swrepresent the between-class and total within-class scatter matrices, respectively. This measure is optimal only for homoscedastic data distributions because all classes are assumed from a normal distribution and identical covariance matrices [131]. Among the direct evaluation criteria, there are also available kernel matrix-based approaches, where the class label information is employed to indicate how an ideal kernel matrix should be for discriminant tasks [77, 82]. For example, Kernel target alignment (KTA) is a measure commonly used to compare two positive semidefinite kernel matrices [132, 133]. The KTA is defining as the normalized Frobenius inner product between a candidate kernel matrix and their ideal kernel. It is interpreted as the cosine distance between two vectors [134], taking values in [−1; 1]. Though high KTA values amount to the kernel being well aligned to the class distribution of the ideal kernel, the opposite is not always true. For two kernel matrix, the objective function can be formulated as follow. max Jf(˜ θf) = hK1, KidealiF qhK1, K1iFhKideal, KidealiF (13) Based on KTA, other kernel evaluation measures have also been derived recently, such as kernel polarization [135, 136], and feature space-based kernel matrix evaluation measure (FSM) [134]. Likewise, there are other measures less known as the α,β, and γmetrics proposed by [137]. In specific, α−measure employ the Pearson correlation coefficient obtained for the triangular kernel matrix candidate and the ideal kernel matrix, both previously normalized. As a result, it is less computationally intensive than KTA. The objective function based on the α−measure can be computed as follow. max Jf(˜ θf) = m X i=1 m X j=i+1 (Kij −K)(Sij −S) v u u t m X i=1 m X j=i+1 (Kij −K)2v u u t m X i=1 m X j=i+1 (Sij −S)2 (14) Sij =1,if yi=yj 0,if yi6=yj(15) where the mean values of Kand Sare calculated over a triangular matrix without main diagonals. K=2 m(m−1) m X i=1 m X j=i+1 Kij (16) S=2 m(m−1) m X i=1 m X j=i+1 Sij (17) JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 7 On the other hand, β−measure based on the t−student statistical test provides an evaluation for the equality of two mean vectors to determine the difference between the elements of the partitioned kernel matrix [82, 137]. For different operating conditions the β−measure may be used to determine the kernel parameter maximizing the separability between classes. max Jf(˜ θf) =  ¯ kw−¯ kb qσ2 w mw+σ2 b mb  (18) kb={Kij|i<j∧yi6=yj}(19) kw={Kij|i<j∧yi=yj}(20) where ¯ kxand σ2 xare the mean and variance of respective vectors kwand kb, with mxobservations. Beyond that, γ−measure use the Brown−Forsythe test to evaluate the equality between the variance in the kernel block matrix, but without the normality assumption. Results shown in [137] suggest that the γ−measure could have good results independently of whether classes are balanced or not. The objective function based on this measure can be computed as follows. min Jf(˜ θf)=(m−2) mw(¯ Zw−¯ Zwb)2+mb(¯ Zb−¯ Zwb)2 Pmw i=1(¯ Zi−¯ Zw)2+Pmb i=1(¯ Zi−¯ Zb)2(21) then considering the Eqs (19-20), and kwb ={kw∪kb}with mwb =mw+mbas the number values in kwb, it can be obtained that for each of them Zx=kx −¯ kx. In this case, the mean values ¯ Zw,¯ Zband ¯ Zwb are calculated using the following transformed variables. ¯ Zx=1 mx mx X i=1 Zix (22) 3.2 Indirect evaluation criteria These strategies choose the kernel parameters that ensure a high monitoring performance according to some objective function. For example, Jia et al. [86] relates several monitoring indicators using three kernel functions (i.e., RBF, Poly, and Sigmoid) and the following fitness function. min Jf(˜ θ) = η(ηA+SPElim ×dηA)×dη(23) where ηis the correct monitoring rate with resolution dη, and SPElim is statistical control limit of SPE. Meanwhile, ηAis referred to as the inverse of the retained component numbers of KPCA with a resolution dηA. In this case, the correct monitoring rate is calculated as: η=ηr,n ηt,f +ηr,f ηt,f (24) where ηt,n is the normal total number of samples by testing; ηr,n is the normal total number of samples by correct testing; ηt,f is the fault total number of samples by testing; ηn,f is the fault total number of samples by correct testing. A similar approach is addressed in Shi et al. [88], which considers the kernel parameter selection by minimizing the ratio between the number of erroneous diagnoses, and the number of total training data. On the other hand, Zhang et al. [138] consider a combination of the false alarm rate dr and detection rate fr using the following fitness function. max Jf(˜ θ) = 2 X i=1 νi(dri−fri)(25) where νiis the weighting factor for the two monitoring indices T2and SPE, respectively. Another fitness function is reported by Bai et al. [31] for the feature selection in rotating machinery using the KECA approach. max Jf(˜ θ) = E(1)/ N X i=1 E(i)(26) where, Edenotes the kernel entropy score of each kernel entropy component. Alternatively, Rong et al. [80] attempt to choose the kernel parameters achieving the lowest misclassification rate by three-fold cross-validation, such as: min Jf(˜ θf)=22+τ/2dσ2(27) where τ={0,1,...,30},drepresent the dimension of the input space, and σ2is the variance of the training data set taking a unitary value after normalization process. On the other hand, Fu et al. [23] propose a framework to estimate the parameter corresponding with the RBF-KPCA and RBF-KPLS models using the following optimization function. min Jf(˜ θf) = 1 mzp m X i=1 ε(i)T(˜ θf)ε(i)(˜ θf)(28) where the prediction zcorresponding to the data set {Xi}m i=1 ∈Rpis defining by the inverse mapping of the kernel model, and the prediction error is ε= (z−ˆz). More recently, Bernal et al. [34] assumed the false alarm rate and missing alarm rate as two measures that can be unifying through the AUC index to select the kernel parameter that best distinguishes normal from faulty process conditions. Then, the following fitness function is formulated. min Jf(˜ θf)=1−1 c c X i=1 AUCi(˜ θf)(29) where there are cmutually exclusive classes representing the fault operation conditions of the process. Beyond that, Pilario et al. [36] raised that a priori fault information could not be available for complex processes. Thereby, the normal operation data sets could be divided into two data sets (i.e., SET1: training data and SET2: validation data) to evaluate the monitoring schemes. In this case, Pilario et al. [36] employed the grid search method to decide the number of states and RBFpoly kernel parameters used for an MK-CVDA model, such that the false alarms in the validation data set are minimized. The following optimization criterion is also a proposal of the present paper and will be used to show the experimental results later. This updated index combines the ideas suggest by [34, 36] both applied for the incipient fault monitoring. min ˜ θf Jf(˜ θf) = FAR(˜ θf|fault-free) + 1−1 c c X i=1 AUCi(˜ θf|fault)!(30) JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 8 where FAR(˜ θf|fault-free)represent the performance of the monitoring scheme during the normal operating conditions and care the monitoring fault conditions. 3.3 Kernel evaluation framework The selection and adjusting of the kernel functions for fault diagnosis applications can be both highly dependent on the problem at hand. However, there are some general guidelines and procedures that might contribute to achieving better use of kernel-based methods in this field. Li et al. [139] provide a framework for kernel-based learning that is extending to fault diagnosis applications in the present work. Figure 2 depicts the modified framework as four consecutive phases: (i) Kernel selection, (ii) Parameters optimization, (iii) Training of methods, and (iv) Validation to fault diagnosis. As already emphasized, the kernel function candidates could be select by the researchers’ experience/experimental studies. Depending on what kind of information it is expecting to extract in the fault diagnosis tasks should decide the best approach and objective function to choose the kernel parameters. Basic kernel selection Construting objetive function of kernel optimi]ation Solving optimized data dependent kernel parameter Objetive function Optimization algorithm Construting optimized kernel Optimized parameter Training Sample Prior knowledge of problem Optimized learning machines /kernel methods Optimized kernel Evaluate the kernel method in fault diagnosis tasks Test sample Result meet performed criteria Historical database of industrial process Kernel optimization Training tools Implementation No Yes Testing tools Optimized parameter of kernel method Fig. 2. Framework for kernel-based learning in faults diagnosis tasks. 4 EXPERIMENTAL STUDIES AND RESULTS As has already mentioned, the main goal of this work is to provide a general overview of kernel-based preprocessing methods, giving particular emphasis to the adjusting of kernel solutions. However, several kernel functions were selected to show experimental results using two different optimization criteria. 4.1 Qualitative analysis of kernel functions Although many studies have recognized the key role that has the kernel functions, there remains a systematic use of the kernel RBF for fault diagnosis applications without exploring other alternative functions. Next, different kernel functions are graphically exploring with the intention to find common rules that can help with their selection. In this context, the synthetic data set shown in Figure 3 is used to obtain each kernel matrices. -4 -2 0 2 4 x1 -4 -2 0 2 4 x2 Unnormalized data -2 -1 0 1 2 x1 -2 -1 0 1 2 x2 Normalized data ~N(0,1) (b)(a) Fig. 3. Free-fault data set used for the qualitative study of kernel functions. (a) Synthetic data, (b) Standardized synthetic data The following assumptions are considered throughout the exploratory study to conduct this from a fault diagnosis perspective. As can see from Figure 3, the first requirement is that the fault monitoring procedures begin with the data standardization by using the mean and covariance describing the normal operating condition. In this context, the centering and scaling of the kernel matrices are also assuming as necessary operations for the kernel-based fault monitoring. Let the kernel matrix Kij;∀i, j ={1, . . . , m}, resulting to evaluate the kernel functions, the centered kernel matrix is obtained as follows: Kc=K−1mK−K1m+ 1mK1m(31) where 1m∈Rm×mis a matrix in which each element is equal to 1/m. Under the conditions stated above, Figure 4(a-b) compares the kernel matrices K(xj,xj)before and after of carry out the operations of centering and scaling. Note that all kernel function in Figure 4(a-b) are tuned with the same kernel parameter θ= 2. For the additive kernel k13, which combining the polynomial and RBF kernel functions, were choose the parameters ω= 0.5, d= 1,α= 1, and θ= 2. It is apparent that the gaussian kernel family k1, k3, k5, and k6 provides values closer to one to indicate a high degree of similarity between the data, and zero for point out the dissimilarities. The above kernels can be easily adjusting through direct evaluation criteria by using the class label information. However, this is less evident to the kernel functions k8, k10, k11, 13, and k14. However, given the representation of kernels k5 and k6 into a more reduced interval, for two different data set could be difficult to distinguish their features. Another unanticipated finding on this exploratory study was that kernel k2 and k3 shown identical results. The JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 9 0 2 0.5 12 k(Xj,Xj) k1-RBF kernel 1 0 1 0 -1 -1 -2 -2 -0.2 212 0.2 k(Xj,Xj) k1-RBF kernel 1 0 0.6 0 -1 -1 -2 -2 0 2 0.5 12 k(Xj,Xj) k3-Rational kernel 1 0 1 0 -1 -1 -2 -2 0 2 0.5 12 k(Xj,Xj) k5-Laplacian kernel 1 0 1 0 -1 -1 -2 -2 0 2 0.5 12 k(Xj,Xj) k6-Exponential kernel 1 0 1 0 -1 -1 -2 -2 -0.2 212 0.2 k(Xj,Xj) k3-Rational kernel 1 0 0.6 0 -1 -1 -2 -2 -0.3 212 0.2 k(Xj,Xj) k5-Laplacian kernel 1 0 0.6 0 -1 -1 -2 -2 -0.2 212 0.2 k(Xj,Xj) k6-Exponential kernel 1 0 0.6 0 -1 -1 -2 -2 -1 212 0 k(Xj,Xj) k8-Multiq kernel 1 0 1 0 -1 -1 -2 -2 2 2 3 12 k(Xj,Xj) k8-Multiq kernel 1 4 00 -1 -1 -2 -2 -15 2 -7.5 12 k(Xj,Xj) k10-Power kernel 1 0 0 0 -1 -1 2-2 -10 2 0 12 k(Xj,Xj) k10-Power kernel 1 0 10 0 -1 -1 -2 -2 -3 2 -1.5 12 k(Xj,Xj) k11-Log kernel 1 0 0 0 -1 -1 -2 -2 -1 2 0.5 12 k(Xj,Xj) k11-Log kernel 1 0 2 0 -1 -1 -2 -2 -3 2 0 12 k(Xj,Xj) k13-RBFpoly kernel 1 0 3 0 -1 -1 -2 -2 -3 2 0 12 k(Xj,Xj) k13-RBFpoly kernel 1 0 3 0 -1 -1 -2 -2 1 2 2 12 k(Xj,Xj) k14kernel 1 0 3 0 -1 -1 -2 -2 -1.5 2 -0.5 12 k(Xj,Xj) k14-kernel 1 0.5 00 -1 -1 -2 -2 -2 -1 0 1 2 Observations 0 0.2 0.4 0.6 0.8 k(Xj,Xj) k1-RBF kernel = 1 = 2 = 5 = 10 -2 -1 0 1 2 Observations 0 0.2 0.4 0.6 0.8 k(Xj,Xj) k3-Rational kernel -2 -1 0 1 2 Observations 0 0.2 0.4 0.6 0.8 k(Xj,Xj) k5-Laplacian kernel -2 -1 0 1 2 Observations 0 0.2 0.4 0.6 0.8 k(Xj,Xj) k6-Exponential kernel -2 -1 0 1 2 Observations 0 0.2 0.4 0.6 0.8 j) k1-RBF kernel -2 -1 0 1 2 Observations 0 0.2 0.4 0.6 0.8 j) k3-Rational kernel -2 -1 0 1 2 Observations 0.2 0.4 0.6 0.8 j) k5-Laplacian kernel -2 -1 0 1 2 Observations 0 0.2 0.4 0.6 j) k6-Exponential kernel -2 -1 0 1 2 Observations 0 0.8 1.6 2.4 k(Xj,Xj) k13-RBFpoly kernel -2 -1 0 1 2 Observations 0.4 0.5 0.6 0.7 0.8 j) k13-RBFpoly kernel -2 -1 0 1 2 Observations -6 -4 -2 0 2 k(Xj,Xj) k14-kernel -2 -1 0 1 2 Observations 1 1.5 2 k14-kernel -2 -1 0 1 2 Observations 0.5 1.5 2.5 3.5 j) k8-multiquadratic -2 -1 0 1 2 Observations -20 -15 -10 -5 0 j) k10-Power kernel -2 -1 0 1 2 Observations -6 -4 -2 0 j) k11-Log kernel -2 -1 0 1 2 Observations 0 2 4 6 8 10 k(Xj,Xj) k11-Log kernel -2 -1 0 1 2 Observations 0 10 20 30 k(Xj,Xj) k10-Power kernel -2 -1 0 1 2 Observations -1.8 -1.1 -0.4 0.2 k(Xj,Xj) k8-multiquadratic (d) (c) (b) (a) Fig. 4. A qualitative comparison. (a) Kernel matrices K(xj, xj)obtained with parameter θ= 2; (b) Kernel matrices K(xj, xj)adjusting with parameter θ= 2, after the operations of centering and scaling; (c) Values of the main diagonal for kernel matrices K(xj, xj)after the operations of centering and scaling, adjusting with different parameters θ= 1,2,5,10; (d) Kernel vectors from K(¯xj, xj)after the operations of centering and scaling, adjusting with different parameters θ= 1,2,5,10. same reasoning applies to the case of kernel k14 and k16. Figure 4(c) compares the variation on the main diagonal values from the above kernel functions, but with different tuning parameters. The parabola shown in the figure gives an idea about the kernel parameter sensibility, and how much the kernel emphasizes the differences between the input data. Note that the minimal variation of the kernel k13 due to their parameters. Figure 4(d) illustrates the kernel vector K(¯x,xj)obtained to compare the mean of the data set with each observation. As in the previous case, a parabola with less dilatation represents data more closely to the center of the cluster. The analysis of similar graphical could help decide if using the approaches based on the inter-cluster distances to choose the kernel parameters. 4.2 Fault detection in study cases The following describes the kernel-based fault monitoring scheme utilized. Figure 5 illustrates the flowchart adopted to evaluate the kernel functions summarized in Table 1. Kernel transformation Eigenvalue decomposition {X}m i=1 principal components (PC) k( Xi , Xj )SPE(X) SPEUCL T2(X) T2 UCL + Fig. 5. Flowchart for KPCA-based fault monitoring scheme. In this work, SPE and Hotelling’s T2statistics were combining to derive the detectability conditions in the systematic and residual subspaces from the KPCA model. More details about the combination of these indices are given in [60, 140, 141]. Moreover, the criterion λi/sum(λi)> 10−4was employed to select the number of eigenvalues [142]. Upper control limits (UCL) with 99% significance were obtaining for these statistics through kernel density estimation (KDE) [36, 143]. The fault detection procedures were evaluated by using several performance indicators: detection delay (DD), false alarm rate (FAR), and missed detection rate (MDR). The detection delay also referred to as latency time, reflects the elapsed time before that faults are continuously detected. In this study, the latency is the time for which five consecutive alarms are detecting from the start of the simulated faults. As earlier stated, FAR and MDR indicators were also computed by using the Eqs (32) and (33). However, the Area under the curve (AUC) was also used to give a joint interpretation of the two above indicators [34]. FAR = No.of samples (J > JUCL|fault −free) total samples (fault −free) (32) MDR = No.of samples (J > JUCL|fault) total samples (fault) (33) On the other hand, the Bayesian optimization algorithm available in the mathematical assistant MatlabrR2018 was utilized to chosen all kernel parameters. This approach is a particular case of the nonlinear optimization where the algorithm decides which point to explore next based on the analysis of distribution over functions [144, 145]. More specifically, the Bayesian optimization with an expected improvement search (i.e., expected-improvement-plus) was employed to escape from local minimums, considering the estimated variables in the range [10−5,105]and sixty evaluations of the objective function. In this study, the JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 16 [74] B. Zhou and X. Gu, “Multi-block statistics local Kernel Principal Component Analysis algorithm and its application in nonlinear process fault detection,” Neurocomputing, vol. 376, pp. 222–231, 2020. [75] Y. Zhang and S. J. Qin, “Fault detection of nonlinear processes using multiway Kernel Independent Component Analysis,” Industrial & Engineering Chemistry Research, vol. 46, no. 23, pp. 7780–7787, 2007. [76] R. T. Samuel and Y. Cao, “Kernel Canonical Variate Analysis for nonlinear dynamic process monitoring,” IFAC-PapersOnLine, vol. 48, no. 8, pp. 605–610, 2015. [77] X. B. He, Y. P. Yang, and Y. H. Yang, “Fault diagnosis based on variable-weighted Kernel Fisher Discriminant Analysis,” Chemometrics and Intelligent Laboratory Systems, vol. 93, no. 1, pp. 27–33, 2008. [78] Z.-B. Zhu and Z.-H. Song, “Fault diagnosis based on imbalance modified Kernel Fisher Discriminant Analysis,” Chemical Engineering Research and Design, vol. 88, no. 8, pp. 936–951, 2010. [79] ——, “A novel fault diagnosis system using pattern classification on kernel FDA subspace,” Expert Systems with Applications, vol. 38, no. 6, pp. 6895–6905, 2011. [80] G. Rong, S.-Y. Liu, and J.-D. Shao, “Fault diagnosis by locality preserving discriminant analysis and its kernel variation,” Computers & Chemical Engineering, vol. 49, pp. 105–113, 2013. [81] Z. Liu, J. Qu, M. J. Zuo, and H.-b. Xu, “Fault level diagnosis for planetary gearboxes using hybrid kernel feature selection and Kernel Fisher Discriminant Analysis,” The International Journal of Advanced Manufacturing Technology, vol. 67, no. 5-8, pp. 1217–1230, 2013. [82] J. M. Bernal-de L´ azaro, A. Moreno-Prieto, O. Santiago-Llanes, and A. J. Silva-Neto, “Optimizing kernel methods to reduce dimensionality in fault diagnosis of industrial systems,” Computers & Industrial Engineering, vol. 87, pp. 140–149, 2015. [83] K. Peng, K. Zhang, G. Li, and D. Zhou, “Contribution rate plot for nonlinear quality-related fault diagnosis with application to the hot strip mill process,” Control Engineering Practice, vol. 21, no. 4, pp. 360–369, 2013. [84] H. Zhang, K. Peng, K. Zhang, and G. Li, “Quality-related process monitoring based on total kernel PLS model and its industrial application,” Mathematical Problems in Engineering, vol. 2013, p. 707953, Dec. 2013. [85] H. Shi, J. Liu, and Y. Zhang, “An optimized Kernel Principal Component Analysis algorithm for fault detection,” IFAC Proceedings Volumes, vol. 42, no. 8, pp. 846–851, 2009. [86] M. Jia, H. Xu, X. Liu, and N. Wang, “The optimization of the kind and parameters of kernel function in KPCA for process monitoring,” Computers & Chemical Engineering, vol. 46, pp. 94–104, 2012. [87] M. Van and H.-J. Kang, “Wavelet kernel local Fisher Discriminant Analysis with Particle Swarm Optimization algorithm for bearing defect classification,” IEEE Transactions on Instrumentation and Measurement, vol. 64, no. 12, pp. 3588–3600, 2015. [88] H. Shi, J. Liu, Y. Wu, K. Zhang, L. Zhang, and P. Xue, “Fault diagnosis of nonlinear and large-scale processes using novel modified Kernel Fisher Discriminant Analysis approach,” International Journal of Systems Science, vol. 47, no. 5, pp. 1095–1109, 2016. [89] M. S. Kan, A. C. Tan, and J. Mathew, “A review on prognostic techniques for non-stationary and non-linear rotating systems,” Mechanical Systems and Signal Processing, vol. 62, pp. 1–20, 2015. [90] Y. Zhang, H. Zuo, and F. Bai, “Classification of fault location and performance degradation of a roller bearing,” Measurement, vol. 46, no. 3, pp. 1178–1189, 2013. [91] S. Dong, D. Sun, B. Tang, Z. Gao, Y. Wang, W. Yu, and M. Xia, “Bearing degradation state recognition based on kernel PCA and wavelet kernel SVM,” Proceedings of the Institution of Mechanical Engineers, Part C: Journal of Mechanical Engineering Science, vol. 229, no. 15, pp. 2827–2834, 2015. [92] J. Sheng, S. Dong, and Z. Liu, “Bearing fault diagnosis based on intrinsic time-scale decomposition and improved Support Vector Machine model,” Journal of Vibroengineering, vol. 18, no. 2, pp. 849–859, 2016. [93] R. Shao, W. Hu, Y. Wang, and X. Qi, “The fault feature extraction and classification of gear using Principal Component Analysis and Kernel Principal Component Analysis based on the Wavelet packet transform,” Measurement, vol. 54, pp. 118–132, 2014. [94] Z. Liu, W. Guo, J. Hu, and W. Ma, “A hybrid intelligent multi-fault detection method for rotating machinery based on RSGWPT, KPCA and Twin SVM,” ISA Transactions, vol. 66, pp. 249–261, 2017. [95] X. Wan, D. Wang, W. T. Peter, G. Xu, and Q. Zhang, “A critical study of different dimensionality reduction methods for gear crack degradation assessment under different operating conditions,” Measurement, vol. 78, pp. 138–150, 2016. [96] N. Sakthivel, S. Saravanamurugan, B. B. Nair, M. Elangovan, and V. Sugumaran, “Effect of kernel function in Support Vector Machine for the fault diagnosis of pump,” Journal of Engineering Science and Technology, vol. 11, no. 6, pp. 826–838, 2016. [97] P. K. Wong, Z. Yang, C. M. Vong, and J. Zhong, “Real-time fault diagnosis for gas turbine generator systems using extreme learning machine,” Neurocomputing, vol. 128, pp. 249–257, 2014. [98] L. Jiang, B. Zeng, F. R. Jordan, and A. Chen, “Kernel function and parameters optimization in kica for rolling bearing fault diagnosis,” Journal of Networks, vol. 8, no. 8, p. 1913, 2013. [99] B. Ma, J. D. Wu, J. Ma, X. D. Wang, and Y. G. Fan, “Fault monitoring and classification method of rolling bearing based on KICA and LSSVM,” in Advanced Materials Research, vol. 971. Trans Tech Publ, 2014, pp. 476–480. [100] Z. Li, X. Yan, Z. Tian, C. Yuan, Z. Peng, and L. Li, “Blind vibration component separation and nonlinear feature extraction applied to the nonstationary vibration signals for the gearbox multi-fault diagnosis,” Measurement, vol. 46, no. 1, pp. 259–271, 2013. JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 17 [101] Z. C. Li, “New detection method for gear faults based on kernel independent component analysis and bp neural network,” in Advanced Materials Research, vol. 909. Trans Tech Publ, 2014, pp. 371–374. [102] L. Jiang, J. Xuan, and T. Shi, “Feature extraction based on semi-supervised Kernel Marginal Fisher Analysis and its application in bearing fault diagnosis,” Mechanical Systems and Signal Processing, vol. 41, no. 1-2, pp. 113–126, 2013. [103] S. W. Choi and I.-B. Lee, “Nonlinear dynamic process monitoring based on dynamic kernel PCA,” Chemical Engineering Science, vol. 59, no. 24, pp. 5897–5908, 2004. [104] I. Jaffel, O. Taouali, M. F. Harkat, and H. Messaoud, “Moving window KPCA with reduced complexity for nonlinear dynamic process monitoring,” ISA transactions, vol. 64, pp. 184–192, 2016. [105] Q. Zhang, P. Li, X. Lang, and A. Miao, “Improved dynamic Kernel Principal Component Analysis for fault detection,” Measurement, p. 107738, 2020. [106] Y. Zhang and Z. Hu, “On-line batch process monitoring using hierarchical Kernel Partial Least Squares,” Chemical Engineering Research and Design, vol. 89, no. 10, pp. 2078–2084, 2011. [107] Y. Zhang, S. Li, Z. Hu, and C. Song, “Dynamical process monitoring using dynamical hierarchical Kernel Partial Least Squares,” Chemometrics and Intelligent Laboratory Systems, vol. 118, pp. 150–158, 2012. [108] Q. Jia and Y. Zhang, “Quality-related fault detection approach based on dynamic Kernel Partial Least Squares,” Chemical Engineering Research and Design, vol. 106, pp. 242–252, 2016. [109] G. Stefatos and A. B. Hamza, “Dynamic Independent Component Analysis approach for fault detection and diagnosis,” Expert Systems with Applications, vol. 37, no. 12, pp. 8606–8617, 2010. [110] J. Mori and J. Yu, “Quality relevant nonlinear batch process performance monitoring using a kernel based multiway non-Gaussian latent subspace projection approach,” Journal of Process Control, vol. 24, no. 1, pp. 57–71, 2014. [111] L. Cai, X. Tian, and S. Chen, “Monitoring nonlinear and non-gaussian processes using gaussian mixture model-based weighted Kernel Independent Component Analysis,” IEEE Transactions on neural networks and learning systems, vol. 28, no. 1, pp. 122–135, 2015. [112] Y. Liu, F. Wang, Y. Chang, F. Gao, and D. He, “Performance-relevant Kernel Independent Component Analysis based operating performance assessment for nonlinear and non-Gaussian industrial processes,” Chemical Engineering Science, vol. 209, p. 115167, 2019. [113] X. Deng, N. Zhong, and L. Wang, “Nonlinear multimode industrial process fault detection using modified Kernel Principal Component Analysis,” IEEE Access, vol. 5, pp. 23 121–23 132, 2017. [114] Y. Chen, X. Deng, and Y. Cao, “Nonlinear soft sensor modeling method based on multimode Kernel Partial Least Squares assisted by improved KFCM clustering,” in Chinese Automation Congress (CAC). IEEE, 2019, pp. 4245–4250. [115] R. Tan, T. Cong, J. R. Ottewill, J. Baranowski, and N. F. Thornhill, “An on-line framework for monitoring nonlinear processes with multiple operating modes,” Journal of Process Control, vol. 89, pp. 119–130, 2020. [116] Y. Hu, H. Ma, and H. Shi, “Enhanced batch process monitoring using just-in-time-learning based Kernel Partial Least Squares,” Chemometrics and Intelligent Laboratory Systems, vol. 123, pp. 15–27, 2013. [117] H. Zhang, X. Tian, X. Deng, and Y. Cao, “Batch process fault detection and identification based on discriminant global preserving Kernel Slow Feature Analysis,” ISA Transactions, vol. 79, pp. 108–126, 2018. [118] W. Li and S. Hongbo, “Improved kernel PLS-based fault detection approach for nonlinear chemical processes,” Chinese Journal of Chemical Engineering, vol. 22, no. 6, pp. 657–663, 2014. [119] X. Deng and J. Deng, “Incipient fault detection for chemical processes using two-dimensional weighted SLKPCA,” Industrial & Engineering Chemistry Research, vol. 58, no. 6, pp. 2280–2295, 2019. [120] X. Deng, P. Cai, Y. Cao, and P. Wang, “Two-step localized Kernel Principal Component Analysis based incipient fault diagnosis for nonlinear industrial processes,” Industrial & Engineering Chemistry Research, vol. 59, no. 13, pp. 5956–5968, 2020. [121] P. Cai and X. Deng, “Incipient fault detection for nonlinear processes based on dynamic multi-block probability related Kernel Principal Component Analysis,” ISA Transactions, 2020. [122] R. Vitale, O. E. de Noord, and A. Ferrer, “A kernel-based approach for fault diagnosis in batch processes,” Journal of Chemometrics, vol. 28, no. 8, pp. S697–S707, 2014. [123] Y. Zhang, R. Sun, and Y. Fan, “Fault diagnosis of nonlinear process based on KCPLS reconstruction,” Chemometrics and Intelligent Laboratory Systems, vol. 140, pp. 49–60, 2015. [124] Z. Li, G. Tan, and Y. Li, “Fault diagnosis based on improved Kernel Fisher Discriminant Analysis.” Journal of Software, vol. 7, no. 12, pp. 2657–2662, 2012. [125] Y. Hu, H. Ma, and H. Shi, “Robust online monitoring based on spherical Kernel Partial Least Squares for nonlinear processes with contaminated modeling data,” Industrial & Engineering Chemistry Research, vol. 52, no. 26, pp. 9155–9164, 2013. [126] A. R. Teixeira, A. M. Tom´ e, K. Stadlthanner, and E. W. Lang, “KPCA denoising and the pre-image problem revisited,” Digital Signal Processing, vol. 18, no. 4, pp. 568–580, 2008. [127] L. Zhang, W. Zhou, P. Chang, J. Liu, Z. Yan, T. Wang, and F. Li, “Kernel sparse representation-based classifier,” IEEE Transactions on Signal Processing, vol. 60, no. 4, pp. 1684–1695, 2012. [128] K.-P. Wu and S.-D. Wang, “Choosing the kernel parameters for Support Vector Machines by the inter-cluster distance in the feature space,” Pattern Recognition, vol. 42, no. 5, pp. 710–717, 2009. [129] R. Ziani, A. Felkaoui, and R. Zegadi, “Bearing fault diagnosis using multiclass Support Vector JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 18 Machines with binary Particle Swarm Optimization and regularized Fishers criterion,” Journal of Intelligent Manufacturing, vol. 28, no. 2, pp. 405–417, 2014. [130] H. Xiong, M. Swamy, and M. O. Ahmad, “Optimizing the kernel in the empirical feature space,” IEEE Transactions on Neural Networks, vol. 16, no. 2, pp. 460–474, 2005. [131] B. Chen, H. Liu, and Z. Bao, “Optimizing the data-dependent kernel under a unified kernel optimization framework,” Pattern Recognition, vol. 41, no. 6, pp. 2107–2119, 2008. [132] N. Cristianini, J. Shawe-Taylor, A. Elisseeff, and J. S. Kandola, “On kernel-Target Alignment,” in Advances in Neural Information Processing Systems, 2002, pp. 367–373. [133] T. Wang, D. Zhao, and S. Tian, “An overview of kernel alignment and its applications,” Artificial Intelligence Review, vol. 43, no. 2, pp. 179–192, 2015. [134] C. H. Nguyen and T. B. Ho, “An efficient kernel matrix evaluation measure,” Pattern Recognition, vol. 41, no. 11, pp. 3366–3372, 2008. [135] Y. Baram, “Learning by kernel polarization,” Neural Computation, vol. 17, no. 6, pp. 1264–1275, 2005. [136] J.-g. WANG, X.-j. CHEN, and W.-x. ZHANG, “A research on the bearing fault diagnosis algorithm of Least Squares Support Vector Machine with multiple kernels,” Modular Machine Tool & Automatic Manufacturing Technique, no. 6, p. 19, 2017. [137] P. Chudzian, “Evaluation measures for kernel optimization,” Pattern Recognition Letters, vol. 33, no. 9, pp. 1108–1116, 2012. [138] N. Zhang, X. Gao, Y. Li, and P. Wang, “Fault detection of chiller based on improved KPCA,” in Chinese Control and Decision Conference (CCDC). IEEE, 2016, pp. 2951–2955. [139] J.-B. Li, Y.-H. Wang, S.-C. Chu, and J. F. Roddick, “Kernel self-optimization learning for kernel-based feature extraction and recognition,” Information Sciences, vol. 257, pp. 70–80, 2014. [140] H. H. Yue and S. J. Qin, “Reconstruction-based fault identification using a combined index,” Industrial & Engineering Chemistry Research, vol. 40, no. 20, pp. 4403–4414, 2001. [141] C. F. Alcala and S. J. Qin, “Analysis and generalization of fault diagnosis methods for process monitoring,” Journal of Process Control, vol. 21, no. 3, pp. 322–330, 2011. [142] J.-M. Lee, S. J. Qin, and I.-B. Lee, “Fault detection of non-linear processes using Kernel Independent Component Analysis,” The Canadian Journal of Chemical Engineering, vol. 85, no. 4, pp. 526–536, 2007. [143] P.-E. P. Odiowei and Y. Cao, “Nonlinear dynamic process monitoring using Canonical Variate Analysis and kernel density estimations,” IEEE Transactions on Industrial Informatics, vol. 6, no. 1, pp. 36–45, 2009. [144] E. Brochu, V. M. Cora, and N. De Freitas, “A tutorial on Bayesian optimization of expensive cost functions, with application to active user modeling and hierarchical reinforcement learning,” arXiv preprint arXiv:1012.2599, 2010. [145] R. Martinez-Cantin, “Bayesopt: A bayesian optimization library for nonlinear optimization, experimental design and bandits,” The Journal of Machine Learning Research, vol. 15, no. 1, pp. 3735–3739, 2014. [146] D. Dong and T. J. McAvoy, “Nonlinear Principal Component Analysis based on principal curves and neural networks,” Computers & Chemical Engineering, vol. 20, no. 1, pp. 65–78, 1996. PLACE PHOTO HERE Jos´ e M. Bernal de L´ azaro is graduated in Automatic Control Engineering from the Universidad Tecnol´ ogica de La Habana Jos´ e A. Echeverr´ ıa (CUJAE, 2010). He received the M.Sc. in Mathematical Modeling applied to Engineering in 2014, and the Ph.D. degree in Applied Sciences from this university in 2016. He is currently working as a Researcher and Full Professor at the Computer Systems and Automation Department of the Universidad Tecnol´ ogica de La Habana. His scientific interests are in the fields of Control and Process Systems Engineering, with emphasis on problems of pattern recognition and fault diagnosis in industrial environments. PLACE PHOTO HERE Orestes Llanes Santiago obtained the Degree of Electrical Engineer from Universidad Tecnol´ ogica de la Habana Jos´ e A. Echeverr´ ıa (CUJAE, 1981). He pursued graduate studies from 1989 to 1994 at the Universidad de Los Andes, M´ erida-Venezuela, where he obtained the degree of Master of Science in Control Engineering in 1990 and the degree of Doctor in Applied Sciences in 1994. He is currently a Researcher and Full Professor at the Computer Systems and Automation Department of the Universidad Tecnol´ ogica de La Habana in Cuba. His areas of interest are the Fault diagnosis in industrial systems, Nonlinear control and Computational intelligence with applications to control. PLACE PHOTO HERE Antˆ onio J. da Silva Neto is a Mechanical/Nuclear Engineer (Universidade Federal do Rio de Janeiro−UFRJ, Brazil, 1983), M.Sc. in Nuclear Engineering (UFRJ, 1989), and Ph.D. in Mechanical Engineering, with a minor in Computational Mathematics (North Carolina State University, USA, 1993). In 1997 he joined the UERJ, being currently a Full Professor at the Polytechnic Institute. His areas of interest are Mechanical Engineering and the Applied and Computational Mathematics, with emphasis in numerical methods and inverse problems. PLACE PHOTO HERE Carlos Cruz Corona is graduated in Automatic Control Engineering from the Universidad Tecnol´ ogica de La Habana Jos´ e A. Echeverr´ ıa (CUJAE, 20??). He received the M.Sc. in Mathematical Modeling applied to Engineering in 20??, and the Ph.D. degree in Applied Sciences from this university in 20??. He is currently working as a Researcher and Full Professor at the Computer Systems and Automation Department of the Universidad Tecnol´ ogica de La Habana.