scieee AI-readable full text Open interactive document viewer

Model-based segmentation of long term SCADA data in the presence of heavy-tailed distributed noise with finite-variance

Shiri, Hamid

Abstract

Pre-print of paper title'' Data-driven segmentation of long term condition monitoring data in the presence of heavy-tailed distributed noise with finite-variance''

Full text

Model-based segmentation of long term SCADA data in the presence of heavy-tailed distributed noise with finite-variance Hamid Shiria,∗, Pawel Zimroza, Jacek Wodeckia, Agnieszka Wyłoma´ nskab, Radoslaw Zimroza aFaculty of Geoengineering, Mining and Geology, Wroclaw University of Science and Technology, Na Grobli 15, 50-421 Wroclaw, Poland bFaculty of Pure and Applied Mathematics, Hugo Steinhaus Center, Wroclaw University of Science and Technology, Wyspianskiego 27, 50-370 Wroclaw, Poland Abstract Machinery condition prognosis systems use long-term historical data to predict the remaining useful life (RUL). One of the critical steps to reach this purpose is to segment long-term data into two or several degradation stages (healthy, unhealthy, critical stage). Finding a changing point between stages may be a crucial preliminary task for the further prediction of the degradation process. However, finding the accurate partition into two or more stages is a challenging task in the actual application when the noise inherent in the observed process exhibits non-Gaussian characteristics. In this paper, a framework for model-based segmentation is presented for the prognosis of long-term machinery data in the presence of heavy-tailed distributed noise with finite variance. It is assumed that three different stages are inherent in the degradation process and that each segment of the data follows a specific trend (constant, linear, exponential, or polynomial). At first, the data is divided into three parts. The trend functions are fitted to the data by using the robust regression method, and the cumulative error is calculated. This process is done iteratively for all possible partitions into three intervals to find the segmentation which minimizes the error. The framework has been tested via empirical analysis of the estimators of the changing points obtained in Monte Carlo simulations. Also, the discussed approaches are applied to the real data. In such measurement, the data which is commonly available (in SCADA systems) is aggregated from the raw signal and sampled at long intervals. Finally, the effectiveness of the segmentation results is assessed by comparing them with envelope frequency analysis of the raw signal to confirm the fact that the detected changing points coincide with the start time of the fault in the machine or not. Keywords: Long-term data, Segmentation, Robust methods, Non-Gaussian noise, Changing points. 1. Introduction Prognostics and Health Management (PHM) is one of the significant tasks in condition-based maintenance (CBM). The CBM systems are developed to assess the machine’s condition by collecting massive amounts of data during the operation of machines. PHM is usually made up of four different processes [1, 2]: 1. Data acquisition is defined as the process of recording and saving different types of monitoring data from various sensors mounted on the monitored5 machine. 2. Construction of a health index (HI): in general, HI is defined as the set of statistical and frequency domainbased features extracted from raw signals to describe machinery’s health condition. 3. Health stage (HS) evaluation and 4. Prediction of RUL. The HS evaluation is taken as an expression of fault detection in the PHM community [3]. Nevertheless, their assignments are different from each other. The diagnosis of faults is determined by the pattern and intensity of the machine fault. However, the primary purpose of the HS evaluation is to divide the continuous10 degradation process into several health stages due to the varying trends of HI. In addition, it is hard and undeserved to predict the RUL during a healthy stage. The RUL estimation should start from the unhealthy stage’s start time, known as the first predicting time (FPT), due to the fact that in the healthy stage observations there is no valuable information about the unhealthy stage and its degradation trend. Therefore, instead of trying too hard to employ the ∗[email protected] Preprint submitted to Mechanical Systems and Signal Processing November 12, 2025 complex model, such as the model with switching regimes or nonlinear trends to predict RUL, breaking the HI into15 the different HS and trying to predict RUL from the last stage may be more accessible and evaluative. Commonly, for HS evaluation, the health index is compared to a limit value that usually is provided by the manufacturer (threshold corresponding to change from ”Good Condition” (healthy stage) to ”Warning” (degradation stage) and ”Warning” to ”Alarm” stage (critical stage). Unfortunately, we do not know the limit values or desired lifetime in many cases, which is often present in the raw materials industry [4], especially when the machine is unique. Also, it should be noted that20 most of these thresholds are introduced by manufacturing industries and are defined based on particular working and environmental conditions standards that may not be valid when the conditions change. In recent years, several papers have been published in the literature to detect changing points in time series in different areas such as financial, medical and meteorology [5–14]. Also, it should be noted that this process is known as regime-switching point detection or signal segmentation in other communities. Signal segmentation is repeatedly25 employed in signal processing applications to separate original data into homogeneous segments or extract the pattern. Gasior et al. [15] utilized segmentation for shock extraction in sieving screen vibrations. Kucharczyk et al. [16] used stochastic modelling for seismic signal segmentation. Grzesiek et al. [17] proposed a methodology to detect regime change when regime A smoothly transforms into regime B. Also, a few papers have been published in the PHM community based on dividing HI into different HS. Some of these articles divide HI into two different stages. Alkan et30 al. [18] presented a methodology for the diagnosis of electromechanical systems faults based on the variance-sensitive adaptive alarm threshold and principal component analysis (PCA). Fink et al. [19] explained the prediction of RUL as a two-stage classification to detect the machine’s condition after the defined time interval. Schlechtingen et al. [20] compared three different model-based approaches for the two-stage division of wind turbines, i.e., a regressionbased model and two artificial neural networks (ANN) based models. Hu et al.[21] used a one-class support vector35 machine and a Gaussian threshold model for condition monitoring of the turbo-pumps. The two-stage division is valuable only for cases where the degradation trends of machinery in the unhealthy stage are consistent and can be represented utilizing a single degradation model. Nonetheless, due to variations of fault patterns or operational conditions, the degradation trends of machinery may change. According to this circumstance, the unhealthy stage should be split into various stages based on the different degradation trends. A few of these research divided HI and40 spectra into several stages, with changing points detection approaches. For instance, Kimotho et al. [22] segmented the degradation trend into five stages due to variation of frequency amplitude in power spectra density. Sutrisno et al. [23] split the bearing degradation process into different stages by employing anomaly detection of frequency spectra. Hu et al. [24], separated the HI into four stages using changing points of confidence levels. Methods based on machine learning techniques are frequently used for long-term data analysis for classification, fault detection,45 prognosis, etc. [4, 25–31]. In this case, unsupervised classification, i.e. clustering, is required because every machine works under different conditions and labeling data is not an easy task. Hence, the clustering approach has a huge potential to divide long-term data into several regimes. The authors of [26] proposed the Long Short Term Memory (LSTM) network with the clustering method for multi-stage predicting RUL. Jaskaran Singh et al. [27] introduced an adaptive data-driven model-based approach to detect regime-changing points using K-means clustering. Also, discrete50 state transition models such as HMMs [32–37] and dynamic state-space models [38, 39] are frequently employed to segment degradation processes into multiple stages. However, most of the mentioned research was performed under the assumption that observation noise has Gaussian distribution, while in most actual applications, especially in the PHM area it is not a proper assumption. In many cases, the impulsive noise can be observed, which is properly described with a heavy-tailed distribution, meaning that the marginal events are likely to occur. This may happen for55 several reasons. Wind inflows are one of the possible sources of impulsive noise in wind turbine [40–42], the ore falling on devices in the mining environment is another [43, 44]. Due to the specific random character of the degradation process observed in the HI data, classical algorithms known from the literature may fail to properly perform the segmentation procedure [45–47]. To solve this issue, in the paper, we propose the model-based approach according to historical data to robustly identify the border of the60 stages (healthy/degradation/critical stage) when the machine changes the condition in the presence of non-Gaussian noise with finite variance. It is considered that the degradation process is composed of three stages and each stage of the data follows a specific trend (constant, linear and exponential or polynomial). First, the long-term historical data is divided into three parts. The trend functions are fitted to the data using the regression method and the cumulative error is calculated. This process is done iteratively for all possible partitions into three intervals to find the changing65 points, which minimizes the error. This procedure is repeated using different regression methods including robust 2 methods such as the least absolute error and fitting the Student’s t-distribution. It should be noted that, the Student’s t-distribution was chosen because it is both heavy-tailed (unlike Gaussian distribution) and at the same time has finite variance (for the considered range of degree of freedom) making it easier to fit than some other theoretically well established distributions. Also, in our previous paper we showed that this distribution can describe a random part as70 proper [48]. Furthermore, for better comparison and to check if the non-Gaussian noise assumption is necessary, along with the mentioned approach we have included two hidden Markov model based methods having the Gaussian noise assumption (These two methods use a relatively different structure for segmentation). We also presented a waterfall plot of the short-time envelope spectrum calculated for available raw vibration data for the real data sets. Such a 3D plot demonstrates how the envelope spectrum changes and allow us to visually validate that detected changing points75 are coinciding with the time the fault in the machine occurs or not. The paper is structured as follows: after the introduction, in Section 2, we reviewed the most common longterm data model and described the model used in this paper. In Section 3, we have reviewed the theoretical and mathematical aspects of the methods used in this paper. In Section 4, we generated data based on Section 2 and employed the proposed structure to segment long-term data in the presence of Gaussian and non-Gaussian noise.80 Finally, in Section 5 the results of applying the proposed approach to three benchmark data sets are presented with an indication of all intermediate steps, and in Section 6 the conclusions are formed. 2. Long term data model Model-based degradation process analysis is a common approach in the PHM community. These models can be described by various deterministic trends (linear, polynomial, exponential and constant), see Fig.1, various noise85 distributions, and structures of dependence for random components; see Fig. 2, 3. According to the pre-knowledge of the degradation process, the model can be selected; for example, the linear model is frequently used for describing the drilling degradation process [1, 49], see panel (b) Fig. 1. Meanwhile, the exponential degradation trend is usually used for expressing the battery degradation process [50–52] (but with a negative slope), see panel (c) in Fig.1. To describe the degradation process of bearings, gears or shafts, different kinds90 of models composed from various deterministic trends may be used [53, 54] (see Fig. 1, panel (d), (e), (f)). In this paper, we assume that the degradation process has complex trends, such as those presented in Fig. 3. This model consists of three stages. At first the degradation process is centered around some constant level which correspond to the healthy stage. The second stage is described with a linear function corresponding to the degradation stage. The last part has either an exponential or a polynomial trend that is referred to as the critical stage. The two95 moments of the transition from one stage to another we will denote as changing points: CP1 (is a border between the healthy stage with constant trend and degradation stage with linear trend) and CP2 (is a border between the degradation stage with linear trend and critical stage with exponential or polynomial trend). Also, this model can include different kinds of noise distribution covering the degradation process. In most of the research in this area, it is assumed the distribution of the noise is Gaussian, see Fig. 2 [55], while this assumption is not100 proper when the machine works in harsh environments [48]. To achieve a model that better reflects the real condition, in our previous work [48] we developed a simulation framework that produces degradation process observations; an exemplary trajectory is shown in Fig.3. In this framework along with the trend of the degradation process, the distribution of the noise and its scale (square root of the variance in the Gaussian case) are also changing in each stage. The main characteristics of our framework is summarised in Table 1. In this paper, we assume of finite variance105 for non-Gaussian distribution. 3 Figure 1: Different types of HI variations and degradation models: (a) good condition (constant trend), (b) good to gradual wear (linear trend), (c) exponential trend, (d) good to accelerated wear (linear trend), (e) good to accelerated wear (exponential trend), (f) three stages model (good, linear progress and exponential(polynomial) progress of degradation) [48]. Figure 2: Three stage degradation model [56]. 4 Figure 3: The proposed long term degradation model with 3 stages [48]. Table 1: Main characteristics of the data for the three stages indicated in Fig. 3. [48]. stage 1 stage 2 stage 3 Trend constant linear exponential or polynomial Scale nearly constant linearly growing lin. or exp. growing Random component Gaussian/non-Gaussian Gaussian/non-Gaussian Gaussian/non-Gaussian 3. Methodology We assume that the long-term health index (HI) data are composed of three distinct stages that expose significantly different behaviors. Based on this assumption, our aim is to divide the time series into these three sections by using robust mathematical regression methods, achieving best fitting to one of the specified models. All of the methods110 that we discuss assume that the health index is a stochastic process (its observations we denote by Xtwhere tis the discrete-time index) and that its trend is a function of time and a set of parameters. Thus, the HI observation process consists of the time-changing trend (expected value of Xt, which we assume to be finite) and residual rt, hence we have: Xt=b Xt+rt,b Xt=EXt.(1) As Θin this section we denote the complete set of parameters to be fitted to the data. It may include the changing115 points (see previous section) and the parameters of a probability distribution if a method employs any. Both terms on the right-hand side of Eq. (1) depend on Θ. We now present the methods in two groups. 3.1. Piece-wise regression models This class of methods allows to divide data into three segments, each having different form of the model. We employ three functional components: constant, linear, and exponential as below:120 b Xt(Θ)=             f1(t, θ1)=θ1,0<t≤τ1, f2(t, θ2, θ3)=θ2t+θ3, τ1<t≤τ2, f3(t, θ4, θ5, θ6)=θ4exp(θ5t)+θ6, τ2<t≤N, (2) where the variables τ1and τ2stand for CP1 and CP2, the number Nis signal length, and θ1, ..., θ6are constants (which need to be fitted) parameters used to model deterministic parts of the HI. 5 3.1.1. Ordinary least squares (OLS) The most basic approach is based on minimizing the sum of squares of the residuals of the model. The procedure designed to find an optimal set of parameters iterates in two loops through possible divisions into three windows125 (which allows us to find optimal τ1and τ2). b Θ = argmin Θ τ1−1 X t=1 (Xt−f1(t, θ1))2+ τ2−1 X t=τ1 (Xt−f2(t, θ2, θ3))2+ N X t=τ2 (Xt−f3(t, θ4, θ5, θ6))2.(3) In the case of observations with Gaussian noise, given each division, optimal parameters can be equivalently estimated using Maximum Likelihood Estimation (MLE) with Gaussian distributed residuals. 3.1.2. Dynamic programming segmentation In this method the optimal segmentation of the time-series of the observations Xt=1,...,Nis done by fitting the model130 parameters in terms of Gaussian distribution MLE by using optimized dynamic programming. The computation cost of this model is high based on using dynamic programming. For more details on this method, please see the following references [57, 58]. 3.1.3. Iteratively reweighted least squares (IRLS) The method of iteratively reweighted least squares (IRLS) is a modification of the OLS method where the weights135 are assigned to the error terms according to changing variance which we assume to be finite (for each time tthe weight wt=1/σtis reciprocal of the corresponding standard deviation σtof the error term of the observation). The weights are modified in an iterative way; thus, in the j-th step, the vector of weights w(j) t,t=1,...,Nis applied, which results in updating the piecewise regression parameters: b Θ(j)=argmin Θ τ1−1 X t=1 w(j) t(Xt−f1(t, θ(j−1) 1))2+ τ2−1 X t=τ1 w(j) t(Xt−f2(t, θ(j−1) 2, θ(j−1) 3))2+ N X t=τ2 w(j) t(Xt−f3(t, θ(j−1) 4, θ(j−1) 5, θ(j−1) 6))2,(4) where b Θ(j)is the current values set of Θin the j-th iteration of the IRLS algorithm. Among many different iterative140 methods of finding the proper weights, we used a method applying to the residuals of the model the Tukey biweight function ϕ(x)=x(1 −x2)2(as defined in [59]) implemented in Matlab robustfit() function. The weights are updated using the formula [60]: w(j) t=ϕ(ε(j) t) ε(j) t ,(5) where ε(j) tis the residual of the model divided by the current estimate of the standard deviation term, that is: ε(j) t=rt σ(j) t .(6) 3.1.4. Least absolute error (LAE)145 The sum of absolute deviations is used as a cost function instead of a sum of squares. It is a widely known alternative to the OLS method, making the fitting of the regression parameters more robust, even in the case of infinite variance [61]. The optimization procedure including LAE cost function with changinng points is defined as follows: b Θ = argmin Θ τ1−1 X t=1|Xt−f1(t, θ1)|+ τ2−1 X t=τ1|Xt−f2(t, θ2, θ3)|+ N X t=τ2|Xt−f3(t, θ4, θ5, θ6)|.(7) The same as in the OLS method we iterate through all the possible divisions into three stages and in that way we find τ1and τ2. This method is beneficial when the errors have rather a heavy-tailed than Gaussian distribution.150 Namely given each division into three windows, it can be equivalently stated as applying MLE estimation with Laplace distributed residuals. 6 3.1.5. Student’s t-distribution estimation (ST) In this method, the residuals of the model in each degradation stage are assumed to have a scaled Student’s t distribution, that is:155 rt=             σ1ut,0<t≤τ1, σ2ut, τ1<t≤τ2, σ3ut, τ2<t≤N, (8) where {ut}t=1,...,Nis a series of independent random variables from the Student’s t distribution with νidegrees of freedom for the i-th segment (i=1,2,3). This enables the heavy tails in the error distribution and makes the fitting more robust allowing proper parametrization of the non-Gaussian observation data. The estimation procedure is based on MLE technique applied to each possible division into three stages. The optimal set of parameters is found to fulfill the following:160 b Θ = argmin Θ τ1−1 X t=1 log pν1Xt−f1(t, θ1) σ1+ τ2−1 X t=τ1 log pν2Xt−f2(t, θ2, θ3) σ2+ N X t=τ2 log pν3Xt−f3(t, θ4, θ5, θ6) σ3,(9) where pν(·) is the probability density function (PDF) of the Student’s t-distribution given by the formula [62]: pν(y)=(1 +y2 ν)−ν+1 2 B(ν 2,1 2)√ν,y∈R,(10) where B(·,·) is the beta function. The parameter νis called the number of the degrees of freedom. We restrict the parameter νto be greater than 2, thus the variance of the underlying random variable is finite [62]. 3.2. Hidden Markov models and Expectation-Maximization algorithm For a shorter notation, in this subsection, we will denote the process of the observations as X={Xt}t=1,...,N.165 Hidden Markov models are the models in which to describe the evolution of the observations process X, an additional unobserved process (hidden states) z={zt}t=1,...,Nis introduced so that the whole system has the Markov property. As the hidden states are unknown, the full likelihood of a system P(X,z|Θ) cannot be calculated directly to obtain the parameters with the MLE method. The standard iterative method that can be used instead is called ExpectationMaximization (EM) algorithm. First, some initial values of Θ(0) are set. Next, the procedure of EM is based on170 repeated steps of calculating the expectation of the likelihood given the distribution of zconditioned with Xand Θ = Θ(u)and reassigning Θto new values that maximize the conditional expectation. The u-th step can be described with the updating formula: Θ(u+1) =argmax Θ Ez|X,Θ=Θ(u)P(X,z)|Θ,(11) where the letter E stands for the conditional expectation. We consider two different models from this class. However, they share the same equation for the evolution of Xgiven z, which is a polynomial regression (of degree p):175 Xt= p X i=0 βi,zkti+σztεt,(12) where t=1, . . . N, the constants β0,zkto βp,zkare polynomial coefficients in the state (degradation stage) zk, and εtis Gaussian white noise. The set of p+1 polynomial coefficients depends on the current stage zt. Hidden states may be interpreted as the degradation stages indexed from 1 to K in the proper order from the healthy state to the complete degradation (so Kdenotes the number of distinct degradation stages). In this paper according to the framework described in Section 2 we put K=3.180 In the HMM approach, CP1 and CP2 are not included in Θ. The result of fitting the HMM model to the data contains the probabilities for each tof the current stage ztbeing equal to 1,...,K. The state with the highest probability indicates in which degradation state the observed machine is at time t, and thus we obtain the division of the degradation process into three stages. 7 3.2.1. Hidden Markov model regression (HMMR)185 This approach employs a mixture of polynomial regressions handled by hidden states of a discrete-time Markov chain. Changing of the stages is therefore fully described by the initial distribution and the one-step transition matrix. A significant restrictions are therefore put to the transition probabilities to better reflect the nature of the process, thus for all l=1,...,Kwe have: Pzt+1=k|zt=l=0,for k<{l,l+1}(13) and generally nonzero otherwise. The parameters are estimated through the Baum-Welch algorithm as described in190 [57]. 3.2.2. Hidden logistic process (HLP) In this method, the process of the unobserved switching stages that affects the set of polynomial regression coefficients is modeled with a multinomial logistic regression model. The distribution of the current stage ztin this model is assumed to be as follows:195 πt,k(β)=P(zt=k|β)=exp(Pp i=0βkiti) PK l=1exp(Pp i=0βliti),for k=1,...,K,(14) where β=[βki]k=1,...,K,i=0,...,pis the matrix containing p+1 regression parameters for all Kdegradation stages. As zis not observed, the total log-likelihood of the model must include the Gaussian distribution of Xtgiven zt. In terms of probability density functions, we have (see [57]): p(X|Θ)= N X t=1 log K X k=1 πt,k(β)pΘ(Xt|zt=k),(15) where the Gaussian PDF denoted with pΘ(Xt|zt=k) comes from Eq. (12). The resulting model is a special case of the time-heterogeneous Gaussian mixture model. A dedicated EM for this approach is described in [57]. As is shown200 there, in a single iteration of EM, the expectation term which is a function of Θ(compare Eq. (11)) can be separated into two terms, one dependent on βand the other dependent on σand β. Thus, fitting of these two sets of parameters can be done independently. The matrix βis estimated utilizing a multi-class iterative reweighted least squares (IRLS) approach. 4. Analysis of simulated data205 Based on the assumption of Section 2, the long-term data model is used to generate HI observations. After the mentioned methodologies are employed to segment the data into three stages in the presence of Gaussian noise and non-Gaussian noise. Additionally, to show the performance of the methodology, the results were compared together. 4.1. Model description Considering the model discussed in Section 2, we propose the following model for generating HI(t) that will be210 used in simulation study: HI(t)=R(t)+D(t),(16) where R(t) and D(t) are, respectively, random and deterministic components. Both these parts consist of three stages, denoted as stage 1, stage 2 and stage 3, which are related to three considered underlying stages (healthy stage/degradation stage/critical stage) and determine the behavior of the process with regard to both trend and noise’s scale (variance in Gaussian distribution). Let us assume that we have a sample signal HI(1),··· ,HI(N). The215 changing point between stages 1 and 2 is denoted as τ1, and the changing point between stages 2 and 3 by τ2, where 1 < τ1< τ2<N. In other words, we can divide the signal segment by segment, so that the sequence HI(1),··· ,HI(τ1) corresponds to stage 1, the sequence HI(τ1+1),··· ,HI(τ2) corresponds to stage 2 and the sequence HI(τ2+1),··· ,HI(N) corresponds to stage 3. 8 In the simulation study presented in the next section, we assume two distributions for the random component,220 namely Gaussian and Student’s t. The aim of using both Gaussian and Student’s t (being an example of a heavytailed distribution) in the simulation is to provide a comparison between them for condition monitoring application, as nonGaussian behavior is widely present in real scenarios. The random component corresponding to R(t) is constructed in the following way: R(t)=SC(t)e R(t),(17) where the function S C(t) represents the time-changing scale and e R(t) is a a series of independent identically distributed225 random variables. For the Gaussian distribution, for each twe put e R(t)∼ N(0,1) and for the Student’s t distribution case e R(t)∼t(ν). For simplicity, we assume that the distribution in each stage is the same, however, as it was mentioned in Section 2, in practice, it may be different for different stages. As it was mentioned, its behavior is different for each stage, namely, we assume that in stage 1 the scale grows linearly, from σ1to σ2(where both values are relatively close to each other), then in stage 2 it also increases linearly,230 from σ2to σ3, and finally, in stage 3 it grows exponentially, from σ3to σ4. The function S C(t) is defined as follows: SC(t)=             a1t+b10<t≤τ1, a2t+b2τ1<t≤τ2, a3exp(b3t)τ2<t≤N, (18) where constants a1,b1,a2,b2,a3,b3are derived in such a way that SC(1) =σ1,SC(τ1)=σ2,S C(τ2)=σ3and SC(N)=σ4[48]. The behavior of the deterministic component D(t) in Eq. (16) has different nature for different stages. Let us recall that in the stage 1 it is at the fixed constant level, denoted here as c1. Then, in stages 2 and 3 we consider linear235 and exponential functions, respectively, with the same growth parameters as for the corresponding stages of the scale function SC(t). Moreover, we assume the function D(t) has no discontinuities in the stage changing points τ1and τ2. Under these assumptions, the deterministic term of the signal has the form D(t)=             c10<t≤τ1, a2t+c2τ1<t≤τ2, a3exp(b3t)+c3τ2<t≤N, (19) where c2and c3are derived in such a way that D(t) is a continuous function. In panel (a) in Fig. 4 we present the deterministic component D(t) and the scale function SC(t) for the following values of the parameters: τ1=1000,240 τ2=1600, N=1700, σ1=1, σ2=2, σ3=7, σ4=25 and c1=10. Here, we assumed four additional values of the νparameter, namely ν∈ {2.1,3,5,10}. As can be seen, for the Gaussian distributed signal we do not observe the outlying observations; see panel (b) in Fig. 4 , while for the non-Gaussian heavy-tailed case, large impulses occur in the data, see panel (c) in Fig. 4 . The smaller the ν, the higher impulses may occur in the signal. It should be noted that the model used for the simulation is different from the one used in the segmentation245 methods. A model used for piecewise regression methods (see Eq. (2)) is a simplified version of the simulated model presented in this section neglecting the scaling factor S C(t), i.e. changing it to a constant. The reason of such is to investigate how complex behaviour may be modeled with a simplified approach. 9 Figure 13: FEMTO test rig. frequency of the data set is 25600 Hz, but the length of the signal is 0.1 second, so it is too low for standard frequency analysis and spectrogram. Similarly to the IMS data set, this data set has been employed in many publications for health index construction [70–77], segmentation of degradation process [22, 78–81] and RUL prediction [82–87]. Bearing1 −1 is selected as a case study in this research from the FEMTO data set. This data set includes 2803 sets of recorded vibration data. Each of these vibration sets is recorded for 0.1 seconds with a 25.6 kHz sampling rate (see335 panel (a) in Fig. 14). This process is repeated every 10 seconds. The shaft speed is approximately 1800 rpm and the load is equal to 4000 N. Also, in this study, the RMS of each vibration set is employed as HI, see panel (b) in Fig.14. (a) (b) Figure 14: FEMTO data set case number 1, (a) Raw bearing run-to-failure vibration signals, (b) health index (RMS). 5.3. Wind turbine data set This data set is collected from a wind turbine (see Fig 15). The sensor has been mounted on a high-speed bearing shaft of the 2.2 MW power wind turbine. This data set was acquired over more than 50 days using two different ways.340 In the first approach, the raw vibration signal is acquired for 6 seconds with a sampling frequency of 100000 Hz per day for +50 days (see panel (a) in Fig.16) and for the second version, the inner race energy of the bearing is calculated every 10 min for +50 days (see panel (b) in Fig. 16). For more information on the methodology used to calculate the inner race energy, see [88]. 16 In the end, the inner race fault has occurred, which has been proven by inspection, see Fig 15. The bearing type345 used during the test is 32222 −J2SKF. It should be mentioned that this data set has been used for prognosis by several papers [88–91]. Figure 15: Wind turbine test rigs. (a) (b) Figure 16: Wind turbine data set, (a) raw bearing run-to-failure vibration signals, (b) health index (Inner race energy). 5.4. Results and discussion In this subsection, we presented the results of our proposed methodology for three real data sets (IMS, FEMTO, and wind turbine data set). We apply our robust approach to every data set and compare the results with conventional350 methods. 17 5.4.1. Results for IMS data set The results of the segmentation methods for this case study are presented in panel (a) of Fig. 17. As it can be seen in panel (a) of Fig. 17, most of the methods detected the first changing point (CP1) (the boundary point between the healthy stage and degradation stage) between Time=500 and 560. By visual check of the HI and results, it looks355 like most of the methods performed the proper results. For the second changing point (the boundary point between degradation stage and critical stage trend) plenty of methods such as OLS, HMMR, DPS, IRLS, and HLP method detected this point around Time=700, where indeed HI presents a big increase of value. In comparison, the LAE and ST methods detected Time=800 and 900 as CP2, respectively. By visually checking the trend of HI, it can be concluded that the result of the LAE and ST method is closer to reality for detecting the second changing point CP2.360 However, to confirm our results, the envelope analysis is applied for each vibration set, and the result is sequentially plotted in panel (b) of Fig. 17. For more details about envelope analysis we refer to [92–95]. As can be seen in panel (b) of Fig. 17 after Time=527 harmonic frequency has appeared, which means the bearing went to the degradation stage. By comparison of the envelope analysis and segmentation methods, it can be concluded that ST and HLP methods provide the best result.365 (a) (b) Figure 17: Changing points detection IMS data set case number 2, (a) result of changing points detection method, (b) envelope spectrum of the raw vibration signal. 5.4.2. Results for FEMTO data set The results of the segmentation methods for this case study are presented in Fig.18. As can be seen, this data set perfectly follows the idea of 3 stages: for Time=0 to c.a. 1300 it is nearly flat, then up to Time=2700 it is linear increase and then rapid growth occurs. Some noise, especially in the middle stage, is seen, and some minor outliers can also be noticed. Thus, we consider it to be a trend with nearly Gaussian noise. As shown in Fig. 18 most of370 the algorithms detected the first changing point (boundary point between healthy and degradation stage) between Time=1200 and 1400, except HMM-based algorithms such as HMMR and HLP that discovered the first changing point much later than it occurs in reality. Based on the visual check first changing point should be located around Time=1300. This means that the ST method could detect this point as well. It should be considered that transferring between the healthy and degradation stages is smooth, and it is not easy to detect this point. According to Fig. 18 the375 second changing point should be somewhere around Time=2700. Therefore, it can be concluded that the ST, LAE, OLS, HMMR, and HLP methods could detect this point as well. However, DPS and IRLS methods have not been able to discover this point properly. 18 Figure 18: Result of changing points detection methods for the FEMTO data set. 5.4.3. Results for wind turbine data set To apply the proposed approach to the wind turbine data set, inner race energy is selected as a health index. As380 can be seen in panel (b) of Fig. 16 this data set includes many outliers. Thus, we consider it as a trend with nonGaussian noise. In addition, a few fluctuations have appeared on HI may that have corresponded to the variation of the load or some phenomena like self-healing. This specific behavior in real cases makes a significant challenge to the segmentation and prediction of RUL. In Fig. 19 the results of the segmentation methods for this case study are presented. The CP1 (changing point between the healthy stage and degradation stage) is detected by all methods385 somewhere between 15 March and 21 March; for instance, OLS, ST, LAE, and IRLS discovered this point on 15 March, while HMMR, DPS, and HLP detected this point on 18 March, 18 March and 19 March, respectively. By looking at Fig. 19, 15 March seems a more logical candidate for CP1 because, after this day, the HI is dramatically growing till 21 April, so it means this point is discovered by HMMR, HLP, and DPS with a delay. For the CP2 (the boundary point between degradation stage and critical stage trend), HMMR, DPS, and HLP were390 detected on 8 April, while ST and IRLS were discovered it on 20 April, except the LAE method which recognized this point on 22 March. Another version of this data set is used to validate our results. As discussed before, this data set version is composed of a raw vibration signal recorded for six seconds every day with a 100k sampling rate, see panel (a) in Fig. 20. So it means it has an excellent potential to apply various frequency analyses to do condition monitoring. The envelope analysis is applied for each vibration set, and the result is sequentially plotted in panel (b)395 of Fig. 20. As it can be seen in panel (b) of Fig. 20 after 14 March, harmonic frequency has appeared, which means the bearing went to the degradation stage. Also, after 22 April, the amplitude of harmonic frequency is dramatically increased, which can be considered CP2. By comparing the results of the segmentation method, it can be found that IRLS and ST methods have the best results for this case. 19 Figure 19: Changing points detection of wind turbine data set. (a) (b) Figure 20: Changing points detection wind turbine, (a) raw vibration data, (b) envelope spectrum map. 5.4.4. Discussion400 The error percentage for detecting changing points in all actual data sets is demonstrated in Fig.21. As we can see in both sub-figures for both CP1 and CP2, the ST method has the lowest percentage error. In contrast, the error percentage for the rest of the methods has fluctuated on different data sets, which can include the fact that considering the non-Gaussian noise distribution improves the process of detecting changing points in long-term condition monitoring data.405 20 (a) (b) Figure 21: Percentage error for real data sets changing points detection, (a) percentage error for CP1, (b) percentage error for CP2. 6. Conclusions In the paper, a segmentation problem of long-term degradation (health index) data from SCADA systems has been investigated. The long-term data describes the degradation process. There are several approaches assuming a model of such degradation. Here, we follow the idea of 3 stages with different properties (see Fig. 3). In practical application, it is a really crucial task to detect when a machine is changing its condition from a healthy stage to a degradation410 stage (warning) and degradation stage to a critical stage (alarm), and it is the base of further analysis like prognostics. Also, in many papers, researchers focus on the last segment, but they don’t mention how they have found the division between stages 2 and 3 and whether this detected point coincides with the true start time of the rapid development of the fault in the machine. Therefore, in this paper, we tested various techniques for CPx finding, as CPx is a basis for identifying data segments with different properties. The achievements of this paper can be summarised as follow:415 •A model of HI data was proposed as three segments sequence with non-Gaussian noise (student t-distributed) to describe the degradation process, which can be used to simulate the artificial data set. •Moreover, we showed that the non-Gaussianity of the model makes detection efficiency worse. •We applied selected methods to simulate and analyse signals with Gaussian and non-Gaussian noise. We used Monte Carlo simulation and box plot-based visualization to provide meaningful results. In the case of Gaussian420 noise, most of the mentioned methods have almost been able to detect changing points, although a few methods, such as HLP and ST, had the lowest error. However, the situation is entirely different in the presence of nonGaussian noise. As we expected, the efficacy of methods derived based on Gaussian distribution decreased with the increase in impulsiveness of HI. In contrast, the efficacy of robust methods such as ST, LAE, and IRLS have not been significantly influenced by increasing impulsiveness.425 •We have noticed that detection of CP1 is much more difficult than CP2. Statistical properties evolve in time for the whole life of the machine. The difference between the end of stage 1 and the beginning of stage 2 is not so clear as between stage 2 and stage 3, as we change the trend from linear to exponential. We applied the procedure to real data sets known in the community as benchmark (reference) data sets (namely IMS, FEMTO, and data from wind turbine drive). Additionally, we presented a waterfall plot of envelope spectrum430 calculated for available raw vibration data for IMS and wind turbine data sets. Such a 3D plot demonstrates how the envelope spectrum changes allow validating visually that detected changing points are correct. Also, the results obtained from applying the mentioned methods to real data sets and comparing them with the results of envelope analysis show that the ST method has the most efficiency in detecting changing points. Furthermore, because the investigated data sets are benchmark data sets, many papers use them as references every year. The researchers can435 use the results of the real data set section, particularly 3D plots to compare research results. 21 Also, it is noteworthy that the ST method, developed based on non-Gaussian distribution, had the proper efficiency in detecting changing points in both simulated and real data sets. This can be a point to consider for the researcher working in the field of long-term condition monitoring data analysis in a harsh environment to choose the noise distribution. Although there are other noise distributions to describe impulsive behavior of the degradation process,440 which may be more compatible with this procedure, this issue deserves more research in the future. Acknowledgments The authors (Hamid Shiri) gratefully acknowledge the European Commission for its support of the Marie Sklodowska445 Curie program through the ETN MOIRA project (GA 955681). Project no. POIR.01.01.01-00-0350/21 entitled ”A universal diagnostic and prognostic module for condition monitoring systems of complex mechanical structures operating in the presence of non-Gaussian disturbances and variable operating conditions” co-financed by the European Union from the European Regional Development Fund under the Intelligent Development Program. The project is carried out as part of the competition of the National450 Center for Research and Development no: 1/1.1.1/2021 (Szybka ´ Scie˙ zka) - Agnieszka Wylomanska and Radoslaw Zimroz Conflicts of interest The authors declare no conflict of interest. References455 [1] Y. Lei, N. Li, L. Guo, N. Li, T. Yan, J. Lin, Machinery health prognostics: A systematic review from data acquisition to rul prediction, Mechanical systems and signal processing 104 (2018) 799–834. [2] Y. Lei, Intelligent fault diagnosis and remaining useful life prediction of rotating machinery, Butterworth-Heinemann, 2016. [3] B. A. Weiss, J. Pellegrino, M. Justiniano, A. Raghunathan, Measurement science roadmap for prognostics and health management for smart manufacturing systems, National Institute of Standards and Technology (2016) 100–2.460 [4] F. Moosavi, H. Shiri, J. Wodecki, A. Wyłoma´ nska, R. Zimroz, Application of machine learning tools for long-term diagnostic feature data segmentation, Applied Sciences 12 (13) (2022) 6766. [5] S. Aminikhanghahi, D. J. Cook, A survey of methods for time series change point detection, Knowledge and information systems 51 (2) (2017) 339–367. [6] A. Prakash, N. James, M. Menzies, G. Francis, Structural clustering of volatility regimes for dynamic trading strategies, Applied Mathematical465 Finance (2022) 1–39. [7] S. Das, Blind change point detection and regime segmentation using gaussian process regression, Ph.D. thesis, University of South Carolina (2017). [8] J. Abonyi, B. Feil, S. Nemeth, P. Arva, Fuzzy clustering based segmentation of time-series, in: International Symposium on Intelligent Data Analysis, Springer, 2003, pp. 275–285.470 [9] V. S. Tseng, C.-H. Chen, P.-C. Huang, T.-P. Hong, Cluster-based genetic segmentation of time series with dwt, Pattern Recognition Letters 30 (13) (2009) 1190–1197. [10] A. Sam´ e, F. Chamroukhi, G. Govaert, P. Aknin, Model-based clustering and segmentation of time series with changes in regime, Advances in Data Analysis and Classification 5 (4) (2011) 301–321. [11] E. J. Keogh, M. J. Pazzani, An enhanced representation of time series which allows fast and accurate classification, clustering and relevance475 feedback., in: Kdd, Vol. 98, 1998, pp. 239–243. [12] S. Lee, Y. Jeong, D. Park, B.-J. Yun, K. H. Park, Efficient fiducial point detection of ecg qrs complex based on polygonal approximation, Sensors 18 (12) (2018) 4502. [13] V. S. Tseng, C.-H. Chen, C.-H. Chen, T.-P. Hong, Segmentation of time series by the clustering and genetic algorithms, in: Sixth IEEE International Conference on Data Mining-Workshops (ICDMW’06), IEEE, 2006, pp. 443–447.480 [14] A. Shirvani, K. Arpe, M. Jahandideh, Analysis of trends and change points in meteorological variables over the south of the caspian sea, Theoretical and Applied Climatology 141 (3) (2020) 959–966. [15] K. Gasior, H. Urba´ nska, A. Grzesiek, R. Zimroz, A. Wyłoma´ nska, Identification, decomposition and segmentation of impulsive vibration signals with deterministic components—a sieving screen case study, Sensors (Switzerland) 20 (19) (2020) 1 – 20. [16] D. Kucharczyk, A. Wyłoma´ nska, J. Obuchowski, R. Zimroz, M. Madziarz, Stochastic modelling as a tool for seismic signals segmentation,485 Shock and Vibration 2016 (2016). [17] A. Grzesiek, K. Gasior, A. Wyłoma´ nska, R. Zimroz, Divergence-based segmentation algorithm for heavy-tailed acoustic signals with timevarying characteristics, Sensors 21 (24) (2021). 22 [18] A. Alkaya, ˙ I. Eker, Variance sensitive adaptive threshold-based pca method for fault detection with experimental application, ISA transactions 50 (2) (2011) 287–302.490 [19] O. Fink, E. Zio, U. Weidmann, A classification framework for predicting components’ remaining useful life based on discrete-event diagnostic data, IEEE Transactions on Reliability 64 (3) (2015) 1049–1056. [20] M. Schlechtingen, I. F. Santos, Comparative analysis of neural network and regression based condition monitoring approaches for wind turbine fault detection, Mechanical systems and signal processing 25 (5) (2011) 1849–1875. [21] L. Hu, N. Hu, X. Zhang, F. Gu, M. Gao, Novelty detection methods for online health monitoring and post data analysis of turbopumps,495 Journal of Mechanical Science and Technology 27 (7) (2013) 1933–1942. [22] J. K. Kimotho, C. Sondermann-W¨ olke, T. Meyer, W. Sextro, Machinery prognostic method based on multi-class support vector machines and hybrid differential evolution–particle swarm optimization, Chemical Engineering Transactions 33 (2013). [23] E. Sutrisno, H. Oh, A. S. S. Vasan, M. Pecht, Estimation of remaining useful life of ball bearings using data driven methodologies, in: 2012 ieee conference on prognostics and health management, IEEE, 2012, pp. 1–7.500 [24] Y. Hu, H. Li, X. Liao, E. Song, H. Liu, Z. Chen, A probability evaluation method of early deterioration condition for the critical components of wind turbine generator systems, Mechanical Systems and Signal Processing 76 (2016) 729–741. [25] P. Tamilselvan, Y. Wang, P. Wang, Deep belief network based state classification for structural health diagnosis, in: 2012 IEEE Aerospace Conference, IEEE, 2012, pp. 1–11. [26] J. Liu, F. Lei, C. Pan, D. Hu, H. Zuo, Prediction of remaining useful life of multi-stage aero-engine based on clustering and lstm fusion,505 Reliability Engineering & System Safety 214 (2021) 107807. [27] J. Singh, A. Darpe, S. P. Singh, Bearing remaining useful life estimation using an adaptive data-driven model based on health state change point identification and k-means clustering, Measurement Science and Technology 31 (8) (2020) 085601. [28] W. Mao, J. He, B. Sun, L. Wang, Prediction of bearings remaining useful life across working conditions based on transfer learning and time series clustering, IEEE Access 9 (2021) 135285–135303.510 [29] S. Sharanya, R. Venkataraman, G. Murali, Estimation of remaining useful life of bearings using reduced affinity propagated clustering, Journal of Engineering Science and Technology 16 (5) (2021) 3737–3756. [30] K. Javed, R. Gouriveau, N. Zerhouni, A new multivariate approach for prognostics based on extreme learning machine and fuzzy clustering, IEEE transactions on cybernetics 45 (12) (2015) 2626–2639. [31] Z. Chen, M. Wu, R. Zhao, F. Guretno, R. Yan, X. Li, Machine remaining useful life prediction via an attention-based deep learning approach,515 IEEE Transactions on Industrial Electronics 68 (3) (2020) 2521–2531. [32] A. Giantomassi, F. Ferracuti, A. Benini, G. Ippoliti, S. Longhi, A. Petrucci, Hidden markov model for health estimation and prognosis of turbofan engines, in: International Design Engineering Technical Conferences and Computers and Information in Engineering Conference, Vol. 54808, 2011, pp. 681–689. [33] E. Ramasso, T. Denoeux, Making use of partial knowledge about hidden states in hmms: an approach based on belief functions, IEEE520 Transactions on Fuzzy Systems 22 (2) (2013) 395–405. [34] F. Sloukia, M. El Aroussi, H. Medromi, M. Wahbi, Bearings prognostic using mixture of gaussians hidden markov model and support vector machine, in: 2013 ACS International Conference on Computer Systems and Applications (AICCSA), IEEE, 2013, pp. 1–4. [35] A. Soualhi, H. Razik, G. Clerc, D. D. Doan, Prognosis of bearing failures using hidden markov models and the adaptive neuro-fuzzy inference system, IEEE Transactions on Industrial Electronics 61 (6) (2013) 2864–2874.525 [36] R. B. Chinnam, P. Baruah, Autonomous diagnostics and prognostics in machining processes through competitive learning-driven hmm-based clustering, International Journal of Production Research 47 (23) (2009) 6739–6758. [37] Q. Liu, M. Dong, W. Lv, X. Geng, Y. Li, A novel method using adaptive hidden semi-markov model for multi-sensor monitoring equipment health prognosis, Mechanical Systems and Signal Processing 64 (2015) 217–232. [38] C. K. R. Lim, D. Mba, Switching kalman filter for failure prognostic, Mechanical Systems and Signal Processing 52 (2015) 426–435.530 [39] L. Cui, X. Wang, Y. Xu, H. Jiang, J. Zhou, A novel switching unscented kalman filter method for remaining useful life prediction of rolling bearing, Measurement 135 (2019) 678–684. [40] K. Gong, X. Chen, Influence of non-gaussian wind characteristics on wind turbine extreme response, Engineering structures 59 (2014) 727–744. [41] K. R. Gurley, M. A. Tognarelli, A. Kareem, Analysis and simulation tools for wind engineering, Probabilistic Engineering Mechanics 12 (1)535 (1997) 9–31. [42] A. Kareem, J. Zhao, Analysis of non-gaussian surge response of tension leg platforms under wind loads (1994). [43] J. Hebda-Sobkowicz, R. Zimroz, A. Wyłoma´ nska, J. Antoni, Infogram performance analysis and its enhancement for bearings diagnostics in presence of non-gaussian noise, Mechanical Systems and Signal Processing 170 (2022) 108764. [44] J. Nowicki, J. Hebda-Sobkowicz, R. Zimroz, A. Wyłoma´ nska, Dependency measures for the diagnosis of local faults in application to the540 heavy-tailed vibration signal, Applied Acoustics 178 (2021) 107974. [45] J. Wodecki, A. Michalak, R. Zimroz, Local damage detection based on vibration data analysis in the presence of gaussian and heavy-tailed impulsive noise, Measurement 169 (2021) 108400. [46] J. Hebda-Sobkowicz, J. Nowicki, R. Zimroz, A. Wyłoma´ nska, Alternative measures of dependence for cyclic behaviour identification in the signal with impulsive noise—application to the local damage detection, Electronics 10 (15) (2021) 1863.545 [47] P. Kruczek, R. Zimroz, A. Wyłoma´ nska, How to detect the cyclostationarity in heavy-tailed distributed signals, Signal Processing 172 (2020) 107514. [48] W. Zulawinski, K. Maraj-Zygmat, H. Shiri, A. Wylomanska, R. Zimroz, Framework for stochastic modelling of long-term non-homogeneous data with non-gaussian characteristics for machine condition prognosis, Mechanical Systems and Signal Processing 184 (2023) 109677. [49] D. Zhang, An adaptive procedure for tool life prediction in face milling, Proceedings of the Institution of Mechanical Engineers, Part J:550 Journal of Engineering Tribology 225 (11) (2011) 1130–1136. [50] L. Zhang, Z. Mu, C. Sun, Remaining useful life prediction for lithium-ion batteries based on exponential model and particle filter, IEEE Access 6 (2018) 17729–17740. 23 [51] P. Ma, S. Wang, L. Zhao, M. Pecht, X. Su, Z. Ye, An improved exponential model for predicting the remaining useful life of lithium-ion batteries, in: 2015 IEEE Conference on Prognostics and Health Management (PHM), IEEE, 2015, pp. 1–6.555 [52] C. Pan, Y. Chen, L. Wang, Z. He, Lithium-ion battery remaining useful life prediction based on exponential smoothing and particle filter, Int. J. Electrochem. Sci 14 (9) (2019) 9537–9551. [53] C. Lim, D. Mba, Switching kalman filter for failure prognostic, Mechanical Systems and Signal Processing 52-53 (1) (2015) 426 – 435. [54] L. Cui, X. Wang, H. Wang, J. Ma, Research on remaining useful life prediction of rolling element bearings based on time-varying kalman filter, IEEE Transactions on Instrumentation and Measurement 69 (6) (2019) 2858–2867.560 [55] L. Reuben, D. Mba, Diagnostics and prognostics using switching kalman filters, Structural Health Monitoring 13 (3) (2014) 296 – 306. [56] L. C. K. Reuben, D. Mba, Diagnostics and prognostics using switching kalman filters, Structural Health Monitoring 13 (3) (2014) 296–306. [57] F. Chamroukhi, A. Sam´ e, G. Govaert, P. Aknin, Time series modeling by a regression approach based on a latent process, Neural Networks 22 (5-6) (2009) 593–602. [58] F. Chamroukhi, A. Sam´ e, G. Govaert, P. Aknin, A regression model with a hidden logistic process for feature extraction from time series, in:565 2009 International Joint Conference on Neural Networks, IEEE, 2009, pp. 489–496. [59] P. J. Huber, Robust statistics, in: International encyclopedia of statistical science, Springer, 2011, pp. 1248–1251. [60] J. O. Street, R. J. Carroll, D. Ruppert, A note on computing robust regression estimates via iteratively reweighted least squares, The American Statistician 42 (2) (1988) 152–154. [61] Least absolute deviation estimates in autoregression with infinite variance 16.570 [62] R. Li, S. Nadarajah, A review of student’s t distribution and its generalizations, Empir. Econ. 58 (3) (2020) 1461–1490. [63] J. B. Ali, B. Chebel-Morello, L. Saidi, S. Malinowski, F. Fnaiech, Accurate bearing remaining useful life prediction based on weibull distribution and artificial neural network, Mechanical Systems and Signal Processing 56 (2015) 150–172. [64] S. Saon, T. Hiyama, et al., Predicting remaining useful life of rotating machinery based artificial neural network, Computers & Mathematics with Applications 60 (4) (2010) 1078–1087.575 [65] H. Liao, W. Zhao, H. Guo, Predicting remaining useful life of an individual unit using proportional hazards model and logistic regression model, in: RAMS’06. Annual Reliability and Maintainability Symposium, 2006., IEEE, 2006, pp. 127–132. [66] A. Widodo, B.-S. Yang, Application of relevance vector machine and survival probability to machine degradation assessment, Expert Systems with Applications 38 (3) (2011) 2592–2599. [67] Q. Li, S. Y. Liang, J. Yang, B. Li, Long range dependence prognostics for bearing vibration intensity chaotic time series, Entropy 18 (1)580 (2016) 23. [68] Y. Qian, R. Yan, R. X. Gao, A multi-time scale approach to remaining useful life prediction in rolling bearing, Mechanical Systems and Signal Processing 83 (2017) 549–567. [69] H. Qiu, J. Lee, J. Lin, G. Yu, Wavelet filter-based weak signature detection method and its application on rolling element bearing prognostics, Journal of sound and vibration 289 (4-5) (2006) 1066–1090.585 [70] A. Mosallam, K. Medjaher, N. Zerhouni, Time series trending for condition assessment and prognostics, Journal of manufacturing technology management (2014). [71] T. H. Loutas, D. Roulias, G. Georgoulas, Remaining useful life estimation in rolling bearings utilizing data-driven probabilistic e-support vectors regression, IEEE Transactions on Reliability 62 (4) (2013) 821–832. [72] K. Javed, R. Gouriveau, N. Zerhouni, P. Nectoux, Enabling health monitoring approach based on vibration data for accurate prognostics,590 IEEE Transactions on industrial electronics 62 (1) (2014) 647–656. [73] R. K. Singleton, E. G. Strangas, S. Aviyente, Extended kalman filtering for remaining-useful-life estimation of bearings, IEEE Transactions on Industrial Electronics 62 (3) (2014) 1781–1790. [74] B. Zhang, L. Zhang, J. Xu, Degradation feature selection for remaining useful life prediction of rolling element bearings, Quality and Reliability Engineering International 32 (2) (2016) 547–554.595 [75] S. Hong, Z. Zhou, E. Zio, W. Wang, An adaptive method for health trend prediction of rotating bearings, Digital Signal Processing 35 (2014) 117–123. [76] Y. Lei, N. Li, S. Gontarz, J. Lin, S. Radkowski, J. Dybala, A model-based method for remaining useful life prediction of machinery, IEEE Transactions on reliability 65 (3) (2016) 1314–1326. [77] Y. Nie, J. Wan, Estimation of remaining useful life of bearings using sparse representation method, in: 2015 Prognostics and System Health600 Management Conference (PHM), IEEE, 2015, pp. 1–6. [78] Z. Liu, M. J. Zuo, Y. Qin, Remaining useful life prediction of rolling element bearings based on health state assessment, Proceedings of the Institution of Mechanical Engineers, Part C: Journal of Mechanical Engineering Science 230 (2) (2016) 314–330. [79] D. Zurita, J. A. Carino, M. Delgado, J. A. Ortega, Distributed neuro-fuzzy feature forecasting approach for condition monitoring, in: Proceedings of the 2014 IEEE Emerging Technology and Factory Automation (ETFA), IEEE, 2014, pp. 1–8.605 [80] L. Guo, H. Gao, H. Huang, X. He, S. Li, Multifeatures fusion and nonlinear dimension reduction for intelligent bearing condition monitoring, Shock and Vibration 2016 (2016). [81] X. Jin, Y. Sun, Z. Que, Y. Wang, T. W. Chow, Anomaly detection and fault prognosis for bearings, IEEE Transactions on Instrumentation and Measurement 65 (9) (2016) 2046–2054. [82] H. Li, Y. Wang, Rolling bearing reliability estimation based on logistic regression model, in: 2013 International Conference on Quality,610 Reliability, Risk, Maintenance, and Safety Engineering (QR2MSE), IEEE, 2013, pp. 1730–1733. [83] Z. Huang, Z. Xu, X. Ke, W. Wang, Y. Sun, Remaining useful life prediction for an adaptive skew-wiener process model, Mechanical Systems and Signal Processing 87 (2017) 294–306. [84] Y. Wang, Y. Peng, Y. Zi, X. Jin, K.-L. Tsui, A two-stage data-driven-based prognostic approach for bearing degradation problem, IEEE Transactions on industrial informatics 12 (3) (2016) 924–932.615 [85] Y. Pan, M. J. Er, X. Li, H. Yu, R. Gouriveau, Machine health condition prediction via online dynamic fuzzy neural networks, Engineering Applications of Artificial Intelligence 35 (2014) 105–113. [86] L. Wang, L. Zhang, X.-z. Wang, Reliability estimation and remaining useful lifetime prediction for bearing based on proportional hazard 24 model, Journal of Central South University 22 (12) (2015) 4625–4633. [87] L. Xiao, X. Chen, X. Zhang, M. Liu, A novel approach for bearing remaining useful life estimation under neither failure nor suspension620 histories condition, Journal of Intelligent Manufacturing 28 (8) (2017) 1893–1914. [88] E. Bechhoefer, R. Schlanbusch, Generalized prognostic algorithm implementing kalman smoother. [89] L. Saidi, J. B. Ali, E. Bechhoefer, M. Benbouzid, Wind turbine high-speed shaft bearings health prognosis through a spectral kurtosis-derived indices and svr, Applied Acoustics 120 (2017) 1–8. [90] L. Saidi, E. Bechhoefer, J. B. Ali, M. Benbouzid, Wind turbine high-speed shaft bearing degradation analysis for run-to-failure testing using625 spectral kurtosis, in: 2015 16th International Conference on Sciences and Techniques of Automatic Control and Computer Engineering (STA), IEEE, 2015, pp. 267–272. [91] J. B. Ali, L. Saidi, S. Harrath, E. Bechhoefer, M. Benbouzid, Online automatic diagnosis of wind turbine bearings progressive degradations under real experimental conditions based on unsupervised machine learning, Applied Acoustics 132 (2018) 167–181. [92] R. B. Randall, J. Antoni, S. Chobsaard, The relationship between spectral correlation and envelope analysis in the diagnostics of bearing630 faults and other cyclostationary machine signals, Mechanical Systems and Signal Processing 15 (5) (2001) 945–962. [93] P. Zimroz, H. Shiri, J. Wodecki, Analysis of the vibro-acoustic data from test rig-comparison of acoustic and vibrational methods, in: IOP Conference Series. Earth and Environmental Science, Vol. 942, IOP Publishing, 2021. [94] V. C. Leite, J. G. B. da Silva, G. F. C. Veloso, L. E. B. da Silva, G. Lambert-Torres, E. L. Bonaldi, L. E. d. L. de Oliveira, Detection of localized bearing faults in induction machines by spectral kurtosis and envelope analysis of stator current, IEEE Transactions on Industrial635 Electronics 62 (3) (2014) 1855–1865. [95] H. Shiri, J. Wodecki, Analysis of the sound signal to fault detection of bearings based on variational mode decomposition, in: IOP Conference Series: Earth and Environmental Science, Vol. 942, IOP Publishing, 2021, p. 012020. 25