Article Adaptation of SVM for novelty detection in long-term diagnostic data with Gaussian and non-Gaussian disturbance
Abstract
This is pre print version of paper "Novelty detection for long-term diagnostic data with Gaussian and non-Gaussian disturbances using a support vector machine"
Full text
Article Adaptation of SVM for novelty detection in long-term diagnostic data with Gaussian and non-Gaussian disturbance Forough Moosavi 1, Hamid Shiri 1, Jacek Wodecki 1, Agnieszka Wyłoma´nska 2and Radosław Zimroz 1,* Citation: Moosavi, F.; Shiri, H.; Wodecki, J.; Wyłoma´nska, A.; Zimroz, R. Adaptation of SVM for novelty detection in long-term diagnostic data with Gaussian and non-Gaussian disturbance. Journal Not Specified 2021,1, 0. https://doi.org/ Received: Accepted: Published: Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. Copyright: © 2025 by the authors. Submitted to Journal Not Specified for possible open access publication under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). 1Faculty of Geoengineering, Mining and Geology, Wroclaw University of Science and Technology, Na Grobli 15, 50-421 Wroclaw, Poland; [email protected] (F.M.); [email protected] (H.S.); [email protected] (J.W.) 2Faculty of Pure and Applied Mathematics, Hugo Steinhaus Center, Wroclaw University of Science and Technology, Wyspianskiego 27, 50-370 Wroclaw, Poland; [email protected] *Correspondence: [email protected] Abstract: Machine diagnostics is simply finding the difference between healthy and faulty compo1 nents. To do this we need threshold value. In the case of the unique machine, we do not have a bad 2 condition example. Then we use so-called Novelty Detection. In the paper, we propose Support 3 Vector Machine (SVM) for novelty detection applied to Health Index data. We applied the moving 4 window approach and simple statistical parametrisation of data inside the window. We build 5 the model in multidimensional (mD) space. The estimated hypersphere is a border describing 6 shape of the mD model. Then we proposed an extension of the procedure - if data is inside the 7 border, we can use it to re-training the model. If data are outside the border - this is a novelty. 8 Basically, it means we do not know these data, and it cannot be recognised as a healthy case; thus, 9 it is faulty. We defined the size of the mD hypersphere (for m=2), describing the location of the 10 good-condition data cloud as a potential feature. If the size of the data cloud is growing, it means 11 more dispersion of the data. We showed the efficiency of the method using simulations and two 12 well-known real data sets. It is essential to consider Gaussian and non-Gaussian data sets.13 Keywords: long-term, diagnostic, data modelling, threshold setting, novelty detection, One class 14 classification, statistical analysis, heavy-tailed distribution, machine learning.15 1. Introduction16 Condition Monitoring systems have become more and more popular in the industry. 17 They are used to acquire some informative signals, process them to extract the features 18 describing the machine condition, and compare values of features with established limit 19 values (thresholds) related to the transition from good condition to slow degradation 20 (Warning) and finally to rapid degradation (Alarm). A key issue is to establish these 21 limited values in practical situations. In some cases, these values are provided by 22 manufacturers. Unfortunately, it should be discovered by experts or advanced data23 driven methods in most cases, particularly for unique machines in the mining industry. 24 One class classification (OCC) is a concept that uses machine learning methods 25 when the data training set is related to one class, mostly a "healthy" condition. Then, a 26 trained system is used to detect anomalous data points. In recent years OCC has been 27 used frequently for outlier detection [ 1 – 4 ] and novelty detection [ 5 – 9 ] . For instance, 28 Bartkowiak et al. [ 1 ] used OCC for outlier analysis of nonstationary signals of the 29 gearbox. Riccardo La Grassa et al. [ 9 ] developed one class minimum spanning tree by 30 employing a convolution neural network for novelty detection. Wang Xiaoyou et.al 31 [ 10 ] used unsupervised one-class classification for the assessment of structural health 32 monitoring of the bridge. Also, it should be noted that the outlier and novelty detection 33 differ in concept and application. In outlier detection, the training data set may be 34 composed of normal and anomaly data. The responsibility is to define a boundary 35 between these two classes and apply this boundary to the test data set, which may 36 Version November 17, 2025 submitted to Journal Not Specified https://www.mdpi.com/journal/notspecified
Version November 17, 2025 submitted to Journal Not Specified 2 of 25 include both normal and anomalous data. However, in novelty detection, the training 37 data set dose not containe anomaly data, and the anomaly data appear in the test data 38 sets [ 11 ]. In general, the classifier’s model could be categorized into three leading groups: 39 density-based, reconstruction-based, and boundary-based.40 Density-based one-class classification approach is developed based on estimating 41 the training data density and comparing it with a threshold. This family of OCC has the 42 most efficiency when a high sample of training data is available. The Gaussian method 43 and the mixture of Gaussians are known as density-based methods. The boundary44 based method tries optimizing the boundary as a modeling challenge to find a closed 45 boundary around the training data. Any sample located outside this boundary is known 46 as an outlier or anomaly. Compared to the density-based approach, the boundary47 based method needs fewer data samples to achieve the same performance. One of the 48 popular boundary-based methods is the one-class support vector machine (OCC-SVM) 49 which is used regularly in the literature. The reconstruction-based method is developed 50 based on prior knowledge about a particular domain of historical data to generate a 51 model. Outlier samples would generally not concede with the historical data assumption 52 embedded in the models, and therefore, any sample with a high reconstruction error 53 is assumed to be an outlier. This approach represents an input pattern as an output 54 and tries to minimize the reconstruction error. Principal component analysis (PCA) [ 12 ], 55 multi-layer perceptron (MLP) [ 13 ], and K-means clustering-based one-class classifier 56 [14], are reconstruction-based models.57 The paper proposes an application of the SVM-based one-class classification ap58 proach. Next, some procedure adaptation is delivered to include new data in the training 59 process. A new feature for decision-making based on OCC-SVM is proposed as the final 60 result.61 1.1. Lifetime curve model62 The lifetime curve shape follows the model known in the literature [ 15 , 16 ]. It 63 consists of 3 regimes: healthy condition, slow degradation (degradation stage), and 64 rapid degradation (critical stage). The data could be modeled as a mixture of trend 65 and random components. The random component usually is Gaussian noise see Fig. 66 1; however, in some cases, the raw data definitely contain random components with 67 non-Gaussian distribution, see Fig. 4.68 Figure 1. Long-term data variation model used in this paper - Gaussian case [15] 2. Methodology69 2.1. Lifetime curve statistical parametrization70 say something about other HIs71 72
Version November 17, 2025 submitted to Journal Not Specified 3 of 25 equations for these statistics73 74 2.1.1. Statistic features75 The raw features obtained from the monitored system can be described in a multi76 dimensional way by descriptive statistics. Such a representation allows the highlighting 77 of some properties of data. Several statistical parameters are calculated using a moving 78 window to preselect a small portion of data for each segment. As a result, one may 79 obtain a data segmentation matrix (DSM) with dimension N segments x M features, see 80 Fig. 281 Figure 2. Long-term data variation parametrization used in this paper 2.2. OCC using SVM82 As discussed in the introduction, OCC-SVM is one of the well-known and influen83 tial OCC techniques developed to address the OCC problem in the literature. In this 84 approach, based on training data, the SVM attempted to discover the hypersphere of a 85 single class and assume all samples outside of this hypersphere are anomalies. Fig. 3, is 86 demonstrated the hypersphere constructed by OCC-SVM to learn the ability to classify 87 out-of-training distribution data based on the hypersphere.88 Figure 3. Schematic of hypersphere constructed by OCC-SVM (green, red, orange points are represented healthy, faulty (anomaly) and support vector points respectively) In Eq. (1) , the mathematical expression of hypersphere with center c and radius r is 89 expressed.90 91
Version November 17, 2025 submitted to Journal Not Specified 4 of 25 min r,cr2,(1) subjected to: ∥ϕ(xi)−c∥2≤r2∀i=1, 2, ..., n.(2) where the function ϕ is the hypersphere transformation of x samples. Eq. (1) at92 tempts to minimize the radius of the hypersphere. However, the mentioned expression 93 is sensitive to the outlier; another flexible representation to tolerate outliers of the hyper94 sphere is given by Eq.(3).95 min r,c,ζr2+1 νn n ∑ i=1 ζi,(3) subject to:96 ∥ϕ(xi)−c∥2≤r2+ζi,∀i=1, 2, ..., n.(4) OCC-SVM can be used for both kinds of anomaly detection applications, i.e., outlier 97 detection and novelty detection.98 2.3. Adaptive OCC using SVM99 For the adaptation of this procedure, we measure the surface of the mentioned 100 hypersphere by receiving the new data and save it theoretically; when data is coming 101 from the healthy stage, this surface should be more or less at the same level; in other 102 hands, when data is coming from the faulty area this surface dramatically increased. 103 Therefore, changing points between regimes can be found by defining a threshold or by 104 visual inspection.105 2.4. Robust variance106 What is robust variance?107 3. Simulation and Results108 In this section, we apply the proposed procedure to simulated data. There are two cases: deterministic trend describing the degradation process mixed with (a) Gaussian, and (b) non-Gaussian noise. The model is described as follows: HI(t) = a b·t c·exp(d·t) t≤1000 1000 <t≤1600 t>1600 + σ1·N(t) σ2·t·N(t) σ3·exp(t)·N(t) t≤1000 1000 <t≤1600 t>1600 , (5) where a , b , c , d are constant parameters that are related to deterministic parts, and 109 σ1 , σ2 , σ3 are parameters responsible for the scales of noise parts. Moreover, {N(t)}110 is a noise. Internal and external Gaussian noise is present in every real system due 111 to various reasons. Unfortunately, the case becomes much more complicated when 112 additional non-Gaussian behavior is present in the signal. Such a model of the signal 113 is inspired by real long-term data we collected from various machines. The source of 114 non-Gaussian noise may be related to machine design, the process related to machine 115 operation, electromagnetic interference, etc.116 3.1. αstable distribution as a model of Non-Gaussian noise117 For non-Gaussian case, the symmetric α -stable distribution has been selected as an 118 example of heavy-tailed, non-Gaussian distribution [17–21].119 The α -stable distribution is defined by its characteristic function and it is charac120 terised by four parameters: α (stability), β (skewness), σ (scale) and µ (location). How121
Version November 17, 2025 submitted to Journal Not Specified 5 of 25 ever, for a symmetric case, it is assumed β=µ= 0, and the corresponding characteristic 122 function takes the form123 E[eitX] = exp{−σα|t|α}. (6) The α is known as the stability index and takes values in the interval ( 0, 2 ] . It should be 124 noted that the α -stable distribution reduces to the Gaussian distribution when α= 2. In 125 case of decreasing α value, the distribution is significantly different from the Gaussian 126 distribution [ 22 ]. In the presented simulation study, we assume the stability index is 127 equal to 1.8.128 Figure 4. Long-term data variation model used in this paper - non-Gaussian case 3.2. Results for simulated signal129 In this subsection, we tried applying our proposed methodology to simulated data 130 based on the mentioned reference.131 3.2.1. OCC and AOCC in presence of Gaussian noise132 In the following, the degradation curve is simulated based on Section 5, and it is 133 shown in panel (a) in Fig. 5in the presence of Gaussian noise. Also, the mean and 134 variance of degradation curves are calculated and illustrated in panels (b) and (c) in Fig. 135 5as selected features for the input of OCC, respectively. It should be noted that the mean 136 and variance are calculated for a window length of 20 data points without overlapping. 137
Version November 17, 2025 submitted to Journal Not Specified 6 of 25 Figure 5. Health Index (HI) in the presence of Gaussian noise, (a) simulated degradation curves or health index (HI), (b) Variance of HI, (c) Mean of HI. In Fig. 6, the results of OCC and AOCC are presented. Panel (a) is shown the 138 health index with blue color and data used for extracting features for training OCC 139 with red color. In conventional OCC approaches, the features extracted from the red 140 part are employed to define the border for novelty detection. The boundary detected 141 based on panel (a) is presented in panel (c). Also, based on the fact more data has more 142 information, we added extra data for training see panel (b); by coming the new data, 143 the OCC retrained again with this data, and this procedure is continued adaptively 144 till reached the critical point that machine is going to the faulty area. By comparing 145 panel (c) and panel (d), it can be seen the detected surface SVM-AOCC boundary is 146 increased rather than classic SVM-OCC see panel (d) (SVM-AOCC detects solid black 147 line, and SVM-OCC detects black dash line in panel (c) that transferred in panel (d) for 148 comparison). This increased surface of the SVM-OCC boundary that results from the 149 adaptive procedure has a few benefits, such as reduced blunder alarm.150
Version November 17, 2025 submitted to Journal Not Specified 7 of 25 Figure 6. The results of the proposed methodology in the presence of Gaussian noise. (a) Health index (HI) and the small data for the input of SVM-OCC, (b) Health index (HI) and the significant amount of data for the input of SVM-OCC, (c) the results of SVM-OCC based on the panel (a) red plus sign are data that used for training the blue circles are the rest of data, and the big black circle is the hypersphere detected by SVM-OCC, (d) the results of SVM-OCC based on the panel (b) red plus sign are data that used for training the blue circles are the rest of data, and the big black circle is the hypersphere detected by SVM-OCC and the black dash the hypersphere detected by SVM-OCC based on panel (a). In Fig. 7, the surface of the boundary detected by SVM-OCC in each step is demon151 strated to detect the critical point. As seen in this case which the observed noise has 152 Gaussian characteristics, the surface of the boundary detected is approximately constant 153 till window 56; after that, the surface of the boundary detected dramatically increased, 154 which can be a promising sign of changing the status of the machine to the faulty area. 155 By comparing the real changing point that happened in window’s number=50 with the 156 detected point by the proposed approach in window’s number=56, it can conclude that 157 the proposed method has a delay in the detected changing point that can be the result 158 of a slight change from the healthy area to the faulty area. Furthermore, in Fig. 7can 159 be seen another jump in the surface value of boundary around window’s number=80 160 related to the changing degradation regime to the critical regime.161
Version November 17, 2025 submitted to Journal Not Specified 8 of 25 Figure 7. Changing the surface of boundary detected by SVM-OCC. To confirm the performance of the proposed method, we have repeated this proce162 dure for to 50 simulated data set and the results are presented in Fig. 8as box plot. As it 163 can be seen based on this fact the scale of the noise is constant all results are the same.164 Figure 8. Box plot for 50 simulated cases in the present of Gaussian noise 3.2.2. OCC and AOCC in the presence of non-Gaussian noise165 In the following, the degradation process is simulated when the observed noise 166 includes non-Gaussian characteristics. For a non-Gaussian case, the symmetric α -stable 167 distribution is selected as an example of heavy-tailed, non-Gaussian distribution [ 17 – 21 ]. 168 It is assumed the α=1.8.169 Panel (a) in Fig. 9is shown the simulated degradation curve (health index(HI)) 170 in the presence of non-Gaussian noise. Furthermore, the mean, variance, median and 171 robust variance of degradation curves are calculated and illustrated in panels (b), (c), 172 (d) and (e) in Fig. 9, respectively. As seen in panels (b) and panel (c), due to noise with 173
Version November 17, 2025 submitted to Journal Not Specified 9 of 25 non-Gaussian characteristics in raw signal, a few spikes have appeared in features space, 174 making a challenge for classic methods that developed based on Gaussian features. 175 However, these spike has not any particular effect on the robust features in panels (d) 176 and (e). It should be noted that the features are calculated for a window length of 20 177 data points without overlapping.178 Figure 9. Health Index (HI) in the presence of non-Gaussian noise with α= 1.8, (a) HI and CP, (b) variance of HI, (c) mean of HI, (d) robust variance of HI, (e) median of HI. In Fig. 10, the results of OCC and AOCC are presented. Panel (a) is shown the 179 health index with blue color and data used for extracting features for training OCC with 180 red color. The features extracted from the red part are used to identify the border for 181 OCC. The boundary detected based on panel (a) is presented in panel (c). Furthermore, 182 we added extra data for training see panel (b); by coming up with the new data, the 183 OCC retrained again with this data, and this procedure is continued adaptively till it 184 reached the critical point at which the machine went to the faulty area. By comparing 185 panel (c) and panel (d), it can be seen that the detected surface SVM-AOCC boundary is 186 increased rather than classic SVM-OCC see panel (d) (SVM-AOCC detects solid black 187 line, and SVM-OCC detects black dash line in panel (c) that transferred in panel (d) 188 for comparison). As discussed, this increased surface of the SVM-OCC boundary that 189 results from the adaptive procedure can help to reduce blunder alarms.190
Version November 17, 2025 submitted to Journal Not Specified 16 of 25 Figure 18. Changing the surface of boundary detected by SVM-OCC for FEMTO data set. In Fig. 19, the surface of the boundary detected by SVM-OCC, which used robust 273 features (median and robust variance) in each step, is demonstrated to detect the critical 274 point. As seen in this case, the surface of the boundary detected is approximately 275 constant till window 72; after that, the surface of the boundary detected smoothly 276 increased till the window’s number=79; after the window’s number=79 can be seen, a 277 dramatically increased in the value of the surface area, which can a promising sign of 278 changing the status of the machine to the faulty area. By comparing the detected point by 279 the proposed approach in window’s number=72 with other papers [ 26 ], it can conclude 280 that the proposed method has detected changing point properly. Furthermore, in Fig. 281 19, another jump in the surface value of boundary around the window’s number=138 282 related to the changing degradation regime to the critical regime can be seen. Also, it 283 can be seen the results of using robust features (median and mean) more or less are the284 same as the results of employing non-robust features (mean and variance) for this case 285 study on the FEMTO data set.286
Version November 17, 2025 submitted to Journal Not Specified 17 of 25 Figure 19. Changing the surface of boundary detected by SVM-OCC for FEMTO data set by using robust features (median and robust variance). 4.4. Results for wind turbine287 In this subsection, we applied the proposed method to one of the wind turbine 288 data set case studies. In this paper, the bearing is set as a case study. The RMS of each 289 vibration set is used as the health index (HI); see panel (a) in Fig. 20. Likewise, the mean, 290 variance, robust variance and median of HI are calculated and shown in panels (b), (c), 291 (d), and (e) in Fig. 20, respectively. It should be noted that the mean and variance are 292 calculated for a window length of 20 data points without overlapping.293
Version November 17, 2025 submitted to Journal Not Specified 18 of 25 Figure 20. Health Index (HI) and features for wind turbine data set, (a) health index (HI) and CP, (b) variance of HI, (c) mean of HI, (d) robust variance of HI, (e) median of HI. In Fig. 21, the results of OCC and AOCC are presented. Panel (a) is shown the 294 health index with blue color and data used for extracting features for training OCC with 295 red color. In conventional OCC approaches, the features extracted from the red part are 296 employed to define the border for OCC. The boundary detected based on panel (a) is 297 presented in panel (c). Also, we added extra data for training see panel (b); by coming 298 up with the new data, the OCC retrained again with this data, and this procedure is 299 continued adaptively till it reached the critical point that the machine is going to the 300 faulty area. By comparing panel (c) and panel (d), it can be seen that the detected 301 surface SVM-AOCC boundary is increased rather than classic SVM-OCC see panel (d) 302 (SVM-AOCC detects solid black line, and SVM-OCC detects black dash line in panel (c) 303 that transferred in panel (d) for comparison). This increased surface of the SVM-OCC 304 boundary that results from the adaptive procedure has a few benefits, such as reduced 305 blunder alarm.306
Version November 17, 2025 submitted to Journal Not Specified 19 of 25 Figure 21. The results of the proposed methodology for wind turbine data set. (a) Health index (HI) and the small data for the input of SVM-OCC, (b) Health index (HI) and the significant amount of data for the input of SVM-OCC, (c) the results of SVM-OCC based on the panel (a) red plus sign are data that used for training the blue circles are the rest of data, and the big black circle is the hypersphere detected by SVM-OCC, (d) the results of SVM-OCC based on the panel (b) red plus sign are data that used for training the blue circles are the rest of data, and the big black circle is the hypersphere detected by SVM-OCC and the black dash the hypersphere detected by SVM-OCC based on panel (a). In Fig. 22, the surface of the boundary detected by SVM-OCC in each step is demon307 strated to detect the critical point. As seen in this case, the surface of the boundary 308 detected is approximately constant till window number 107; after that, the value of the 309 surface area is dramatically increased, which can be a promising sign of changing the 310 status of the machine to the faulty area. By comparing the detected point by the pro311 posed approach in Compare with which paper? RZ: Any reference is needed window’s 312 number=100 with other papers, it can conclude that the proposed method has a delay in 313 the detected changing point that can be the result of a slight change from the healthy 314 area to the faulty area.315
Version November 17, 2025 submitted to Journal Not Specified 20 of 25 Figure 22. Changing the surface of boundary detected by SVM-OCC for wind data set In Fig. 23, the surface of the boundary detected by SVM-OCC, which used robust 316 features (median and robust variance) in each step, is demonstrated to detect the critical 317 point. As seen in this case, the surface of the boundary detected is approximately 318 constant till window 108; after that, the surface area value is dramatically increased, 319 which can be a promising sign of changing the status of the machine to the faulty area. By 320 comparing the detected point by the proposed approach in Compare with which paper? 321 window’s number=117 with other papers, it can conclude that the proposed method 322 has a delay in the detected changing point that can result from a slight change from 323 the healthy area to the faulty area. Like, the previous results of using robust features 324 (median and mean) more or less are the same as the results of employing non-robust 325 features (mean and variance) for this case study on the wind turbine data set.326
Version November 17, 2025 submitted to Journal Not Specified 21 of 25 Figure 23. Changing the surface of boundary detected by SVM-OCC for wind turbine data set by using robust features(median and robust variance) 5. Discussion327 From the simulation data sets section, we demonstrated the effect of the non328 Gaussian noise on the performance of SVM-OCC when it uses a classical and robust 329 statistic. Based on the results of Section 3.2, we saw that when we have noise with non-330 Gaussian characteristics, using classical statistics such as mean and variance sensitive to 331 outliers for training the SVM-OCC led us to misestimate the border of SVM-OCC, which 332 is made early or late alarm for the condition monitoring system. In contrast, we have 333 illustrated that this problem does not happen if robust statistics such as median and 334 robust variance are employed to train SVM-OCC, and we can get more reliable results.335 By looking at the results performed for the FEMTO data set by using the classical statistic 336 (mean and variance) and robust statistics (median and robust variance) can be concluded 337 the results are close together, which come from the fact the noise observed from FEMTO 338 data set is more relative to the noise with Gaussian characteristics. This fact is proved in 339 our last paper [ 15 ]. Also, by comparing the results with the actual changing point from 340 this data set in literature [ 26 ], we can see that both methods detected this point properly. 341 It should be noted that we are aware that OCC is always worse than supervised training. 342 Furthermore, the results performed for wind turbines using classical and robust statistics 343 are closer together. Nevertheless, in our previous study, we proved the observed noise is 344 closer to the non-Gaussian noise but not so far from Gaussian noise [ 15 ]. The comparison 345 results are presented in the table.1346 Table 1: Comparison the real data sets results Data set detected CP with classical statistic detected CP with robust statistic reference FEMTO 1420-1440 1440-1460 1429 [26] Wind turbine 2140-20160 2140-20160
Version November 17, 2025 submitted to Journal Not Specified 22 of 25 6. Conclusions347 The paper proposes introducing an adaptive technique for long-term diagnosis 348 based on one class classification (OCC) and also clarifies the effect of noise with non349 Gaussian characteristics on OCC application. One of the challenging subjects in the 350 OCC application is selecting enough data for the training data; if we do not use enough 351 of amount data that will not be updated, the model will not be trained properly, so it 352 may cause an early and false alarm. Unfortunately, initially, we have to assume a small 353 portion of data for training to be sure that will be a data set related to good condition. 354 Also, because most of the machines are working in harsh areas such as mining, we can be 355 seen noise with non-Gaussian characteristics, which can make challenges for traditional 356 classic approaches that assume the noise has Gaussian distribution; we have tried to 357 clarify the effect on the performance of the OCC application. To do this work based 358 on reference [ 15 ], it assumed the degradation model follows 3 stage regime (Healthy, 359 degradation, and critical regime), and it has simulated the case with Gaussian and non360 Gausian noise. Also, we propose the statistical parameterization of the data performed361 locally, i.e., using short moving windows ( the mean and variance are selected as input of 362 OCC), and then SVM-OCC and the proposed adaptive SVM-OCC technique are applied. 363 It can be seen in the case of Gaussian noise when we used the adaptive version, which 364 means we used a large amount of data with training SVM-OCC; the surface of the 365 hypersphere detected as the border of healthy and faulty was increased that it can help 366 to a reduced early alarm. However, the situation in the case of non-Gaussian noise is 367 entirely different; for SVM-OCC ultimately depends on whether the is data selected for 368 training includes the outlier or not; if yes, it can make underestimate the fault, and for 369 the adaptive SVM-OCC, the surface of hypersphere detected as the border of healthy 370 and faulty has fluctuated during the procedure; its difficult to estimate the changing 371 point correctly that can be coming to this fact, in the case of non-Gaussian noise, the 372 idea of moving windows needs to be revised because one outlier will deliver several 373 features with bias. Also, simple statistics that represent data with Gaussian noise are 374 not optimal for non-Gaussian data. For instance, some features like mean and variance 375 are sensitive to the outlier; in addition, for such a heavy-tailed process, the variance is 376 not defined. This means, we need to use robust features like median instead of mean 377 and robust variance instead of classical one which can handle the effect of non-Gaussian 378 noise. Finally, we applied the proposed method to two real datasets named FEMTO and 379 wind turbine to evaluate the method’s performance.Because the observed noise of these 380 two data sets is close to the Gaussian, the result of using robust features (median and 381 robust variance) and non-robust features are more or less the same and could detect the 382 changing points by acceptable error.383 Author Contributions:384 Data Availability Statement385 Archived data sets cannot be accessed publicly according to the NDA agreement 386 signed by the authors.387 Acknowledgments388 The authors (Hamid Shiri) gratefully acknowledge the European Commission for 389 its support of the Marie Sklodowska Curie program through the ETN MOIRA project 390 (GA 955681).391 Project no. POIR.01.01.01-00-0350/21 entitled "A universal diagnostic and prognostic 392 module for condition monitoring systems of complex mechanical structures operating 393 in the presence of non-Gaussian disturbances and variable operating conditions" co394 financed by the European Union from the European Regional Development Fund under 395 the Intelligent Development Program. The project is carried out as part of the competition 396 of the National Center for Research and Development no: 1/1.1.1/2021 (Szybka ´ Scie˙ zka) 397
Version November 17, 2025 submitted to Journal Not Specified 23 of 25 -Forough Moosavi, Agnieszka Wylomanska, and Radoslaw Zimroz.398 Jacek Wodecki Supported by the Foundation for Polish Science (FNP).399 Conflicts of interest400 The authors declare no conflict of interest.401 402 1. Bartkowiak, A.; Zimroz, R. Outliers analysis and one class classification approach for 403 planetary gearbox diagnosis. Journal of Physics: Conference Series. IOP Publishing, 2011, 404 Vol. 305, p. 012031.405 2. Janssens, J.H.; Flesch, I.; Postma, E.O. Outlier detection with one-class classifiers from ML 406 and KDD. 2009 International Conference on Machine Learning and Applications. IEEE, 2009, 407 pp. 147–153.408 3. Li, H.; Qiu, H.; Sun, S.; Chang, J.; Tu, W. Credit scoring by one-class classification driven 409 dynamical ensemble learning. Journal of the Operational Research Society 2022,73, 181–190.410 4. Dreiseitl, S.; Osl, M.; Scheibböck, C.; Binder, M. Outlier detection with one-class SVMs: an 411 application to melanoma prognosis. AMIA annual symposium proceedings. American 412 Medical Informatics Association, 2010, Vol. 2010, p. 172.413 5. Sabokrou, M.; Khalooei, M.; Fathy, M.; Adeli, E. Adversarially learned one-class classifier 414 for novelty detection. Proceedings of the IEEE conference on computer vision and pattern 415 recognition, 2018, pp. 3379–3388.416 6. Oosterlinck, D.; Benoit, D.F.; Baecke, P. From one-class to two-class classification by incor417 porating expert knowledge: Novelty detection in human behaviour. European Journal of 418 Operational Research 2020,282, 1011–1024.419 7. Pimentel, M.A.; Clifton, D.A.; Clifton, L.; Tarassenko, L. A review of novelty detection. Signal 420 processing 2014,99, 215–249.421 8. Perera, P.; Nallapati, R.; Xiang, B. Ocgan: One-class novelty detection using gans with 422 constrained latent representations. Proceedings of the IEEE/CVF Conference on Computer 423 Vision and Pattern Recognition, 2019, pp. 2898–2906.424 9. La Grassa, R.; Gallo, I.; Landro, N. OCmst: One-class novelty detection using convolutional 425 neural network and minimum spanning trees. Pattern Recognition Letters 2022,155, 114–120. 426 10. Wang, X.; Li, L.; Tian, W.; Du, Y.; Hou, R.; Xia, Y. Unsupervised one-class classification for 427 condition assessment of bridge cables using Bayesian factor analysis. Smart Structures and 428 Systems 2022,29, 41–51.429 11. Seliya, N.; Abdollah Zadeh, A.; Khoshgoftaar, T.M. A literature review on one-class classifi430 cation and its potential applications in big data. Journal of Big Data 2021,8, 1–31.431 12. Bishop, C.M.; others. Neural networks for pattern recognition; Oxford university press, 1995.432 13. Japkowicz, N.; Myers, C.; Gluck, M.; others. A novelty detection approach to classification. 433 IJCAI. Citeseer, 1995, Vol. 1, pp. 518–523.434 14. Jiang, M.F.; Tseng, S.S.; Su, C.M. Two-phase clustering process for outliers detection. Pattern 435 recognition letters 2001,22, 691–700.436 15. Framework for stochastic modelling of long-term non-homogeneous data with non-Gaussian 437 characteristics for machine condition prognosis. Mechanical Systems and Signal Processing 438 2023,184, 109677.439 16. Moosavi, F.; Shiri, H.; Wodecki, J.; Wyłoma´nska, A.; Zimroz, R. Application of Machine 440 Learning Tools for Long-Term Diagnostic Feature Data Segmentation. Applied Sciences 2022, 441 12, 6766.442 17. Grzesiek, A.; Gasior, K.; Wyłoma´nska, A.; Zimroz, R. Divergence-based segmentation 443 algorithm for heavy-tailed acoustic signals with time-varying characteristics. Sensors 2021, 444 21.445 18. Hebda-Sobkowicz, J.; Zimroz, R.; Wyłoma´nska, A.; Antoni, J. Infogram performance analysis 446 and its enhancement for bearings diagnostics in presence of non-Gaussian noise. Mechanical 447 Systems and Signal Processing 2022,170, 108764.448 19. Wodecki, J.; Michalak, A.; Wyłoma´nska, A.; Zimroz, R. Influence of non-Gaussian noise on 449 the effectiveness of cyclostationary analysis–Simulations and real data analysis. Measurement 450 2021,171, 108814.451 20. Kruczek, P.; Zimroz, R.; Antoni, J.; Wyłoma´nska, A. Generalized spectral coherence for 452 cyclostationary signals with α -stable distribution. Mechanical Systems and Signal Processing 453 2021,159, 107737.454
Version November 17, 2025 submitted to Journal Not Specified 24 of 25 21. Khinchine, A.Y.; Lévy, P. Sur les lois stables. CR Acad. Sci. Paris 1936,202, 374–376.455 22. Burnecki, K.; Wyłoma´nska, A.; Beletskii, A.; Gonchar, V.; Chechkin, A. Recognition of stable 456 distribution with Lévy index αclose to 2. Physical Review E 2012,85, 056711.457 23. Nectoux, P.; Gouriveau, R.; Medjaher, K.; Ramasso, E.; Chebel-Morello, B.; Zerhouni, N.; 458 Varnier, C. PRONOSTIA: An experimental platform for bearings accelerated degradation 459 tests. IEEE International Conference on Prognostics and Health Management, PHM’12. IEEE 460 Catalog Number: CPF12PHM-CDR, 2012, pp. 1–8.461 24. Liu, Z.; Zuo, M.J.; Qin, Y. Remaining useful life prediction of rolling element bearings based 462 on health state assessment. Proceedings of the Institution of Mechanical Engineers, Part C: Journal 463 of Mechanical Engineering Science 2016,230, 314–330.464 25. Kimotho, J.K.; Sondermann-Wölke, C.; Meyer, T.; Sextro, W. Machinery Prognostic Method 465 Based on Multi-Class Support Vector Machines and Hybrid Differential Evolution–Particle 466 Swarm Optimization. Chemical Engineering Transactions 2013,33.467 26. Zurita, D.; Carino, J.A.; Delgado, M.; Ortega, J.A. Distributed neuro-fuzzy feature forecasting 468 approach for condition monitoring. Proceedings of the 2014 IEEE Emerging Technology and 469 Factory Automation (ETFA). IEEE, 2014, pp. 1–8.470 27. Guo, L.; Gao, H.; Huang, H.; He, X.; Li, S. Multifeatures fusion and nonlinear dimension 471 reduction for intelligent bearing condition monitoring. Shock and Vibration 2016,2016.472 28. Jin, X.; Sun, Y.; Que, Z.; Wang, Y.; Chow, T.W. Anomaly detection and fault prognosis for 473 bearings. IEEE Transactions on Instrumentation and Measurement 2016,65, 2046–2054.474 29. Mosallam, A.; Medjaher, K.; Zerhouni, N. Time series trending for condition assessment and 475 prognostics. Journal of manufacturing technology management 2014.476 30. Loutas, T.H.; Roulias, D.; Georgoulas, G. Remaining useful life estimation in rolling bear477 ings utilizing data-driven probabilistic e-support vectors regression. IEEE Transactions on 478 Reliability 2013,62, 821–832.479 31. Javed, K.; Gouriveau, R.; Zerhouni, N.; Nectoux, P. Enabling health monitoring approach 480 based on vibration data for accurate prognostics. IEEE Transactions on industrial electronics 481 2014,62, 647–656.482 32. Singleton, R.K.; Strangas, E.G.; Aviyente, S. Extended Kalman filtering for remaining-useful483 life estimation of bearings. IEEE Transactions on Industrial Electronics 2014,62, 1781–1790.484 33. Zhang, B.; Zhang, L.; Xu, J. Degradation feature selection for remaining useful life prediction 485 of rolling element bearings. Quality and Reliability Engineering International 2016,32, 547–554. 486 34. Hong, S.; Zhou, Z.; Zio, E.; Wang, W. An adaptive method for health trend prediction of 487 rotating bearings. Digital Signal Processing 2014,35, 117–123.488 35. Lei, Y.; Li, N.; Gontarz, S.; Lin, J.; Radkowski, S.; Dybala, J. A model-based method for 489 remaining useful life prediction of machinery. IEEE Transactions on reliability 2016,65, 1314– 490 1326.491 36. Nie, Y.; Wan, J. Estimation of remaining useful life of bearings using sparse representation 492 method. 2015 Prognostics and System Health Management Conference (PHM). IEEE, 2015, 493 pp. 1–6.494 37. Li, H.; Wang, Y. Rolling bearing reliability estimation based on logistic regression model. 2013 495 International Conference on Quality, Reliability, Risk, Maintenance, and Safety Engineering 496 (QR2MSE). IEEE, 2013, pp. 1730–1733.497 38. Huang, Z.; Xu, Z.; Ke, X.; Wang, W.; Sun, Y. Remaining useful life prediction for an adaptive 498 skew-Wiener process model. Mechanical Systems and Signal Processing 2017,87, 294–306.499 39. Wang, Y.; Peng, Y.; Zi, Y.; Jin, X.; Tsui, K.L. A two-stage data-driven-based prognostic 500 approach for bearing degradation problem. IEEE Transactions on industrial informatics 2016, 501 12, 924–932.502 40. Pan, Y.; Er, M.J.; Li, X.; Yu, H.; Gouriveau, R. Machine health condition prediction via 503 online dynamic fuzzy neural networks. Engineering Applications of Artificial Intelligence 2014, 504 35, 105–113.505 41. Wang, L.; Zhang, L.; Wang, X.z. Reliability estimation and remaining useful lifetime 506 prediction for bearing based on proportional hazard model. Journal of Central South University 507 2015,22, 4625–4633.508 42. Xiao, L.; Chen, X.; Zhang, X.; Liu, M. A novel approach for bearing remaining useful life 509 estimation under neither failure nor suspension histories condition. Journal of Intelligent 510 Manufacturing 2017,28, 1893–1914.511
Version November 17, 2025 submitted to Journal Not Specified 25 of 25 43. Saidi, L.; Ali, J.B.; Bechhoefer, E.; Benbouzid, M. Wind turbine high-speed shaft bearings 512 health prognosis through a spectral Kurtosis-derived indices and SVR. Applied Acoustics 513 2017,120, 1–8.514 44. Saidi, L.; Bechhoefer, E.; Ali, J.B.; Benbouzid, M. Wind turbine high-speed shaft bearing 515 degradation analysis for run-to-failure testing using spectral kurtosis. 2015 16th International 516 Conference on Sciences and Techniques of Automatic Control and Computer Engineering 517 (STA). IEEE, 2015, pp. 267–272.518 45. Ali, J.B.; Saidi, L.; Harrath, S.; Bechhoefer, E.; Benbouzid, M. Online automatic diagnosis of 519 wind turbine bearings progressive degradations under real experimental conditions based 520 on unsupervised machine learning. Applied Acoustics 2018,132, 167–181.521 46. Bechhoefer, E.; Schlanbusch, R. Generalized Prognostic Algorithm Implementing Kalman 522 Smoother. IFAC-PapersOnLine 2015,48, 97–104.523