scieee AI-readable full text Open interactive document viewer

A Hyper-heuristic Inspired Methodology for Failure Prediction in the Context of Industry 4.0

Navajas Guerrero, Adriana

Abstract

Los capítulos 5 y 6 están sujetos a confidencialidad por la autora. 158 p.

Full text

A Hyper-heuristic Inspired Methodology for Failure Prediction in the Context of Industry 4.0 by Adriana Navajas Guerrero Doctor of Philosophy May of 2023 Faculty of Engineering of Bilbao University of the Basque Country, UPV/EHU (cc)2023 ADRIANA NAVAJAS GUERRERO (cc by-nc-nd 4.0) A Hyper-heuristic Inspired Methodology for Failure Prediction in the Context of Industry 4.0 by Adriana Navajas Guerrero Supervisors: PhD. Eva Portillo & PhD. Diana Manjarres Department of Automatic Control and Systems Engineering Doctor of Philosophy 2023 Faculty of Engineering of Bilbao University of the Basque Country, UPV/EHU Author’s right 2023 by Navajas Guerrero, Adriana Acknowledgments No negar´e que me est´a costando saber c´omo empezar con estas l´ıneas. ´ Estas, que ve´ıa tan lejos por tener que escribir y ya han llegado a este libro. Tampoco voy a negar que han sido a˜nos de esfuerzo, trabajo, noches largas, d´ıas eternos, sufrimiento, roces, lloros, desesperaci´on, dudas, risas, vivencias, personas y serendipias. Aun as´ı, sobre todo esta interesante y nutritiva etapa de mi vida ha sido cambio, madurez y vivencias. Cambio en todos los sentidos, tanto en el nuevo rumbo que tom´o mi vida viniendo a Bilbao y tomando la decisi´on de comenzar esta aventura, como en mi forma de pensar, razonar y modificar mis formas de intereactuar con todos los est´ımulos que nos da la vida. Pero sobre todo, a nivel personal he librado una de las mayores batallas de mi vida contra mi peor enemiga, yo misma. Echando la vista atr´as, dir´e que he madurado enormemente y que puedo decir, que me siento orgullosa de lo que soy hoy en d´ıa. Creo que, si bien ha sido una etapa dura, esta etapa la recordar´e como la que me ayud´o a pensar, razonar, tener criterio, querer vivir, ser tolerante, honesta, emp´atica, cari˜nosa . . . Sin duda, esta etapa ha sido el resultado de un esfuerzo personal, pero sobre todo ha sido el fruto del trabajo en equipo y el apoyo incondicional de muchas personas que ya formaban parte de mi vida, que llegaron para quedarse o que se han ido. Es por ello que, sin todas ellas, no estar´ıa aqu´ı sentada escribiendo estas l´ıneas hoy en d´ıa. Por esta raz´on, deseo que todas estas personas sean parte de mi libro, de mi historia y de esta tesis. En primer lugar, agradecer a mis directoras Eva Portillo y Diana Manjarres, por ser apoyo en buenos y malos momentos, por confiar en mi incluso cuando yo dudaba de mi misma, por ser cr´ıticas, por no permitirme tirar la toalla y por animarme cuando estaba cruzando un enorme t´unel negro. Junto a ellas he aprendido muchas cosas en estos a˜nos, tanto en el plano acad´emico como en el personal, que han i hecho que a d´ıa de hoy tambi´en sea como soy. Por todo, GRACIAS. Tambi´en quiero agradecer a Tecnalia por darme la oportunidad de empezar la que sin duda va a ser la etapa m´as relevante en mi vida. Especialmente dar las gracias a todo el equipo de OPTIMA, a mis AIdeanos y a toda la fauna animal´ıstica (cabras, vacas, toritos, koalas, perezosos...) que me ha acompa˜nado en los ´ultimos meses de esta tesis y me han sacado una sonrisa cada d´ıa. Debo hacer una menci´on especial a Sergio Gil, quien siempre ha estado a mi lado apoy´andome, d´andome consejos y siendo parte implicada del camino. Me gustar´ıa extender mi m´as sincero agradecimiento a I˜naki Olabarrieta, quien ha sido de gran ayuda durante el proceso y me ha ense˜nado valiosas lecciones. Tambi´en quisiera agradecer a Isidoro ciri´on por sus charlas inspiradoras y su apoyo. No puedo dejar de mencionar a Iraide. Ella siempre ha estado presente con sus consejos sabios, su apoyo incondicional, sus abrazos reconfortantes, sus mensajes de ´animo, sus risas contagiosas y sus charlas inspiradoras. Su presencia ha sido fundamental en momentos de incertidumbre y dificultades, y su apoyo me ha dado la fuerza y la motivaci´on necesarias para seguir adelante. Desde luego, no puedo olvidarme de mi gente. Agradezco enormemente a mis padres todo lo que tengo, soy y he conseguido desde que tengo uso de raz´on. A ellos, por inculcarme buenos valores, por apoyarme incondicionalmente en todo, por quererme, por ser como son... sencillamente por absolutamente todo, os quiero. Esta tesis no hubiera sido posible sin las buenas amistades, las risas tomando unas cervezas, los viajes, las interminables llamadas de tel´efono, las sesiones de reflexi´on para cambiar el mundo, los atardeceres, los abrazos, en definitiva, sin las personas detr´as de todos esos momentos. . . por eso Gracias a Igor, Idoia, Dani, Mara, Elena, Paula, Iris, Nani, Mar´ıa, Valle, Bego, Pedro, Ibai . . . ii Quiero hacer dos menciones especiales, en primer lugar, a Aritz. A ´el le tengo que agradecer ser un amigo incondicional en estos a˜nos, ser apoyo y ayuda continuos y por compartir m´usicas, entrenamientos, Heardles, aventuras, paseos y amistad. En segundo lugar, a mi compa˜nera y amiga desde el d´ıa en el que me embarqu´e en esta aventura, a mi hortelana favorita, a mi compa˜nera de viaje, a Iratxe. . . GRACIAS. Finalmente, no quer´ıa acabar sin mencionar a dos personas, que no son de toda la vida, pero espero que sean para toda la vida. Dos serendipias que han llegado como un rallo de luz en plena oscuridad, dos lianas en las que en los peores momentos me he podido colgar y han estado ah´ı sin pedirlo, dos personas incre´ıbles que me han regalado momentos maravillosos, a ellos dos, Iraia y Alberto GRACIAS por todo. iii Table 5.6 Results of MRC, F1-score and TM in MV databases by β, where; DB: Database, IQR: Interquartile range, CS: Clustering Solution, TM: Trustworthiness metric, D: Divorce, I: Iris, Io: Ionosphere, W: Wine, C: Cervix, B: Bupa, L: Lymphography. . . . . . . . . . . . . . . . . . . 104 Table 5.7 Characteristics of UCR databases. . . . . . . . . . . 106 Table 5.8 Results of MRC, F1-score and TM in MV databases by β, where; DB: Database, IQR: Interquartile range, CS: Clustering Solution, TM: Trustworthiness metric, S: Sony, A: Arrow, G: Gun, ECG: Electrocardiogram and M: Mote. 107 Table 5.9 Obtained datasets in the case of study. . . . . . . . 112 Table 5.10 Parameter values for PLAHS in the cold stamping case study, where; Nd: Total number of element in dataset, NdL: number of labelled breakage-stops, NdnL: number of unlabelled breakage-stops, Nk: number of know classes. 113 Table 5.11 Results of PLAHS in the cold stamping process in terms of MRC, F1-score and TM and labelling proposal, where; IQR: Interquartile range, CS: Clustering Solution, TM: Trustworthiness metric. . . . . . . . . . . . . . . . . 114 Table 5.12 Time of execution per process signal, where tis the meantimevalue........................ 115 Table 6.1 Comparison of the literature and the proposed hyperheuristic inspired approach for AD. LER: Level of Expertise Requirement, Wz: time-window size, Th: Threshold, C: Calculated, A: Adaptive, Fx: Fixed, O: Optimized, R: Random, CV: Cross Validation, GS: Grid Search, metaOpt *: Comparison between PSO, SHO, KH, SSA, metaOpt**: Comparison between PSO, DE, GA, GWO. . . . . 121 Table 6.2 Features or statistics extracted from TSM, where Fq: Frequency-domain and T: Time-domain, SNR: Signal to Noise Ratio, Std: Standard deviation. . . . . . . . . . . 127 Table 6.3 Definition and formulation of the features used in time-domain, where Wzj: time-window . . . . . . . . . . 134 x Table 6.4 Definition and formulation of the features used in frequency-domain, where Wzj: time-window . . . . . . . 134 Table 6.5 Corresponding Note for each Feature, where F: Frequencydomain and T: Time-domain, SNR: Signal to Noise Ratio, Std: Standard deviation. . . . . . . . . . . . . . . . . . . . 146 Table 6.6 Feasible values for notes where nris the initial random value for note, nnis a new value for note . . . . . . . 148 Table 6.7 Case of study: types of failure, number of TSM (M) and number of SBD (N). . . . . . . . . . . . . . . . . . . . 157 Table 6.8 Optimal HS operator configuration. . . . . . . . . . 161 Table 6.9 Mean AUC MOD value per HS operator set for FailureA.............................. 162 Table 6.10 Mean AUC MOD value per HS operator set for FailureB............................ 164 Table 6.11 Mean AUC MOD value per HS operator set for FailureC............................ 165 Table 6.12 Mean iteration convergence for best set operators’s values.............................. 166 Table 6.13 Mean AUC MOD value per HS operator set for FailureD............................ 167 Table 6.14 Mean AUC MOD value per HS operator set in FailureE. ............................. 168 Table 6.15 Mean AUC MOD value per HS operator set for FailureF. ........................... 170 Table 6.16 Results for failure type A: breakage hole punch. H: Harmony, A-M: AUC MOD, A-R: AUC ROC, SNT: Sensitivity, SPC: Specificity, FPR: False Positive Rate . . 173 Table 6.17 Results for failure type B: breakage hole punch. H: Harmony, A-M: AUC MOD, A-R: AUC ROC, SNT: Sensitivity, SPC: Specificity, FPR: False Positive Rate . . 176 Table 6.18 Results for failure type C: breakage calibration punch. H: Harmony, A-M: AUC MOD, A-R: AUC ROC, SNT: Sensitivity, SPC: Specificity, FPR: False Positive Rate . . 182 xi Table 6.19 Results for failure type D: breakage punch. H: Harmony, A-M: AUC MOD, A-R: AUC ROC, SNT: Sensitivity, SPC: Specificity, FPR: False Positive Rate. . . . . . . 185 Table 6.20 Results for failure type E: wear Calibration Punch. H: Harmony, A-M: AUC MOD, A-R: AUC ROC, SNT: Sensitivity, SPC: Specificity, FPR: False Positive Rate. . . 186 Table 6.21 Results for failure type F: wear nugget evacuation tube. H: Harmony, A-M: AUC MOD, A-R: AUC ROC, SNT: Sensitivity, SPC: Specificity, FPR: False Positive Rate.192 Figure 1.1 Level of digitalization of the reviewed manufacturing companies, source: [2]. . . . . . . . . . . . . . . . . . . 3 Figure 1.2 Degree of Industry 4.0 technologies implementation in SMS companies [3]. . . . . . . . . . . . . . . . . . . . 3 Figure 1.3 Types of data sources in the Manufacture Industry [2]. .............................. 4 Figure 1.4 Type of analysis used in Manufacture Industry [2]. 5 Figure 2.1 State of the art scheme. . . . . . . . . . . . . . . . 13 Figure 2.2 Industrial revolutions timeline. . . . . . . . . . . . 15 Figure 2.3 Enabling technologies of Industry 4.0, source: [4] re-illustrated and modified for this Thesis. . . . . . . . . . 16 Figure 2.4 Improvement actions to the 8 main areas of interest and its corresponding enabling technologies in Industry 4.0. Source :[5] re-illustrated and modified for this Thesis. CPS: Cyber Physical System, RTO: Real-time optimization, ML: ML, AR: Augmented Reality, IoT: Internet of Things, AM: Additive manufacturing. . . . . . . . . . . . 22 Figure 2.5 Categorization and definition of prognosis methods. 26 Figure 2.6 Outlier anomaly example [6]. . . . . . . . . . . . . 28 Figure 2.7 Contextual anomalies example [6]. . . . . . . . . . 28 Figure 2.8 Collective anomalies example [6]. . . . . . . . . . . 29 xii Figure 2.9 Semi-supervised taxonomy for AD. . . . . . . . . . 31 Figure 2.10 Conceptual scheme of the proposed taxonomy for SSLclustering. ........................ 34 Figure 2.11 Taxonomy of hyper-heuristics. . . . . . . . . . . . 37 Figure 2.12 Relevant parameters in the AD context. . . . . . 44 Figure 3.1 Hyper-heuristic inspired methodology applied in thisThesis. .......................... 57 Figure 4.1 Sheet metal cold-stamping process. . . . . . . . . . 67 Figure 4.2 Wire rod cold-stamping process. . . . . . . . . . . 68 Figure 4.3 Examples of sheet metal cold-stamping pieces. . . 69 Figure 4.4 Examples of wire rod cold-stamping pieces. . . . . 69 Figure 4.5 Die holding and riveting pin blocks in cold stamping machine............................. 70 Figure 4.6 Feed rollers and straighteners. . . . . . . . . . . . . 71 Figure 4.7 Bolt cold forming process. . . . . . . . . . . . . . . 72 Figure 4.8 Flow diagram of the cold forming process under study. 73 Figure 4.9 Signals captured from the cold forming process in twostamps........................... 75 Figure 4.10 Process signal 1. . . . . . . . . . . . . . . . . . . . 75 Figure 4.11 Process signal 2. . . . . . . . . . . . . . . . . . . . 76 Figure 4.12 Process signal 3. . . . . . . . . . . . . . . . . . . . 76 Figure 4.13 Process signal 4. . . . . . . . . . . . . . . . . . . . 77 Figure 4.14 Process signal 5. . . . . . . . . . . . . . . . . . . . 77 Figure 4.15 Process signal 6. . . . . . . . . . . . . . . . . . . . 78 Figure 4.16 Flow diagram of datasets creation process. . . . . 78 Figure 5.1 Relation between amount of labels and learning methods. ........................... 82 Figure 5.2 Flow diagram of the proposed PLAHS. . . . . . . . 86 Figure 5.3 Example of the Bagging scheme proposed in PLAHS. 88 Figure 5.4 Example of clustering solutions. . . . . . . . . . . . 96 Figure 5.5 Flow diagram of the PLAHS labelling system. . . . 98 Figure 5.6 TM process evaluation. . . . . . . . . . . . . . . . 100 xiii Figure 5.7 Correlation of TM and F1-score for unsupervised elements in UCI databases. . . . . . . . . . . . . . . . . . 105 Figure 5.8 Correlation of TM and F1-score for unknown elements in UCR databases. . . . . . . . . . . . . . . . . . . 108 Figure 5.9 Separability of databases, where: S: Sony, ECG: Electrocardiogram, A: Arrow, G: Gun, M: Mote, I: Iris, W: Wine, D: Divorce, Io: Ionosphere, B: Bupa, C: Cervix and L: Lymphography. . . . . . . . . . . . . . . . . . . . . 110 Figure 5.10 PLAHS implementation in the real use case. . . . 113 Figure 5.11 Voting solutions for BS8-BS9. . . . . . . . . . . . 115 Figure 6.1 Parameters related with the failure prediction. . . 122 Figure 6.2 Conceptual scheme of the proposed HIMAFP. . . . 126 Figure 6.3 Examples of two possible set of heuristic parameters [Wz, Fe, ThH, ThL] for a specific type of failure in TSM1. The upper figure shows the expected behaviour that detects the anomaly in the previous time-window to the failure (Wz1). The lower figure shows a set of parameters that is not predictor of the type of failure of interest. 136 Figure 6.4 Flow diagram of the proposed HIMAFP for collectiveAD............................. 138 Figure 6.5 Reorganization and clipping procedure. First row shows the input data format before segmentation, second row shows the reorganization per TSM, last row shows how data is clipped and finally adapted for the HIMAFP. . . . 144 Figure 6.6 Threshold parameter explanation. . . . . . . . . . 146 Figure 6.7 Different time-window simulations of  EV (E1 to E6) and its evaluation with AUC MOD and AUC ROC. . 153 Figure 6.8 Evaluation process of AUC MOD vs AUC ROC for casesE1toE6......................... 153 Figure 6.9 Example of HIMAFP visualization results. . . . . . 155 Figure 6.10 Signal monitoring for failure type A. . . . . . . . 158 Figure 6.11 Signal monitoring for failure type B. . . . . . . . 158 Figure 6.12 Signal monitoring for failure type C. . . . . . . . 159 Figure 6.13 Signal monitoring for failure type D. . . . . . . . 159 xiv Figure 6.14 Signal monitoring for failure type E. . . . . . . . 160 Figure 6.15 Signal monitoring for failure type F. . . . . . . . 160 Figure 6.16 Metric evolution AUC MOD in Failure type A for all combination of HS operator values for TSM6. . . . . . 163 Figure 6.17 Metric evolution AUC MOD in Failure type B for all combination of HS operator values for TSM1. . . . . . 164 Figure 6.18 Metric evolution AUC MOD for Failure type C for all combination of HS operator values for TSM2. . . . . . 166 Figure 6.19 Metric evolution AUC MOD for Failure type D for all combination of HS operator values for TSM2. . . . . . 168 Figure 6.20 Metric evolution AUC MOD for Failure type E for all combination of HS operator values for TSM5. . . . . . 169 Figure 6.21 Metric evolution AUC MOD for Failure type F for all combination of HS operator values in TSM1. . . . . . . 170 Figure 6.22 HIMAFP SBDs solutions for Harmony H2.1 in FailureA(I) ......................... 175 Figure 6.23 HIMAFP SBDs solutions for Harmony H2.1 in FailureA(II)......................... 176 Figure 6.24 HIMAFP SBDs solutions for Harmony H6.1 in FailureA(I) ......................... 177 Figure 6.25 HIMAFP SBDs solutions for Harmony H6.1 in FailureA(II)......................... 178 Figure 6.26 HIMAFP SBDs solutions for Harmony H3.4 in FailureB ........................... 180 Figure 6.27 HIMAFP SBDs solutions for Harmony H5.1 in FailureB ........................... 181 Figure 6.28 HIMAFP SBDs solutions for Harmony H3.1 in FailureC............................ 183 Figure 6.29 HIMAFP SBDs solutions for Harmony H3.3 in FailureC............................ 184 Figure 6.30 HIMAFP SBDs solutions for Harmony H1.1 for FailureD ........................... 187 Figure 6.31 HIMAFP SBDs solutions for Harmony H6.1 for FailureD ........................... 188 xv Figure 6.32 HIMAFP SBDs solutions for Harmony H2.2 in FailureE ........................... 190 Figure 6.33 HIMAFP SBDs solutions for Harmon H2.4 in FailureE ............................. 191 Figure 6.34 HIMAFP SBDs solutions for Harmon H3.1 in FailureF ............................. 193 Figure 6.35 HIMAFP SBDs solutions for Harmon H5.1 in FailureF ............................. 194 Figure 6.36 Boxplot graphic for the distribution of TTF through the different harmonies per TSM in failure type D. . . . . 196 Figure 6.37 RTS for harmony H2.1 of TSM2per SBD where RTS: Real Time Simulation. . . . . . . . . . . . . . . . . . 197 Figure 6.38 RTS for harmony H4.1 of TSM4per SBD where RTS: Real Time Simulation. . . . . . . . . . . . . . . . . 198 Figure A.1 Metric evolution AUC MOD in Failure type A in TSM1for all combination of HS operator values. . . . . . 215 Figure A.2 Metric evolution AUC MOD in Failure type A in TSM2for all combination of HS operator values. . . . . . 216 Figure A.3 Metric evolution AUC MOD in Failure type A in TSM3for all combination of HS operator values. . . . . . 216 Figure A.4 Metric evolution AUC MOD in Failure type A in TSM4for all combination of HS operator values. . . . . . 217 Figure A.5 Metric evolution AUC MOD in Failure type A in TSM5for all combination of HS operator values. . . . . . 217 Figure A.6 Metric evolution AUC MOD in Failure type A in TSM6for all combination of HS operator values. . . . . . 218 Figure A.7 Metric evolution AUC MOD in Failure type B in TSM1for all combination of HS operator values. . . . . . 218 Figure A.8 Metric evolution AUC MOD in Failure type B in in TSM2for all combination of HS operator values. . . . . 219 Figure A.9 Metric evolution AUC MOD in Failure type B in in TSM3for all combination of HS operator values. . . . . 219 Figure A.10 Metric evolution AUC MOD in Failure type B in in TSM4for all combination of HS operator values. . . 220 xvi Figure A.11 Metric evolution AUC MOD in Failure type B in TSM5for all combination of HS operator values. . . . . 220 Figure A.12 Metric evolution AUC MOD in Failure type B in TSM6for all combination of HS operator values. . . . . 221 Figure A.13 Metric evolution AUC MOD in Failure type C in TSM1for all combination of HS operator values. . . . . 221 Figure A.14 Metric evolution AUC MOD in Failure type C in TSM2for all combination of HS operator values. . . . . 222 Figure A.15 Metric evolution AUC MOD in Failure type C in TSM3for all combination of HS operator values. . . . . 222 Figure A.16 Metric evolution AUC MOD in Failure type C in TSM4for all combination of HS operator values. . . . . 223 Figure A.17 Metric evolution AUC MOD in Failure type C in TSM5for all combination of HS operator values. . . . . 223 Figure A.18 Metric evolution AUC MOD in Failure type C in TSM6for all combination of HS operator values. . . . . 224 Figure A.19 Metric evolution AUC MOD in Failure type D in TSM1for all combination of HS operator values. . . . . 224 Figure A.20 Metric evolution AUC MOD in Failure type D in TSM2for all combination of HS operator values. . . . . 225 Figure A.21 Metric evolution AUC MOD in Failure type D in TSM3for all combination of HS operator values. . . . . 225 Figure A.22 Metric evolution AUC MOD in Failure type D in TSM4for all combination of HS operator values. . . . . 226 Figure A.23 Metric evolution AUC MOD in Failure type D in TSM5for all combination of HS operator values. . . . . 226 Figure A.24 Metric evolution AUC MOD in Failure type D in TSM6for all combination of HS operator values. . . . . 227 Figure A.25 Metric evolution AUC MOD in Failure type E in TSM1for all combination of HS operator values. . . . . . 227 Figure A.26 Metric evolution AUC MOD in Failure type E in TSM2for all combination of HS operator values. . . . . . 228 Figure A.27 Metric evolution AUC MOD in Failure type E in TSM3for all combination of HS operator values. . . . . . 228 xvii Figure A.28 Metric evolution AUC MOD in Failure type E in TSM4for all combination of HS operator values. . . . . . 229 Figure A.29 Metric evolution AUC MOD in Failure type E in TSM5for all combination of HS operator values. . . . . . 229 Figure A.30 Metric evolution AUC MOD in Failure type E in TSM6for all combination of HS operator values. . . . . . 230 Figure A.31 Metric evolution AUC MOD in Failure type F in TSM1for all combination of HS operator values. . . . . . 230 Figure A.32 Metric evolution AUC MOD in Failure type F in TSM2for all combination of HS operator values. . . . . . 231 Figure A.33 Metric evolution AUC MOD in Failure type F in TSM3for all combination of HS operator values. . . . . . 231 Figure A.34 Metric evolution AUC MOD in Failure type F in TSM4for all combination of HS operator values. . . . . . 232 Figure A.35 Metric evolution AUC MOD in Failure type F in TSM5for all combination of HS operator values. . . . . . 232 Figure A.36 Metric evolution AUC MOD in Failure type F in TSM6for all combination of HS operator values. . . . . . 233 xviii NOMENCLATURE αProportion of labelled data c βProportion of known elements per class λNumber of bootstrapped subdatasets τNumber of iterations τLS Number of iterations in local search AD Anomaly Detection AL Active Learning AMOSA Multi-objective Simulated Annealing Algorithm AUC MOD Area Under Curve ROC Modified AUC ROC Area Under Curve ROC C1 Entropy of class proportions CB Constraint-Based CG Constraints Guidance CM Condition and Maintenance CN Constraints Contr Contribution CPPS Cyber-Physical Production Systems CPS Cyber Physical System CS Clustering Solution CV Cross Validation xix Chapter 1. Introduction to where it is supposed to be. With these results, it is clear that not only Industry 5.0 is far from being implemented, but also in some cases, Industry 4.0 is still an emerging challenge for many companies. Regarding the types of data sources, Figure 1.3 shows that the main source of data continues to be the traditional ones, such as historical production records (54.8%), work orders for products and processes (50.7%), data recorded by operators (48.6%) and machine parameters (46.6%). Figure 1.3: Types of data sources in the Manufacture Industry [2]. However, in industrial environments, it is frequent that the available databases have certain shortcomings, such as being partially labelled. To have a fully labelled database is specially important to facilitate the prediction of failures and the implementation of predictive maintenance operations. It is crucial to identify the work order or label that corresponds to the machine’s operation at any particular time which enables a comprehensive analysis without disrupting the machine’s operating conditions. Finally, it can be noticed in Figure 1.4 that the most commonly applied techniques for the analysis of the data are 1) trend analysis and 2) statistical analysis, while the use of Machine Learning (ML) techniques lags behind. Statistical methods are preferred over ML primarily because ML 4 1.1. Motivation Figure 1.4: Type of analysis used in Manufacture Industry [2]. models based on data tend to be black boxes, making them challenging to interpret. This lack of interpretability is a significant drawback for users, who may have limited experience with both ML techniques and the interpretation of outcomes. In general, both large and SMS companies are willing to take advantage of the benefits that Industry 4.0 offers to 1) improve the quality of their products and customer satisfaction, and 2) increase the efficiency of their production systems [9, 2]. However, looking at the degree of Industry 4.0 implementation, it can be observed that the difficulties for changing the culture of traditional manufacturing companies is one of the existing problems. This creates 1) a lack of qualified operators and domain experts in technologies related to Industry 4.0, 2) lack of knowledge in tasks that are useful for the improvement of their performance, 3) distrust in the recommendations provided by certain technologies (such as ML), and 4) low interest in investing capital for transforming the industry towards the new technological era [3]. In this sense, the general motivation of this Thesis is twofold. Firstly, to help manufacturing companies to take a step further into the technological era, losing their fear and distrust of solutions based on new technologies and techniques, and making this step easier for them. This step must be useful, easy to develop, understandable, with some kind of trust value, cost-effective and the effect or benefit should be achieved quite early. In this way, Industry would gain confidence 5 Chapter 1. Introduction and motivation in the process. Secondly, a key Industry 4.0 action to help manufacturing companies in particular is to develop predictive maintenance tasks. In this way, machine downtimes could be reduced and malfunctions that lead to breakdowns could be predicted, thereby improving productivity and increasing profits. In this sense, the design and development of predictive maintenance solutions based on ML techniques that enable to identify anomaly patterns in process variables (time series) related to failures is of great interest. Furthermore, such solutions should have the following characteristics: •Capable to adapt to several use cases, i.e., be generic. •Easy to use and understand by an inexperienced user. •The results should be provided in an understandable way. This can be achieved through the use of variables or parameters that are directly related to the operation of the machinery or the process. •Capability to work with databases that contain few information or little amount of labelled samples delving into a semisupervised environment. •Provide the user with a trust metric in the system itself and the results it delivers. In this way the user can estimate how reliable the output is, alleviating uncertainty and making the model more understandable. 1.2 Objectives After analysing the current state of Industry 4.0, the position of companies in this respect, the real level of integration of this technological era and therefore detecting the needs in the context of Industry 6 1.2. Objectives 4.0, it has been observed that the tasks related to predictive maintenance, i.e. the prediction of failures and breakdowns, are important and necessary in Industry 4.0. Taking all this into account, the main objective of this Thesis focuses on proposing a methodology capable of predicting failures in industrial machinery using ML techniques in the context of Industry 4.0. In order to develop a solution, it is necessary first of all to explore and analyse the most appropriate techniques, without ignoring understandability, explainability and user-friendliness for potential users who are unfamiliar with ML technologies. In order to develop this main objective, a number of particular objectives must be met: Obj.1 Study advanced data analysis techniques focused on the detection of anomalies in Industry 4.0 in order to predict failures in the industrial processes. Obj.2 Identify relevant features of process variable time series in the context of Industry 4.0 that help the detection of anomalies to predict a failure in a machinery system. Among others, distance and similarity metrics and their influence on ML techniques are to be studied. Obj.3 Obtain a methodological basis, which provides a) a valuable Know-How on the analysis of process variable time series (giving a better insight into the process, and thus improving it) and b) explainability and a user-friendly system for the end user. Obj.4 Design and develop a general methodology that can work with both labelled and partially labelled databases, i.e. designing a solution capable of inferring knowledge from the labelled part of the databases to label unknown samples. Obj.5 Study evaluation metrics in semi-supervised and supervised environments for the identification of patterns in process variable 7 Chapter 1. Introduction time series. Obj.6 Study existing trustworthiness metrics and implement them in the proposed solution. This way the user can be provided with an estimation about how reliable the solution is. 1.3 Structure This section outlines the structure of this Thesis, which is divided into 6 additional chapters. In Chapter 2, a literature review about the different aspects addressed in the Thesis is done. This state of the art focuses on Industry 4.0 and the enabling technologies as well as the necessary action points, such as artificial intelligence and predictive maintenance in order to improve the performance of industrial processes. With this in mind, the main objective is to analyse and compare solutions in the literature that use ML techniques to detect anomalies in industrial environments. Chapter 3 presents the general strategy proposed for the automatic and user-friendly methodology for predictive maintenance in the context of Industry 4.0. The proposed strategy is mainly based on a hyper heuristic inspired approach supported by meta-heuristics, in particular Harmony Search (HS). In this sense, the two main contributions of this Thesis presented in Chapter 5 and Chapter 6, respectively, will exploit this general strategy. In order to demonstrate the Thesis proposal based on the aforementioned hyper-heuristic strategy, Chapter 4 presents a real industrial case consisting in a cold stamping press for bolt forming. Chapters 5, and 6 present the different contributions made as part of the hyper-heuristic inspired methodology. Chapter 5 presents a 8 1.3. Structure hyper-heuristic solution for labelling partially labelled databases together with a trustworthiness metric. Chapter 6 proposes a hyperheuristic strategy that aims to provide the user with a set of predictive parameters for future breakages and failures in systems by detecting anomaly patterns in the associated process variable time series. Finally, Chapter 7 concludes the Thesis by analysing the results, determining the main contributions of this Thesis and outlining future research directions. 9 Chapter 1. Introduction 10 CHAPTER 2 STATE OF THE ART “Even though nothing changes, if I change, everything changes.” — Honor´e de Balzac This section contextualises the scope of the Thesis, does extensive research on the related literature,describes the literature gap in which the present Thesis is developed, and finally presents the contribution of the Thesis. Figure 2.1 shows an outline of the reviewed topics in the state of the art as well as the relationship between them. As Figure 2.1 shows, in Section 2.1, basic definitions and knowledge about Industry, its evolution over the years and the different industrial revolutions that exist are introduced. Specially for Industry 4.0, the key points, areas and action points for industrial enhancement are described together with the associated enabling technologies. In Section 2.2 one of the improvement actions, named predictive maintenance, is introduced. It will explain why it is necessary and the types that exist, as well as the techniques or methods and tech11 Chapter 2. State of the Art nologies needed to make predictive maintenance a reality in Industry. Specifically, this state of the art focuses on ML techniques applied to the detection of anomalies. In section 2.3, certain characteristics of the data associated with anomalies are considered, such as the type of anomalies to be detected, the type of data and the degree or percentage of labelling. In particular, an extensive analysis of collective Anomaly Detection (AD) techniques common to Time Series (TS) in industrial environments is developed. Depending on the percentage of labelling, this section is divided into 2 subsections, giving rise to 1) an exploration and analysis of the existing techniques for the autolabelling of Partially Labelled Databases (PLD)(Subsection 2.3.1), and 2) an extensive analysis of the techniques and solutions proposed in the literature to develop a collective type of AD in TS with fully labelled databases (Subsection 2.3.2). 12 Industry Key points Areas and actions of improvement Predictive Maintenance Model-based Data drivenbased Hybrid models Anomaly Detection Proposed solutions Proposed solutions SSL Classification SSL Clustering SSL Regression Heuristic Metaheuristic Hiper-heuristic Level of Expertise Required (LER) Low Medium High Feature Engineering Approach Heuristic Metaheuristic Industry evolution Industry 4.0Industry 1.0 Industry 2.0 Industry 3.0 Enabling Technologies Diagnosis Prognosis Fault detection and isolation Fault identification Remaining Useful Life Machine Learning (ML) Fault prediction Methods Industry 5.0 Real-time optimization Machine Learnign tasks Classification Regression Clustering 1 2 Initial data considerations Type of anomaly Percentage of Labels Type of data Learning approach Supervised Labelled databases Unsupervised Unlabelled databases Semi-supervised Partially labelled databases Time Series Multivariate Point anomaly Contextual anomaly Collective anomaly 3 Strategy Methodology Approach Type of knowledge Constraints Labels Filter Wrapper Distance-based Constraint-based Hybrid Strategy Method for parameter estimation 2.1 2.2 2.3.1 2.3.2 2.2 Hiper-heuristic 25% LABELS 1 100% LABELS 2 2.3 Figure 2.1: State of the art scheme. 13 Chapter 2. State of the Art in manufacturing processes. One task for improvement is the use of robots in collaboration with humans. In this way, tasks that are heavy for users and can cause injuries will be performed by robots. Therefore, the digitalization of knowledge and tasks will make it easier for users to manage and monitor production. The data provided by companies that have already taken these measures is an increase of 45-55% in productivity [5]. 4. Nowadays, in the inventories there is an excess of both purchased materials and those produced by the manufacturing company itself. It is therefore necessary to carry out stock monitoring actions to prevent the company from having a large capital stock. These improvement actions are not only focused on the exact accounting of stock, but also on the adequacy and planning of the stock necessary for production, thus eliminating excesses, the arrangement of stocks in the warehouses to use as little space as possible or overproduction. Through tasks such as real-time supply chain optimisation, Industry 4.0 can typically reduce inventory holding costs by 20-50% [5]. 5. It is indisputable that quality improvement is an important research area for the industry, not only at the product level but also because of the additional costs (in time, materials and labour) involved in the reprocessing of waste. These quality problems are often caused by unstable processes, poor packaging and malfunctioning or broken machinery. This can be solved by real-time monitoring and data analysis tools that detect abnormal behaviour and alert the worker to stop or modify the industrial process in time. By applying such levers, costs related to suboptimal quality can be reduced by 10 to 20 % [5]. 6. To achieve an optimal match between supply and demand it is necessary for the industry to understand the demand both in terms of quantity and characteristics of the desired product. Advanced and adequate demand analysis based on data models makes it possible to increase the accuracy of demand by 85% per 20 2.1. Industry 4.0 week and to understand the characteristics most sought after by customers [5]. 7. In terms of time to market, pioneering a product brings extra benefits and also makes it easier to respond earlier to potential problems. Enabling actions to improve this area include concurrent engineering or rapid experimentation/prototyping (e.g. through 3D printing). This can reduce time to market by 3050% [5]. 8. Nowadays, offering a good after-sales service and remote maintenance is a key element with great potential in the industry. Here, one of the most demanded actions is remote maintenance. These are software solutions that allow technicians to establish a secure remote connection to industrial equipment to carry out a diagnosis without the need to visit the site. A reduction in maintenance costs of between 10-40% has been observed thanks to remote and predictive maintenance using data analytics techniques [5]. 21 Chapter 2. State of the Art Service/ aftersales Resource/ process Asset utilization Labor Inventories Quality Supply/ demand match Virtually guided self-service Real-time yield optimization Smart energy consumption Remote monitoring and control Predictive maintenance Human-robot collaboration Remote monitoring and control Digital management Stock management Real-time supply chain optimization Digital quality management Data-driven demand prediction Industry 4.0 improvement actions to the 8 main areas Time to market Concurrent engineering Rapid simulation and experimentation 10 - 40% reduction of maintenance costs Productivity increase by 3 - 5% 20 - 50% reduction in time to market Forecasting accuracy increased to 85+% Costs for quality reduced by 10 - 20% Costs for inventory holding decreased by 20 - 50% 45 - 55% increase of productivity in technical professions through automation of knowledge work 30 - 50% reduction of total machine downtime Predictive Maintenance Remote Maintenance Smart resources consumption Intelligent IoTs Real-time yield optimization Automation of knowledge work Real-time data analisys Data-driven demand prediction Data-driven demand prediction Data-driven product design Service/ aftersales Time to market Supply/ demand match Quality Remote Maintenance Predictive Maintenance Virtually guided self-service Concurrent engineering Rapid simulation and experimentation Data-driven demand prediction Optimization of warehouse space Legend: Enabling Industry 4.0 technologies IoT ML Cloud CPS CobotsRTO AR Big Data AM Figure 2.4: Improvement actions to the 8 main areas of interest and its corresponding enabling technologies in Industry 4.0. Source :[5] re-illustrated and modified for this Thesis. CPS: Cyber Physical System, RTO: Real-time optimization, ML: ML, AR: Augmented Reality, IoT: Internet of Things, AM: Additive manufacturing. 22 2.2. Predictive maintenance As can be seen, these improvement actions in the Industry 4.0 have a direct effect on 1) enhancing the value of Industry, 2) improving production and product quality and 3) increasing the profits of companies. As referenced above, predictive maintenance is one of the key improvement actions that is directly correlated with driving the correct and optimal use of machinery. This is of crucial importance for companies, especially those with heavy and expensive machinery. In order to perform a good predictive maintenance it is necessary to have a good history of data captured by different sensors and intelligent devices (IoT and CPS) placed in the different key points of the machines. After obtaining the data and storing them, it is usually necessary to apply intelligent analysis techniques (ML) that create models capable of identifying patterns and detecting anomalies to either foresee undesired situations, such as breakages, future stoppages or wear in any element of the machine. With this analysis it is also possible to perform optimisation in real time (real-time optimisation and cloud computing) and therefore, improve productivity. 2.2 Predictive maintenance Maintenance techniques, together with Industry 4.0, have developed new skills and improvements over the years. The earliest maintenance technique is basically corrective maintenance, which takes place only after a failure has occurred. A later maintenance technique is preventive maintenance (also called planned maintenance), which sets a periodic interval to perform preventive maintenance regardless of the health status of a physical asset. With the rapid development of modern technology, products are becoming more complex and higher quality and reliability are demanded, which increase the cost of preventive maintenance. Therefore, more efficient maintenance approaches, such as predictive maintenance [19], are needed. Predictive maintenance is based on the continuous monitoring of a machine or a process, allowing maintenance to be performed only when it is needed. It al23 Chapter 2. State of the Art lows the early detection and prediction of failures thanks to predictive tools based on historical data (e.g. ML), statistical inference methods and engineering approaches [20]. As mentioned before, predictive maintenance brings different benefits to production environments like productivity improvement, reduction of system failures, minimisation of unscheduled machinery downtimes, increased efficiency in the use of financial and human resources, and scheduling optimisation of maintenance interventions [21]. 2.2.1 Predictive maintenance categories Predictive maintenance focuses on two main parts: fault diagnosis and fault prognosis. Diagnosis involves all those tasks that focus on the detection, isolation and identification of faults in a machine or industrial process. Thus, the steps to be followed in diagnosis [19] are 1) detect the faults by determining when something in the monitored system is wrong, 2) isolate the faults by determining which element is at fault and 3) identify the fault by determining the natural cause of failure. After the diagnostic process, the user usually has an idea of what faults are occurring and why. With this information, prognostic tasks can be performed on the systems, which refer to the task of predicting that a failure will occur in the system. Two main tasks can be observed at this point, failure prediction and Remaining Useful Life (RUL) estimation. Fault or failure prediction determines that a failure is imminent and estimates when it may occur, i.e. the Time To Failure (TTF). The latter task, RUL estimation, assesses how long a machining component has until it cannot function anymore in accordance with its intended purpose, given the current machine age and condition, and the past operation profile [19]. Prognosis is said to be much more efficient than diagnosis for fewer downtimes in industrial processes [19], however it can sometimes be more uncertain. In this respect, in the scope of the present Thesis it is considered of great 24 2.2. Predictive maintenance interest to provide the user with a confidence value about how much the user can rely on such predictions. 2.2.2 Methods for predictive maintenance development The way in which predictive maintenance tasks, both diagnosis and prognosis, are performed can be done in different ways depending on 1) the type of information available, i.e. whether there is data recorded over time (sensor data), information on previous machine conditions and maintenance (CM) [22] or the operating conditions/work order (WO) that a particular machine is running, i.e., to know the operating parameters of a machine [22, 23], and 2) the degree of expert knowledge required or available. Having this in mind, and as can be seen in Figure 2.5, three different estimation methods can be chosen [19, 20, 22, 24, 25]. •White-box or model-based model. This method requires a lot of expert knowledge as it is based on the physical laws governing the analysed elements. Knowing the conditions of the machine is essential, but having experimental sensorised data is not necessarily required. In this case the result is usually concise and understandable for experts but it is time-consuming, requires a lot of prior knowledge and valuable information related to the machine conditions and maintenance. •Black box or data-driven model. This methodology requires mainly a good historical sensor data [22] while data from condition maintenance is not needed. This type of model is mainly characterised by the fact that it does not require expert knowledge and is based on mathematical models such as ML or statistical algorithms. They are called black box models because they are commonly difficult to understand to domain experts, as knowledge is extracted from the data and not from physical laws. 25 Chapter 2. State of the Art •Grey box or hybrid models. These models are a mixture of the two previous ones, trying to use the benefits of each one. With the advent of Industry 4.0 and the ease of massive data capture, data-driven models are of great interest for their application to predictive maintenance, especially when explicit relations between the failure and the related features are unknown. If properly managed, this can result in a fast and reliable way for the prediction of failures and the optimized management of maintenance tasks, fulfilling all above mentioned advantages about a good predictive maintenance in Industry 4.0 and even in the already imminent Industry 5.0. Specifically, this Thesis aims to design and develop a new methodology for industrial fault prediction, as detailed later. WO Low Physics modeling Model Type AI approach Statistical approach Data driven methods Sensor data Physical approach Model Based methods CM data AI, statistical & physical approaches Hybrid methods CM data Sensor data WO WO High Expert knowledge level Data modeling Figure 2.5: Categorization and definition of prognosis methods. After evaluating 1) the importance of predictive maintenance in improving productivity in companies within the context of Industry 4.0, 2) the categories of predictive maintenance identified in the literature and 3) the methodologies used for its development, it can be concluded that performing an adequate and efficient prognosis is of vital importance for the Industry. Furthermore, given the immense 26 2.3. Anomaly detection for fault prediction amount of data available nowadays, it makes sense to use methods based on historical data together with ML enabling technologies to develop a correct failure prediction system. 2.3 Anomaly detection for fault prediction Within ML, AD can be applied to analyse signal patterns and identify unexpected behaviours [6] that can predict a subsequent failure. When developing ML algorithms for AD, it is necessary to pay attention to several concepts, such as: the type of data, the type of anomaly and the available labels. The first key aspect for the proper selection for AD method is the nature of the data. Each data describing an instant in time is composed of one attribute or variable (univariate) or a set of attributes (multivariate). Regarding multivariate datasets, these may be of the same data type or a mixture of several types, e.g. categorical or continuous. Finally, there is the possibility that the variables may or may not be related to each other. Specifically in Industry, the most commonly used datasets contain a) individual multivariate data, i.e., data of different but unrelated variables (hereafter referred to as multivariate data) or b) TS, i.e. data related to one or more variables monitored over a period of time t[6, 19]. In terms of the anomaly type that can occur and therefore those for which the AD algorithm is to be designed, three types of anomalies are distinguished [6]: •Abnormal points or outliers: these are random points that are out of the pattern of the rest of the points. They usually appear in a drastic way and can be identified with simpler algorithms, such as the implementation of rules or limits based on the statistical distribution of the data. Figure 2.6 shows how there are 27 Chapter 2. State of the Art certain points such as O1and O2that are outside the rest of the majority groups [6]. Figure 2.6: Outlier anomaly example [6]. •Contextual anomalies: these are points or instances that may arise along a TS which, depending on the environment or context in which they are located, are called anomalies with respect to their surroundings. Over the last few years, these have been the most studied by researchers in the field. As can be seen in Figure 2.7, where the evolution of the temperature over a year is shown, there are two specific moments in time where the value of the temperature is the same (t1and t2). The great difference between both is the context in which they are, since for the moment t1, it is an expected or normal value. For t2its value leaves the context of the environment in which it is [6]. Figure 2.7: Contextual anomalies example [6]. 28 2.3. Anomaly detection for fault prediction •Abnormal patterns or collective anomalies: these are formed when a set of small fragments of the TS does not fit the other patterns in the data set. As can be seen in Figure 2.8, there is an area (marked in red) in the TS where the pattern differs from the rest of the series [6]. This type of anomaly is the most commonly found in the industrial environments. Figure 2.8: Collective anomalies example [6]. Finally, the third key point for selecting an AD strategy is related to the data labels. It is especially important in the industrial environment for predictive maintenance to know the labels of the data because, a bad or missing label can lead to an erroneous or inaccurate solution. However, it is worth noting that this is usually a very costly process for domain experts to do, not only in terms of quality but also in terms of time. Additionally, in AD commonly there is a large imbalance in the databases, with usually more normal tags than anomalous cases. Based on the extent to which the labels are available, AD techniques can operate in one of the following three modes: •Supervised learning. Techniques trained in supervised mode assume that the database is completely labelled for normal behaviour as well as for the anomaly class. The typical approach in this case is to build a prediction model for the normal or 29 Chapter 2. State of the Art Finally, hyper-heuristics allow for the handling of large, computationally complex search problems instead of only solving single or lower complexity problems. In fact, hyper-heuristics provide a superior, automated alternative to meta-heuristics that typically require fine-tuning of related parameters for optimal performance on a specific problem [51]. Hyper-heuristics automate a combination of heuristics and/or meta-heuristics as another search process for finding the best parameters to control heuristic or meta-heuristic behavior, with the overall objective of balancing explorative and exploitative search to achieve a solution. Accordingly, hyper-heuristics based approaches are are based on hyper-heuristic algorithms [52]and are the less referred in the literature. These are also understood as “a combined meta-heuristic search method that through the automatization of its stochastic parameters, such as: the selection, generation, combination or adaptation of several components to efficiently solve computational NP-hard search problems [53, 54]”. In the literature different definitions of this type of algorithms can be found but mainly are defined as “a highlevel heuristic approach that, given a particular problem and a set of low-level heuristics, its main objective is to find the best heuristic or sequence of heuristics to solve the problem rather than to provide the solution to the problem directly [55, 54]”. Similar to other domains, automatization is a key challenge to promote the penetration of a given technology. While heuristics and meta-heuristics require human intervention to optimize their free parameters for obtaining the best balance in their search capabilities, hyper-heuristics provide an automated solution for this challenge. As hyper-heuristics are gaining prominence due to its properties, different taxonomies are proposed in the literature which are differentiated in Figure 2.11. The taxonomy proposed in [55] differentiates between two domains 1) the nature of the heuristics to use and 2) the source of knowledge or feedback in the process. Regarding the 36 2.3. Anomaly detection for fault prediction Nature of the space search Hyper-heuristics Feedback Techniques Offline learning Online Learning No-learning Generation Selection Construction Perturbation Construction Perturbation Greedy and peckish Random selection Meta-heuristic based By learning mechanisms Figure 2.11: Taxonomy of hyper-heuristics. nature of the heuristics and how they are implemented in the hyperheuristic two categories can be distinguished [56]: a) selection which are methodologies that search the best configuration of given heuristics or b) generation which, in contrast, are methodologies that generate hyper-heuristics from a set of given heuristics. Within this classification, two types of low-level methodologies can be distinguished: 1) constructive heuristics, which gradually creates a suitable final solution from start to finish. The goal is to intelligently choose the most appropriate heuristic for the current problem state. The process continues until a complete solution is achieved, and since the problem has a fixed size, there is a natural endpoint to the construction process, or 2) perturbative or local search heuristics, which develop a partial modification of a heuristic solution previously created. The aim is to iteratively choose and select the heuristics based on the current solution that is already complete. Delving into the knowledge used as feedback from the search process, it can be find three main categories of hyper-heuristics 1) online learning, where the learning process occurs during the resolution of a given search problem, while the hyper-heuristic interacts with a de37 Chapter 2. State of the Art fined environment, such as utilizing reinforcement learning for heuristic selection and employing meta-heuristics as high-level search strategies over a heuristic search space. As a result, the high-level strategy can leverage task-specific local properties to identify the most suitable low-level heuristic to apply, 2) offline learning, in this case the idea is to accumulate knowledge, often in the form of rules or programs. The main idea is to use a given set of training instances to learn and then, be generalized to solve unseen instances. Offline learning techniques allow the algorithm to learn from previous problem-solving experiences, resulting in a more efficient and effective methodology. Some examples of offline learning approaches using hyper-heuristics include case-based reasoning, and genetic programming. By leveraging these techniques, the algorithm can identify patterns and apply previously learned knowledge to improve its problem-solving abilities, where one learns from one training set to apply it to another, and 3) no-learning refers to those methodologies where there is no feedback during the search process. Regarding the techniques used to develop hyper-heuristics, four methods can be distinguished [57]: 1) Hyper-heuristic based on random selection, which is the most basic and straightforward of the hyper-heuristic family, as it randomly selects a low-level heuristic from a given set at each decision point without considering past performance, 2) Greedy and peckish hyper-heuristic, which selects and applies the low-level heuristic that produces the largest improvement to the current objective value at each decision point. If no improving low-level heuristic is available, the heuristic chooses the one that leads to the smallest deterioration. Greedy hyper-heuristic requires a preliminary evaluation of each low-level heuristic in the set to select the best one, which makes it a more computational complex approach, 3) Meta-heuristic based hyper-heuristic, wherein a meta-heuristic is the method of global search that operates in the solution space of a problem and employs strategies to escape local optima and 4) Hyperheuristic by learning mechanisms, that uses various techniques to learn the performance history of low-level heuristics. At each decision point, 38 2.3. Anomaly detection for fault prediction a hyper-heuristic chooses a promising low-level heuristic based on the effectiveness of each one gathered from earlier stages or previous runs. Once the proposed taxonomy in Figure 2.11 is presented, the related state of the art can be analysed accordingly. Table 2.1 summarises the classification and comparison done of the literature references. Table 2.1: Comparison of the literature. TS: Time Series, MV: Multivariate, L: Labels, CN: Constraints, F: Filter, W: Wrapper, CB: Constraint-Based, DB: Distance-metric Based, CG: Constraints Guidance, S: Seeding, OF: Optimization Function, H: Heuristics, MH: Meta-heuristics, HH: Hyper-Heuristics. Reference Data Knowledge Strategy Methodology Approach [41, 58, 59] MV L W CB (S) H [42, 60] W CB (CG) H [38] W DB MH [61] W CB (OF) H [62] F CB (OF) H [50, 63, 64, 65] F CB (OF) MH [39] CN W Hyb (DB + OF) H [66, 67] W CB (CG) H [68] W CB (CG) H ensemble [69, 43, 70, 71] W DB H [72, 73] F CB (OF) H [37] F CB (OF) H ensemble [74, 75] TS LW CB(CG) H [48] F CB (OF) MH [76, 77] CN W DB H [35] W CB (CG) H [78] F CB (OF) H Multivariate data In the context of MV data, it can be observed that there are solutions that use labels and constraints. In both types of knowledge, wrapper strategies are more common and have more diversity in the methodologies than filter ones. Filter strategies are mainly constraint39 Chapter 2. State of the Art based by means of an objective function. In regards to the solutions that use labels and wrapper strategies, both heuristic and meta-heuristic algorithms are used. Concerning heuristic algorithms, labels are introduced as a) initial seeds of a Kmeans [41, 58, 59] with predefined parameters, b) a guiding method in the formation of clusters in DBSCAN [42, 60] or c) a evaluating method of membership agreement in Fuzzy-C-Means [61]. In contrast, the use of meta-heuristics is developed by using a distance-based methodology and SA as the main algorithm [38]. Regarding the solutions that use labels and filter strategies, it is observed that meta-heuristics are commonly employed. These optimise an objective function either using GA [64, 63] or a multi-objective optimisation with AMOSA [65, 50] of different evaluation indexes with predefined parameters. Finally, with regards to filter strategy and heuristics few references use an objective function as in [62] which employs a fuzzy algorithm. Delving into solutions using constraints, it can be observed that no reference employing meta-heuristics is encountered. However, the use of heuristic ensemble methods has proliferated in this field [68, 37]. Specifically, for the solutions using constraints together with a wrapper strategy, hybrid [39] and distance-based [69, 43, 70, 71] methodologies can be found. These solutions are complex and problemdependent wherein the learning metric of the clustering itself can be modified easily as in hierarchical [69, 43], EM [39, 70] or K-means [71] clustering algorithms. Additionally, there are solutions that use constraints to guide cluster formation and improve the quality of the solution by means of both heuristics (K-means [67], DBSCAN [66]) and ensemble methods [68]. Additionally, using this type of knowledge (constraints) and filter strategies, the solutions are based on constraint-based methods 40 2.3. Anomaly detection for fault prediction employing objective functions using heuristic algorithms (K-means) and AL layer [72, 73] or an ensemble structure [37] of hierarchical clustering methods. Time series data Throughout the literature review, it can be observed that the solutions proposed for TS data are scarce due to their complexity both in terms of size and noise [76, 35]. There are solutions that jointly use semi-supervised clustering prior to develop a classification system [79, 26]. In [74, 75], it can be found a solution based on wrapper strategy together with heuristics (Graph based clustering) guided by the use of labels. By contrast, authors in [48] combine two meta-heuristics (GA and SA) using a filter strategy, thus leveraging the strengths of both and obtaining promising results in the context of fault diagnosis of bearings. If knowledge is introduced in the form of constraints, authors usually develop not an holistic solution but a partial one adaptable to different solutions. In this respect, it is common to use filter strategies applied together with an AL technique [78] or wrapper strategies using distance-based methods [76]. Finally, some complete solutions based on wrapper strategies and heuristics can be found. In [35] authors use the user-supplied constrains as guidance (AL) for two types of clustering, while [77] uses a complex metric learning integrating the distance requirements according to the constraints of the elements. Discussion As this literature review has shown, there are a wide variety of references using SSL clustering based solutions for labelling PLD. In gen41 Chapter 2. State of the Art eral, it can be observed that constraint-based methodologies are the most commonly used, as well as solutions that use a specific heuristic. Many of the proposed solutions are rich in knowledge and techniques, ranging from the simplest heuristic algorithms to the utilisation of a) complex integrated learning metrics to evaluate constraints, b) ensemble algorithms or c) meta-heuristic systems with internal parameterization and optimization. However, it has not been found yet a proposal capable of: a) integrating a method applicable to both types of data (MV, TS), b) being completely autonomous in terms of optimization of internal parameters of the algorithms and of the method itself for the correct labelling of semi-supervised databases, c) being user-friendly and d) providing a trustworthiness metric for improving the user confidence on the system. The latter is very important since in semi-supervised systems there is no certainty about the accuracy of the solution, specially due to the lack of labelled samples. 2.3.2 Collective anomaly detection solutions in time series data in the context of Industry 4.0 Throughout this subsection a description and analysis of the characteristics of the anomalies that can be found in industrial environments and the parameters of the signals that can be predictors of failures in industrial processes or machinery systems are presented. In this way and together with AD techniques based on ML, predictive maintenance systems can be developed. As mentioned before, many industrial processes suffer from a degradation of the normal behaviour that concludes on a failure of machinery, tool or process. This degradation of the normal behaviour is understood as an anomaly in the process, that is recorded in the process signals (TS) as collective anomaly. Actually, the common AD procedure starts from generating an experimental database with preselected variables associated to Nfailure cases. Afterwards, the 42 2.3. Anomaly detection for fault prediction experimental database is analysed and the abnormal patterns that enable to predict the failures are usually identified based on heuristic rules. Finally, an algorithm is designed and developed to implement a real-time detection of such anomaly patterns [80, 81, 82, 83, 84]. However, these techniques are costly and tedious as they are centred on the specific use case. The detection of this degradation is key to predict the failure and in this Thesis it is understood as the period of time or time-window immediately prior to the failure where the behaviour stops following the normal trend. This time-window depends on the dynamics and characteristics of the type of failure in a particular process and may be unknown beforehand. Likewise, it is also common to ignore other important parameters as: (a) The particular process variable or variables that enable the prediction of the failure among all the Mpre-selected process variables, hereinafter referred to as a Time Series Measurement (TSM ). (b) The general statistic or extracted feature (Fe) of the process variables within the degradation time-window size that enable the prediction of the failure. (c) The size of the time-window (Wz) to perform the feature extraction. (d) The threshold value (Th) that the statistic feature value should exceed or not to be categorized as a failure. An example of the above-mentioned parameters can be seen in Figure 2.12. Firstly, there is a process variable or TSM recorded until a breakage of an element occurs. In this case, the “maximum” Fe is extracted through the time-windows of size Wz giving rise to the extracted FeW z(TSM1) values. It is observed that with the esti43 Chapter 2. State of the Art Wz FeWz (TSM1) Th TSM1 Failure Figure 2.12: Relevant parameters in the AD context. mated Th, based in some rules, in the degradation window Wz, the anomalous points that predict the future breakage can be detected. Regarding the detection and diagnosis of faults in industry, a multitude of algorithms and methodologies related to collective datadriven AD have been developed over the years. Among the solutions that use data-driven methods, two main approaches can be distinguished: a) the rule based heuristic, that provides the user with a single possible solution [85] with a low computational cost at the expense of a high level parameterisation and expert knowledge [86], and b) the meta-heuristic [87, 86, 88], which follows a problem-independent optimization strategy for adjusting the parameters of the meta-heuristic algorithm, that obtains an optimal solution [89]. However, these latter tend to be too problem-specific or knowledge-intensive to be implemented in cheap, easy-to-use computer systems. Following this idea, hyper-heuristic builds systems which can handle classes of problems rather than solving just one problem delving with the issues of meta-heuristics. Using hyper-heuristic can be interesting in AD as it is able to optimise both parameters and algorithms at the same time. 44 2.3. Anomaly detection for fault prediction In the scope of this Thesis, no reference has been found in the literature that performs a classification of proposals evaluating these concepts, i.e., heuristics, meta-heuristics and hyper-heuristics based approaches, in terms of failure prediction. First of all, it is worth noting that the vast majority of references in the literature uses a heuristic method. In these cases the prior knowledge required about the problem at hand is high. This can be demanded either 1) in the initial steps, where data is conscientiously preprocessed and selected according to the problem [90], 2) in the parameter setting process for a given problem, such as Wz and thresholds, to adjust the internal hyper-parameters of the algorithms, or to make crucial decisions necessary for the AD system to accurately detect anomalies [91, 92, 93]. In this literature review, a classification is proposed based on the amount of knowledge required by each proposal. The categorization is decided depending on 3 actions required by the user: a) the user has to fix both types of parameters (problem parameters and hyperparameters of the algorithm); b) the user has to perform a deep data preprocessing; c) the user has to interact with the failure prediction system. In this context, three categorizations are defined: •High Expert Knowledge, when the three actions above are required. •Medium Expert Knowledge, when two of the three actions above are required. •Low Expert Knowledge, when one of the three actions above is required. Table 2.2 summarises the classification and comparison done between the related literature in the context of collective AD. The table classifies the solutions by its expert knowledge requirements and indicates the methods used for parameter estimation differentiating between fixed, calculated and optimized parameters. 45 Chapter 2. State of the Art Once a database is available in which the work orders or operating conditions are fully identified, the predictive maintenance tasks can be performed. As has been observed in the industrial paradigm, it is common to find the situation where the explicit relationships between process variables and failures are unknown. There is a long history of ML techniques, by which it can be stated that they are a good solution to determine the implicit relationships between process variables and parameters and failures enabling to predict failures and breakdowns, as they are based on historical data and experience. A system applied to label databases in the context of Industry 4.0, for predictive maintenance tasks, should fulfil the following requirements: •Integrate a solution applicable to the common types of data in industrial environments (MV, TS). •Be completely autonomous in terms of optimization of internal parameters of the algorithms and of the solution itself for the correct labelling of PLD. •Be user-friendly and understandable. •Provide a trustworthiness metric for improving the user confidence on the system. In this case, based on data models and applying ML techniques, the AD task can provide a proper solution. Section 2.3.2 shows that the detection of collective type anomalies in the industrial environment is gaining relevance. In the literature review, it has been observed, firstly, that heuristic solutions that require a medium-high level of expert knowledge are the main employed. These solutions become problem-dependent, with a large number of parameters to be known and predefined and, consequently, with poor generalisation capabilities. Secondly, the use of meta-heuristics to alleviate the burden 52 2.4. Conclusions of parameterisation has not been widely exploited in the literature. In fact, most proposals use them in hybrid or complex models to estimate hyper-parameters but with the limitation of using predefined preset parameters. In fact, few references are found based on hyperheuristic inspired approaches so as, some of them try to be either 1) an integrated and optimised ad-hoc solution for fault prediction by means of innovative Deep-learning structures, or 2) a best model selection framework that gives to the final user the best statistical model and some extracted features of the TS that detect the anomaly. However, these approaches are not user friendly, understandable and have a reduced number of pre-designed algorithms, i.e., some Deeplearning and statistical algorithms respectively, and extracted features to choose from, limiting the best possible solutions and making them less generalized. In conclusion, in the scope of AD the proposed solutions are limited by several conditions: 1) the need of input data preprocessing, 2) the need of expert domain knowledge to classify the faults, 3) the suboptimal definition of the relevant parameters of the problem, 4) the evaluation in benchmark or synthetic datasets [110] which does not guarantee its validity for specific real-world use cases, and 5) the lack of understandable and interpretable information to the final user about the TTF, important features [111] or relevant parameters that help to detect the anomalies, making the proposals a black-box model for them. After the analysis of the literature in the context of failure prediction in industrial environments through the detection of collective anomalies, which is the common type of anomaly identified in the industrial environments, it can be determined that there is no proposed solution that jointly addresses the following points: 1) use ML techniques, 2) be generic and not problem-dependent, 3) be easy to use and understandable to the user, 4) provide the user with the optimal set of parameters that are intrinsically related to the faults in order to predict them, 5) be able to set the necessary parameters ensuring the 53 Chapter 2. State of the Art best results through optimisation techniques, 6) search for the best set of heuristics or rules that predict failures rather than the solution directly, and 7) be trained and tested over a real industrial use case. All these points provide the user with an optimal solution that identifies the potential parameters and heuristics which are predictors of the related failures. This system is understandable strengthening the user’s knowledge and easy to use. A holistic system based on collective AD techniques applied to predictive maintenance tasks, should fulfil the following requirements: •Use ML and optimization techniques. •Be generic and no problem-dependant. •Be easy to use and understandable to the user, which should only introduce the data and read the results, understanding them and taking the relevant actions. •Provide the user with the optimal set of fault predictor parameters for detecting collective anomalies. •Be able to set the parameters it needs. Look for the best set of heuristics that predict the faults instead of the solution directly. •Be trained and tested over an industrial use case. 54 CHAPTER 3 A HYPER-HEURISTIC INSPIRED METHODOLOGY FOR FAILURE PREDICTION This chapter describes the integral hyper-heuristic inspired methodology proposed in this Thesis for failure prediction. Specifically, this chapter illustrates how the full methodology is developed and unified to tackle two main and relevant issues in the context of Industry 4.0. Figure 3.1 shows the conceptual scheme of the proposed hyperheuristic inspired methodology. Specifically, this methodology has two main objectives to tackle the gaps found in the literature specified in Chapter 2: 1) to automatically label PLD and 2) to obtain the best set of parameters that are predictors of a failure from a FLD in the context of Industry 4.0. In the former this methodology, named PLASH, has the following objectives: 1) be a user-friendly and autonomous system for labelling PLD, 2) evaluate and optimize the different clustering solutions performed by a new semi-supervised metric PSOM and 3) provide the user with the results in an understandable way together with a trustworthiness metric. Regarding the task of collective anomaly detection for failure prediction, named HIMAFP, 55 Chapter 3. A hyper-heuristic inspired methodology for failure prediction the aim is to find the optimal set of parameters that are predictors of a specific type of failure that allows to reliably and robustly predict the failure parameters providing the final user with information to detect the anomaly and anticipate the failure. As shown in Figure 3.1 the hyper-heuristic methodology can be divided into three main stages. The initial component, i.e., Data processing focuses on developing the required data treatment for different types of data inputs to be properly processed in the following component. The central part of the methodology, i.e. HS procedure, focuses on an optimization search process that employs a well-known and established metaheuristic algorithm called Harmony Search. Finally, the methodology includes a stage for results visualization and interpretation to the user in a clear and understandable manner. 56 Hyper-heuristic Methodology Initialization nIter = 1 nIter = ζ ? STOP HMCR PAR RSR Metric evaluation Sorting and selection of the best harmonies Yes Improvisation nIter = nIter+1 No Best Harmonies Selected Is the database completely labelled? Initial Database NO YES Fully labelled database (FLD) Partially labelled database (PLD) X%<100% Data processing 100% HS procedure A Hyper-heuristic Inspired Approach for Automatic Failure Prediction in the Context of Industry 4.0 Q1 PLAHS: a Partial Labelling Autonomous Hyper-heuristic System for Industry 4.0 with application on classification of cold stamping process Q1 Results visualization and interpretation Final solution with the most representative parameters that predict the failure User approves labelling proposal Real time failure detection system Decision Making towards Predictive Maintenance Methodology PLD FLD Figure 3.1: Hyper-heuristic inspired methodology applied in this Thesis. By following this approach, the user simply needs to input the database for performing the analysis. There are two potential scenarios based on the database at the beginning. •If the database is PLD, firstly the system will automatically label the unlabelled samples of the database. As stated before, this part of the methodology (in orange color in the Figure 3.1) is 57 Chapter 3. A hyper-heuristic inspired methodology for failure prediction called PLAHS. After the database has been labelled and therefore it is a FLD, an AD task is developed to obtain a set of failure predictor parameters. The journal paper associated with this methodology is found in: Applied Soft Computing (Q1) “Minor revision”. Navajas Guerrero, A., Portillo, E. & Manjarres, D., PLAHS: A Partial Labelling Autonomous Hyper-Heuristic System for Industry 4.0 with Application on Classification of Cold Stamping Process. Available at SSRN 4288726. •If the database is FLD, the system can directly develop the AD task to obtain a set of parameters that are potential predictors of the failure of interest. As mentioned before, this part of the methodology (in purple in the Figure 3.1) is called HIMAPF. The journal paper associated with this methodology is found in: Computers and Industrial Engineering (Q1). NavajasGuerrero, A., Manjarres, D., Portillo, E., & Landa-Torres, I. (2022). A hyper-heuristic inspired approach for automatic failure prediction in the context of industry 4.0. Computers & Industrial Engineering, 171, 108381. 3.1 Data processing In this stage an initial treatment and processing of the data is conducted, i.e., it is adjusted to the necessary characteristics for each of the main objectives addressed in this Thesis through the proposed hyper-heuristic inspired approach. In both cases, i.e., the auto-labelling (PLAHS) and the automatic failure prediction (HIMAPF), a unique data processing methodology is proposed. In general it is necessary to organize the data in a coherent way with the proposed hyper-heuristic inspired methodology, and preprocess the data according to the requirements of the HS-based 58 3.2. Harmony search procedure solution. Specifically, for PLAHS, the labeled and unlabeled data should be organized properly to extract reliable information, while in HIMAFP, data processing is necessary to isolate failure cases for independent consideration by the HS. In Chapter 5, Subsection 5.2.1, an extensive explanation of the data processing techniques used for the auto-labelling task is described in depth. It details how the data is utilized and integrated for analysis purposes. Similarly, Chapter 6, Section 6.3 also offers an in-depth analysis of how the data is processed for the automatic failure prediction task. 3.2 Harmony search procedure In the second stage, an optimisation process takes place for both scenarios, which will provide the basis and coherence to the hyperheuristic strategy. Optimization methods are widely employed for various purposes, such as industrial planning, econometrics, scheduling, decision making, engineering, and computer science applications. The optimization field is an active area of research, with new techniques continuously being developed [112]. Optimisation involves choosing the most suitable option from a given set of alternatives, subject to relevant constraints. The process entails minimizing or maximizing the objective or cost function of the problem. The process iteratively selects values from a permissible set until the optimal outcome is attained, or the stopping criterion is met. Meta-heuristic algorithms are well-established and effective methods for solving optimization problems with satisfactory outcomes [113]. These algorithms possess several features, such as simplicity, robustness, and flexibility, which make them an attractive area of research to efficiently solve computational NP-hard search problems [114, 53]. Many meta-heuristic algorithms draw inspiration from natural phe59 Chapter 3. A hyper-heuristic inspired methodology for failure prediction nomena, such as Particle Swarm Optimization (PSO), SA, GA, and HS. These algorithms are intelligently designed and can produce effective solutions to a wide range of optimization problems [112]. Chapter 2, Section 2.3.1 states that using meta-heuristic algorithms is a common method to develop hyper-heuristic algorithms. Thus, the proposed hyper-heuristic inspired methodology to solve both tasks is developed by means of a meta-heuristic algorithm. In particular, the proposal of hyper-heuristic inspired methodology is implemented by a meta-heuristic algorithm, specifically the HS algorithm, to explore and exploit the resolution space for an optimal solution to a problem. The HS algorithm is a population-based algorithm that uses a set of solutions, called harmonies, stored in the Harmony Memory (HM). It is achieved through an iterative process, applying several improvisation operators to find the best fitness value of a solution vector, called harmony. HS algorithm emulates the collaborative behavior of musicians who adjust their instruments’ pitches to achieve a harmonious sound. HS is a highly effective meta-heuristic algorithm that generates a solution vector intelligently through the exploration and exploitation of the search space. Since its emergence in 2001 in [115] by Geem et al., the HS algorithm has been recognized as a highly efficient population-based metaheuristic algorithm for combinatorial optimization. It has reached significant interest from researchers across diverse fields, who have enhanced its performance by fine-tuning its parameters and integrating its components with other meta-heuristic algorithms [112]. This algorithm has shown in several pieces of research [116, 117, 118] that thanks to its internal operators (HMCR memory consideration, PAR adjustment rate and RSR randomness) an optimal interaction between exploration and exploitation in the search of the best solution is achieved. As shown in Figure 3.1, this process is mainly composed of five 60 3.2. Harmony search procedure main parts i.e., 1) initialization, 2) improvisation, 3) the proper methodology depending on the scenario, 4) metric evaluation and 5) sorting and selection of the best harmonies. 3.2.1 Harmony search algorithm Initialization. This step is only performed during the first iteration and the initial harmony is proposed. To create the HM, two factors must be considered. Firstly, the Size of the Harmony Memory (HMS) that determines the number of harmonies that will constitute the overall HM. Secondly, the choice of encoding system is crucial. The encoding consists of a certain number of notes (nx), which serve as parameters to be optimized and form the basis of the harmony. The harmonies of the HM are generated randomly among the feasible values of each note. An example of the structure of the HM is shown above. In the HM 3.1 it is shown a matrix of different nxnotes per harmony and the respective fitness function f(nx) for a length of HSM harmonies. HM =          n11 n12 ··· n1xf(n1) n21 n22 ··· n2xf(n2) n31 n32 ··· n3xf(n3) . . .. . ..... . .. . . nHSM1nHSM2··· nHSMx f(nHSM )          (3.1) Improvisation. This is a process wherein a new HM is generated by applying consecutively three different probabilistic parameters that modifies the initial HM. These probabilistic parameters are: HMCR (Harmony Memory Considering Rate), PAR (Pitch Adjusting Rate) 61 Chapter 4. Case study: A cold-stamping press for bolt manufacturing Figure 4.2: Wire rod cold-stamping process. and malleable, such as low alloy steel, aluminium alloys (preferably magnesium alloys without copper), brass, silver and gold. Commonly, different elements are added to the steel contributing to the final characteristics of the material to be used. Some of the elements most commonly employed in alloys are [121]: •Copper. It improves the corrosion resistance of the material [122]. •Nickel. It reduces hardening and distortion temperature when hardened. The nickel alloy extends the critical temperature level, without carbides or oxides. This increases strength without decreasing ductility [122]. •Chromium. It increases hardenability and improves wear and corrosion resistance. It results in the creation of chromium carbides that are very hard. However, they are more ductile than a steel of the same hardness produced simply by increasing its carbon content. The addition of chromium extends the critical temperature range of the steel [122]. 68 4.1. Cold-stamping process Figure 4.3: Examples of sheet metal cold-stamping pieces. Figure 4.4: Examples of wire rod cold-stamping pieces. 4.1.2 Parts of a cold forming machine for bolt manufacturing. As mentioned before, the case under study is a stamping process that works with a wire rod placed on a coil as raw material and is commonly used for the manufacture of bolts and nuts. The machinery that develops this work is composed mainly of 3 blocks as Figures 4.5 and 4.6 show. 69 Chapter 4. Case study: A cold-stamping press for bolt manufacturing •Die-holding block or Transfer: This is the part in which the dies are housed, as well as the entire set of tools of the cold press. It has a cylindrical shape, so that the dies are adjusted to the diameter of the workpiece. This block is mainly composed of two parts: a module (commonly known as a transfer) which can be moved and folded down, and a lower part where the dies are housed. By moving the transfer, the dies can be accessed to change or repair them. This part is also responsible for moving the material during stamping by means of cams and different mechanisms. •Riveting pin block: This block, unlike the die set, consists of a single piece. The clamping system is similar to that of the inner block. At the back of the piece there are wedges, which allow the user to modify the position of each riveting pin independently of the others. Riveting pin block Die holding block Riveting pin Die Figure 4.5: Die holding and riveting pin blocks in cold stamping machine. •Feed rollers and straighteners: These rollers have the function of dragging the material (wire rod). They are mounted in such a way that the material is held under pressure between the two rollers and when the rollers rotate, the material advances. Prior to passing through the rollers, the material is passed through a 70 4.1. Cold-stamping process straightener so as not to encounter any problems when pulling the material. The rollers have a groove of the diameter of the material in which the wire fits. In order to avoid marks on the material, the edges between the groove and the outside are rounded. Feed rollers and straighteners raw material Figure 4.6: Feed rollers and straighteners. 4.1.3 Steps in the cold-forming process for bolts manufacturing. As Figure 4.7 shows, the bolt manufacturing process consists of three main stages. First, the material, which is in the form of wire on a coil, is fed into the machinery through the rollers and straighteners. Once the wire enters, it first goes through the cutting process, where the parameters are set to cut it to the desired size or length. After cutting, the piece of material is transported until the first die and is then placed in the transfer system. It is then moved through continuous die and tool openings to displace and form the working metal into the desired product. The riveting pin (mobile tool in charge of striking the piece), pushes the material into the die, thus giving the shape of the die. Finally, after passing through all the stations of the process, the bolt is ready to fulfil its function. 71 Chapter 4. Case study: A cold-stamping press for bolt manufacturing The cold stamping process is a high-speed manufacturing process, so the temperature and pressure must be just right at each step. It can be used to reduce or increase the diameters and lengths of the raw material and can remove small amounts of material by punching and trimming. All of these cold stamping processes and operations are carried out in a continuous planned process [51]. 6 stages forming processCutting stage 4 intermediate stages Wire rod introduction Figure 4.7: Bolt cold forming process. 4.2 Cold forming database The use case employed in this Thesis performs the steps described above. In this particular case, as Figure 4.8 shows, the process consists of six stations where each one performs a different operation on the material. In order to collect data from the process, a series of piezoelectric load cell sensors monitoring the force of stamping have been installed at each of the six process stations. The process is monitored over the course of a year and the data is stored in text files, each of which contains records of the strength of the six signals in 5-minute periods (acquisition frequency 1 kHz). Finally, these files are stored in a data file. 72 4.2. Cold forming database Data Storage File system Raw Material 6 Stage cold stamping machine Piezoelectric load cell sensors Material introduction Cold stamping bolt forming Monitorization and Data adquisition Figure 4.8: Flow diagram of the cold forming process under study. 4.2.1 Data contained in the database Throughout the monitoring period, production stoppages have occurred for three different reasons: 1) machine stoppage, which can be caused by unexpected breakage or wear of different components, 2) voluntary stoppage, that are due to various parts replacement or scheduled maintenance or 3) stoppage due to brankamp, which is related to a condition monitoring system that exists in the production plant in which some limits in the effort curves are set so as under some specific rules the machine is stopped. For this case study, machine stoppage records will be used and analysed, which are those that appear unexpectedly due to breakage or wear and tear and therefore necessary to be detected in advance. Thus, it is intended to develop a predictive maintenance on these tools and thus anticipate the occurrence of breakage or wear that lead to a machine stoppage.. Furthermore, a document containing manually recorded data by the machine operator, hereafter named Stops Registers file (SR-file), is available to document and classify any production stoppages. The record includes the cause of the stoppage and the associated WO in which the machine is currently engaged. For the domain expert is particularly important to know in which WO each breakage-stop occurs for a proper analysis of the process. This means that the machine must work with certain operating parameters that characterise each WO. Specifically, the concepts that are registered in this file are the 73 Chapter 4. Case study: A cold-stamping press for bolt manufacturing followings: •Type of stoppage (voluntary, machine or bankramp). •Date and time of the recorded breakdown. •Reason for the stoppage depending on the type of the stoppage. •Tools involved. •Other observations. However, as discussed in the introduction, the reality in real databases is that they contain noise, outliers, and are partially labelled. This case study presents the following drawbacks: •In the manual recording of stoppages, data about WOs, the reason or element affected by the stop among other parameters are missing in some cases. This could be caused by different factors such as operator shift changes or human errors. Therefore it is a partially labelled database. •There is a time lag between the end of the time series representing a stop and their manual recording in the SR-file. Despite having access to a vast historical database, one of the main challenges faced is the matching between the monitored signal and its corresponding manual registration. As a result, the number of accurately recorded stoppages that can be correctly attributed to the specific WO and the corresponding broken or worn element has been reduced. As previously mentioned, there are six different types of signals belonging to the force records of the different stations. Figure 4.9 illustrates the shape of the time series resulting from monitoring two stamps. Considering each of the signals of two stamps, the Figures 4.10 to 4.15 are obtained. 74 4.2. Cold forming database Figure 4.9: Signals captured from the cold forming process in two stamps. Figure 4.10: Process signal 1. 75 Chapter 4. Case study: A cold-stamping press for bolt manufacturing Figure 4.11: Process signal 2. Figure 4.12: Process signal 3. 76 4.2. Cold forming database Figure 4.13: Process signal 4. Figure 4.14: Process signal 5. 77 Chapter 7. Conclusions and Future Work friendly and comprehensible solutions that incorporate various algorithms and features. Therefore, a comprehensive ML solution, that is generic and not problem-dependent, is easy to use and understand, provides optimal parameters related to faults, and can be self-parameterized with optimization techniques is needed. The solution should also search for the best set of heuristics or rules that predict failures rather than the solution directly, and be trained and tested on real datasets. This integrated solution would offer the optimal failure prediction parameters while being user-friendly and promoting user knowledge. 7.2 Contributions The key contribution of this Thesis is the creation of a hyperheuristic inspired methodology that leverages optimization algorithms and ML to predict failures in industrial systems within the context of Industry 4.0. This methodology is hyper-heuristic inspired in terms of: a) reducing the expert knowledge required, b) providing an easyto-use computer system, c) operating on a range of related problems rather than on one narrow class of problems. Chapter 3 thoroughly elucidates this methodology, which consists of two distinct main stages and successfully addresses the primary goal of the thesis, which is the development of a ML-based solution for predicting industrial system failures. As mentioned and described in Chapter 3 this methodology has 2 main objectives: 1) to automatically label PLD and 2) to obtain the best set of parameters that are predictors of a failure from a FLD in the context of Industry 4.0. In the task of labelling PLD this methodology has the following objectives: 1) being a user-friendly and autonomous system for labelling PLD, 2) evaluate and optimize the different clustering solutions by a new semi-supervised metric PSOM 204 7.2. Contributions and 3) provide the user with both the results in an understandable way and a trustworthiness metric. The methodology applied over this field, in this Thesis is called PLAHS. Regarding, the task of collective AD for failure prediction, the aim is to find the optimal set of parameters that are predictors for a specific type of failure that allows to reliably and robustly predict the failure parameters providing the final user with information to detect the anomaly and anticipate the failure. In this respect, the methodology name to this field is HIMAFP. In particular, the subsequent points describe the contributions of this Thesis, which are directly linked to the accomplished and proposed objectives. Contr.1 The thesis presents a comprehensive and detailed study of two specific areas. Firstly, the available techniques for labeling unlabeled samples of databases are investigated. This is particularly significant in the context of predictive maintenance, as having more information leads to more accurate ML-based models. To this end, a thorough survey of semi-supervised clustering-based techniques in the literature is conducted and a new taxonomy is proposed. Notably, the evaluation of the solutions based on whether they are heuristic, meta-heuristic, or hyper-heuristic, which has not been done in previous research, is part of the contributions of this thesis. Secondly, ML solutions for detecting collective anomalies in the context of Industry 4.0 are explored. The contribution is twofold: 1) proposing a classification of these solutions based on the LER to implement them, and 2) evaluating and classifying the solutions based on whether they are heuristic, meta-heuristic, or hyper-heuristic approaches. It should be noted that no prior research has considered this classification. This contributions accomplish the objective 1 (Obj. 1) proposed in Chapter 1. 205 Chapter 7. Conclusions and Future Work Contr.2 Through the design and development of the integral hyperheuristic inspired methodology, it has been possible to develop an extensive study not only about the extracted Fe from the TSM that may be predictors of failure, but also other predictor parameters, such as the Wz prior to the break where anomalous behaviour is found and the Th that detects it. To develop this study, the part of the integral methodology that has been designed is HIMAPF. In summary, the results obtained from the proposed hyperheuristic inspired methodology HIMAFP for collective AD and related fault prediction applied over a real database shows interesting results. The HIMAFP has demonstrated its usefulness in terms of: a) providing domain experts with valuable knowledge about the behaviour of significant process variables in real cases with lack of information, b) developing an easy-to-use userfriendly methodology, c) adapting to different types of failures and its corresponding anomalies. The study has involved testing six different types of breakages in a cold stamping process used for bolt manufacturing. The findings demonstrate that HIMAFP is effective in identifying the optimal combination of relevant features from the TSM and their associated thresholds, which can indicate the imminent occurrence of a failure. HIMAFP can also handle the dynamic behaviors of abnormal TSM Fe and Wz to obtain TTF information. This information is valuable for developing real-time AD systems and making important corrective, preventive, or prescriptive decisions in industrial processes. The study highlights the importance of obtaining information about features in both the time and frequency domains, as evidenced by the testing of 19 different features. The optimal results have been obtained using frequency domain features in 30% of the total harmonies, emphasizing their significance and validating their inclusion in the study. This contribution accomplish objectives 2 (Obj. 2) and 3 (Obj. 206 7.2. Contributions 3) proposed in Chapter 1. Contr.3 In the context of designing and developing a general solution that can work with both labelled and partially labelled databases, a solution capable of inferring knowledge from the labelled part of the databases to label unknown samples has been developed. Specifically, the design and development towards an intelligent hyper-heuristic inspired system (PLAHS) to autonomously label PLD, that enable to classify signals from industrial processes in the context of Industry 4.0, has been proposed and is part of the contributions of this Thesis. PLAHS is a complete system that combines a bootstrapping data processing technique, the HS meta-heuristic algorithm, and an ensemble voting method for partially labeled databases. This system is innovative for several reasons. Firstly, it provides a labeling solution for time series and multivariate data. Secondly, it integrates the HS algorithm to optimize the internal parameters of three commonly used clustering algorithms (K-means, DBSCAN, and HAC), while also delivering the best clustering and optimized solution for each case study. Thirdly, it introduces a new evaluation metric (PSOM) that considers both supervised and unsupervised parts, weighted by the percentage of labels. Fourthly, it is user-friendly, allowing users to input their dataset and easily understand the results. Finally, the system provides a novel trustworthiness metric (TM) to measure its behavior and the user’s level of confidence in the results obtained. PLAHS has demonstrated its ability to adapt, evaluate, and provide the best labelling solution for various types of data, including multivariate and time series data, and scenarios with different percentages of labelled data ranging from 15 to 90%. The system’s effectiveness has been shown through several experiments: 1) seven databases from the UCI repository have been used to test different percentages of knowledge and degrees of separability, 2) six databases from the UCR repository 207 Chapter 7. Conclusions and Future Work have been used to test different degrees of labelled elements and intrinsic separability and 3) a real case study involving a cold stamping press for bolt manufacturing has been conducted. In addition, this approach has demonstrated its hyper-heuristic inspiration in various ways: a) a low level of expert knowledge requirement, b) an easy to use system, c) its ability to operate on different types of problem. This contribution accomplish objective 4 (Obj. 4) proposed in Chapter 1. Contr.4 Throughout the development of this Thesis, different studies and analyses have been developed related to the evaluation metrics in the different types of learning within the ML techniques. In this way, the following experiments have been developed. In the first trials and studies, a system of optimization of clustering parameters and selection of the best clustering algorithms has been proposed in HSOCC, where the internal validation metrics are investigated and analyzed, specifically the well-known Silhouette metric. In the contribution and labeling proposal, PLAHS, an analysis and bibliographic search is performed on different evaluation metrics in semi-supervised environments. The general rule when developing labelling systems is to use external validation indices where a confusion matrix is constructed with the known samples and extrapolated to the global solution. Other metrics used in the literature measure the degree of cohesion of the groups that are formed. In the proposal and contribution of this Thesis, a dual metric called PSOM has been developed based on a supervised part evaluated by a proposed and novel metric called MRC and an unsupervised part evaluated by Silhouette. In PSOM both parts are weighted by the percentage of labels. With regard to the results in the experimentation developed in PLAHS, the proposed metrics to evaluate semi-supervised (PSOM) and supervised (MRC) environments have shown to be able to give 208 7.2. Contributions a correct and accurate assessment. In this way, by using the MRC metric and the F1-score of the supervised part, it has been possible to develop a trustworthiness metric (TM) that is not only able to give the end-user an evaluation of how reliable the proposed labeling solution is, but also, it is able to tentatively calculate the value of the F1-score for the unknown part. This correlation has an R2fit value for multivariate bases of 0.813 and for time series of 0.755 as the results in Chapter 5 reveal. No existing reference has been found proposing such a metric for assessing reliability in semi-supervised environments and its application to labeling, as determined through a comprehensive review of the state of the art literature on this topic. Finally, in fully supervised environments and in the context of the detection of collective anomalies in time series, there is a wide range of external validation metrics that report the degree of satisfaction of the solutions. Some of the most widely implemented are AUC ROC , Sensitivity, Specificity, Recall, Precision, or Accuracy among others. The analysis of these metrics has been performed in the HIMAPF methodology. As part of the contribution of this thesis, an evaluation metric based on AUC ROC , called AUC MOD , has been developed and validated. Specifically, for cases where it is important to know when the anomalies appear, since the AUC ROC does not take it into consideration. By contrast, as shown and mentioned, the proposed AUC MOD metric enhances the value of the HIMAFP for its application to the industry where the time of occurrence of an anomaly is vital to predict the failures. As a summary, this thesis has put forward four contributions with regards to evaluation and trustworthiness metrics. These include: 1) a semi-supervised evaluation metric called PSOM, 2) a supervised evaluation metric named MRC specifically designed for partially supervised environments, 3) supervised evaluation metrics intended for collective AD in the context of Industry 4.0 called AUC MOD, and 4) a trustworthiness metric (TM) 209 Chapter 7. Conclusions and Future Work tailored for semi-supervised environments. This contributions accomplish objectives 5 (Obj. 5) and 6 (Obj. 6) proposed in Chapter 1. 7.3 Results From the developments and contributions done in this Thesis, the following journal articles have been published: •Computers and Industrial Engineering (Q1). NavajasGuerrero, A., Manjarres, D., Portillo, E., & Landa-Torres, I. (2022). A hyper-heuristic inspired approach for automatic failure prediction in the context of industry 4.0. Computers & Industrial Engineering, 171, 108381. •Applied Soft Computing (Q1) “Minor revision”. Navajas Guerrero, A., Portillo, E., & Manjarres, D. PLAHS: A Partial Labelling Autonomous Hyper-Heuristic System for Industry 4.0 with Application on Classification of Cold Stamping Process. Available at SSRN 4288726. Additionally, the following conference paper and presentation have been made: •In 14th International Conference on Soft Computing Models in Industrial and Environmental Applications (SOCO 2019). Navajas-Guerrero, A., Manjarres, D., Portillo, E., & Landa-Torres, I. (2020). A novel heuristic approach for the simultaneous selection of the optimal clustering method and its internal parameters for time series data. In 14th International Conference on Soft Computing Models in Industrial and Environmental Applications (SOCO 2019) Seville, Spain, May 210 7.4. Future work 13–15, 2019, Proceedings 14 (pp. 179-189). Springer International Publishing. •I Congreso Anual de Estudiantes de Doctorado (CAED). In this case a presentation entitled “Automatic system for failure prediction in Industry ” was presented. 7.4 Future work This Thesis has made it possible to identify several potentially interesting areas for future research. The most significant ones are: •In Chapter 5 the selected clustering algorithms are among the most frequently utilized ones in the literature,i.e. K-Means, DBSCAN and HAC. As a potential avenue for future investigation, it would be valuable to expand the number of algorithms considered and cover a broader range of potential solutions. Furthermore, this system boasts the advantage of being easily extendable and generalizable in terms of its configurations. It could be worthwhile to incorporate a greater variety of preconfigured distance measures for cluster formation, such as cosine or Mahalanobis distances, among others. •Chapter 6 describes the part of the hyper-heuristic methodology that is dedicated to collective AD for predicting failures in industrial environments (HIMAPF), it would be interesting extend the proposed methodology by integrating in the optimisation process a sliding window that, with the initial parameters obtained in HIMAPF, can fine-tune these parameters. In this way, the objective is to find the new set of fine-tuned predictor parameters such as: the size of the sliding window (Ws) used to extract the feature (Fe) sequentially to obtain a time series of the characteristic (Fews), the size of the degradation window (Wz) where the expected anomaly prior to failure is located, and 211 Chapter 7. Conclusions and Future Work the thresholds (ThHand ThL) that detect anomalous behavior in FeW s within the degradation window (Wz). These fine-tuned parameters could make it possible to detect collective anomalies that predict failure in a more precise, adjusted, and accurate manner. •In the context of Chapter 6, it would be worth exploring a potential future line of research which focuses on identifying which set or sets of features extracted from the signals have the potential to predict faults. This could be achieved by conducting a study within the optimization methodology, aimed at determining which combinations of feature parameters work well together to provide an accurate prediction. Rather than simply providing a list of individual feature parameters, the study would seek to identify combinations of features that, when used together, have the potential to provide a more comprehensive and reliable fault prediction. This approach could be more effective in identifying faults early, before they develop into more serious problems, ultimately improving the overall reliability and efficiency of the system under consideration. 212 Appendices 213 Appendix A. Results of Chapter 6 0 5 10 15 20 25 30 Iterations 0.66 0.68 0.70 0.72 0.74 0.76 0.78 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.10: Metric evolution AUC MOD in Failure type B in in TSM4for all combination of HS operator values. 0 5 10 15 20 25 30 Iterations 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.11: Metric evolution AUC MOD in Failure type B in TSM5 for all combination of HS operator values. 220 A.1. Harmony Search operators analysis 0 5 10 15 20 25 30 Iterations 0.70 0.75 0.80 0.85 0.90 0.95 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.12: Metric evolution AUC MOD in Failure type B in TSM6 for all combination of HS operator values. A.1.3 Type C 0 5 10 15 20 25 30 Iterations 0.825 0.850 0.875 0.900 0.925 0.950 0.975 1.000 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.13: Metric evolution AUC MOD in Failure type C in TSM1 for all combination of HS operator values. 221 Appendix A. Results of Chapter 6 0 5 10 15 20 25 30 Iterations 0.825 0.850 0.875 0.900 0.925 0.950 0.975 1.000 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.14: Metric evolution AUC MOD in Failure type C in TSM2 for all combination of HS operator values. 0 5 10 15 20 25 30 Iterations 0.84 0.86 0.88 0.90 0.92 0.94 0.96 0.98 1.00 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.15: Metric evolution AUC MOD in Failure type C in TSM3 for all combination of HS operator values. 222 A.1. Harmony Search operators analysis 0 5 10 15 20 25 30 Iterations 0.84 0.86 0.88 0.90 0.92 0.94 0.96 0.98 1.00 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.16: Metric evolution AUC MOD in Failure type C in TSM4 for all combination of HS operator values. 0 5 10 15 20 25 30 Iterations 0.800 0.825 0.850 0.875 0.900 0.925 0.950 0.975 1.000 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.17: Metric evolution AUC MOD in Failure type C in TSM5 for all combination of HS operator values. 223 Appendix A. Results of Chapter 6 0 5 10 15 20 25 30 Iterations 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.18: Metric evolution AUC MOD in Failure type C in TSM6 for all combination of HS operator values. A.1.4 Type D 0 5 10 15 20 25 30 Iterations 0.80 0.85 0.90 0.95 1.00 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.19: Metric evolution AUC MOD in Failure type D in TSM1 for all combination of HS operator values. 224 A.1. Harmony Search operators analysis 0 5 10 15 20 25 30 Iterations 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.20: Metric evolution AUC MOD in Failure type D in TSM2 for all combination of HS operator values. 0 5 10 15 20 25 30 Iterations 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.21: Metric evolution AUC MOD in Failure type D in TSM3 for all combination of HS operator values. 225 Appendix A. Results of Chapter 6 0 5 10 15 20 25 30 Iterations 0.86 0.88 0.90 0.92 0.94 0.96 0.98 1.00 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.22: Metric evolution AUC MOD in Failure type D in TSM4 for all combination of HS operator values. 0 5 10 15 20 25 30 Iterations 0.64 0.65 0.66 0.67 0.68 0.69 0.70 0.71 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.23: Metric evolution AUC MOD in Failure type D in TSM5 for all combination of HS operator values. 226 A.1. Harmony Search operators analysis 0 5 10 15 20 25 30 Iterations 0.80 0.85 0.90 0.95 1.00 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.24: Metric evolution AUC MOD in Failure type D in TSM6 for all combination of HS operator values. A.1.5 Type E 0 5 10 15 20 25 30 Iterations 0.75 0.80 0.85 0.90 0.95 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.25: Metric evolution AUC MOD in Failure type E in TSM1 for all combination of HS operator values. 227 Appendix A. Results of Chapter 6 0 5 10 15 20 25 30 Iterations 0.88 0.90 0.92 0.94 0.96 0.98 1.00 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.26: Metric evolution AUC MOD in Failure type E in TSM2 for all combination of HS operator values. 0 5 10 15 20 25 30 Iterations 0.75 0.80 0.85 0.90 0.95 1.00 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.27: Metric evolution AUC MOD in Failure type E in TSM3 for all combination of HS operator values. 228 A.1. Harmony Search operators analysis 0 5 10 15 20 25 30 Iterations 0.825 0.850 0.875 0.900 0.925 0.950 0.975 1.000 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.28: Metric evolution AUC MOD in Failure type E in TSM4 for all combination of HS operator values. 0 5 10 15 20 25 30 Iterations 0.88 0.90 0.92 0.94 0.96 0.98 1.00 Fitness average of 10 montecarlo ['0.5', '0.1', '0.1'] ['0.5', '0.1', '0.3'] ['0.5', '0.3', '0.1'] ['0.5', '0.3', '0.3'] ['0.5', '0.5', '0.1'] ['0.5', '0.5', '0.3'] ['0.7', '0.1', '0.1'] ['0.7', '0.1', '0.3'] ['0.7', '0.3', '0.1'] ['0.7', '0.3', '0.3'] ['0.7', '0.5', '0.1'] ['0.7', '0.5', '0.3'] ['0.9', '0.1', '0.1'] ['0.9', '0.1', '0.3'] ['0.9', '0.3', '0.1'] ['0.9', '0.3', '0.3'] ['0.9', '0.5', '0.1'] ['0.9', '0.5', '0.3'] Figure A.29: Metric evolution AUC MOD in Failure type E in TSM5 for all combination of HS operator values. 229 Bibliography [6] V. Chandola, A. Banerjee, and V. Kumar, “Anomaly detection: A survey,” ACM computing surveys (CSUR), vol. 41, no. 3, pp. 1–58, 2009. [7] H. Lasi, P. Fettke, H.-G. Kemper, T. Feld, and M. Hoffmann, “Industry 4.0,” Business & information systems engineering, vol. 6, no. 4, pp. 239–242, 2014. [8] M. Breque, L. De Nul, A. Petridis et al., “Industry 5.0: Towards a sustainable, human-centric and resilient european industry,” 2021. [9] A. G. Gonzalez, D. R. Quinonero, and S. F. Vega, “Assessment of the degree of implementation of industry 4.0 technologies: Case study of murcia region in southeast spain,” Engineering Economics, vol. 32, no. 5, pp. 422–432, 2021. [10] Y. Koren, The global manufacturing revolution: product-processbusiness integration and reconfigurable systems. John Wiley & Sons, 2010. [11] R. N. Langlois and P. L. Robertson, “Explaining vertical integration: Lessons from the american automobile industry,” The Journal of Economic History, vol. 49, no. 2, pp. 361–375, 1989. [12] M. T. Okano, “Iot and industry 4.0: the industrial new revolution,” in International Conference on Management and Information Systems, vol. 25, 2017, p. 26. [13] X. Xu, Y. Lu, B. Vogel-Heuser, and L. Wang, “Industry 4.0 and industry 5.0—inception, conception and perception,” Journal of Manufacturing Systems, vol. 61, pp. 530–535, 2021. [14] P. K. R. Maddikunta, Q.-V. Pham, B. Prabadevi, N. Deepa, K. Dev, T. R. Gadekallu, R. Ruby, and M. Liyanage, “Industry 5.0: A survey on enabling technologies and potential applications,” Journal of Industrial Information Integration, vol. 26, p. 100257, 2022. [15] J. Gantz, D. Reinsel et al., “Extracting value from chaos,” IDC iview, vol. 1142, no. 2011, pp. 1–12, 2011. [16] J. Lee, B. Bagheri, and H.-A. Kao, “A cyber-physical systems architecture for industry 4.0-based manufacturing systems,” Manufacturing letters, vol. 3, pp. 18–23, 2015. 236 Bibliography [17] P. Flach, Machine learning: the art and science of algorithms that make sense of data. Cambridge university press, 2012. [18] R. Azuma, Y. Baillot, R. Behringer, S. Feiner, S. Julier, and M. Blair, “Recent advances in augmented reality. ieee computer graphics and applications,” IEEE Computer Graphics and Applications, vol. 21, no. 6, pp. 34–47, 2001. [19] A. K. Jardine, D. Lin, and D. Banjevic, “A review on machinery diagnostics and prognostics implementing conditionbased maintenance,” Mechanical systems and signal processing, vol. 20, no. 7, pp. 1483–1510, 2006. [20] T. P. Carvalho, F. A. Soares, R. Vita, R. d. P. Francisco, J. P. Basto, and S. G. Alcal´a, “A systematic literature review of machine learning methods applied to predictive maintenance,” Computers & Industrial Engineering, vol. 137, p. 106024, 2019. [21] J. Dalzochio, R. Kunst, E. Pignaton, A. Binotto, S. Sanyal, J. Favilla, and J. Barbosa, “Machine learning and reasoning for predictive maintenance in industry 4.0: Current status and challenges,” Computers in Industry, vol. 123, p. 103298, 2020. [22] M. Paolanti, L. Romeo, A. Felicetti, A. Mancini, E. Frontoni, and J. Loncarski, “Machine learning approach for predictive maintenance in industry 4.0,” in 2018 14th IEEE/ASME International Conference on Mechatronic and Embedded Systems and Applications (MESA). IEEE, 2018, pp. 1–6. [23] F. Arellano-Espitia, M. Delgado-Prieto, A.-D. Gonzalez-Abreu, J. J. Saucedo-Dorantes, and R. A. Osornio-Rios, “Deepcompact-clustering based anomaly detection applied to electromechanical industrial systems,” Sensors, vol. 21, no. 17, p. 5830, 2021. [24] P. Czop, G. Kost, D. S lawik, and G. Wszo lek, “Formulation and identification of first-principle data-driven models,” Journal of Achievements in materials and manufacturing Engineering, vol. 44, no. 2, pp. 179–186, 2011. [25] D. An, N. H. Kim, and J.-H. Choi, “Practical options for selecting data-driven or physics-based prognostics algorithms with reviews,” Reliability Engineering & System Safety, vol. 133, pp. 223–236, 2015. 237 Bibliography [26] G. Forestier and C. Wemmert, “Semi-supervised learning using multiple clusterings with limited labeled data,” Information Sciences, vol. 361, pp. 48–65, 2016. [27] X. J. Zhu, “Semi-supervised learning literature survey,” 2005. [28] E. Swana and W. Doorsamy, “An unsupervised learning approach to condition assessment on a wound-rotor induction generator,” Energies, vol. 14, no. 3, p. 602, 2021. [29] J. Chen, C. Lu, and H. Yuan, “Bearing fault diagnosis based on active learning and random forest,” Vibroengineering PROCEDIA, vol. 5, pp. 321–326, 2015. [30] N. Piroonsup and S. Sinthupinyo, “Analysis of training data using clustering to improve semi-supervised self-training,” Knowledge-Based Systems, vol. 143, pp. 65–80, 2018. [31] J. E. Van Engelen and H. H. Hoos, “A survey on semi-supervised learning,” Machine Learning, vol. 109, no. 2, pp. 373–440, 2020. [32] G. Kostopoulos, S. Karlos, S. Kotsiantis, and O. Ragos, “Semisupervised regression: A recent review,” Journal of Intelligent & Fuzzy Systems, vol. 35, no. 2, pp. 1483–1500, 2018. [33] E. Bair, “Semi-supervised clustering methods,” Wiley Interdisciplinary Reviews: Computational Statistics, vol. 5, no. 5, pp. 349–361, 2013. [34] Y. Qin, S. Ding, L. Wang, and Y. Wang, “Research progress on semi-supervised clustering,” Cognitive Computation, vol. 11, no. 5, pp. 599–612, 2019. [35] T. Van Craenendonck, W. Meert, S. Dumanˇci´c, and H. Blockeel, “Cobras ts: A new approach to semi-supervised clustering of time series,” in International Conference on Discovery Science. Springer, 2018, pp. 179–193. [36] N. Grira, M. Crucianu, and N. Boujemaa, “Unsupervised and semi-supervised clustering: a brief survey,” A review of machine learning techniques for processing multimedia content, vol. 1, pp. 9–16, 2004. [37] T. Yang, N. Pasquier, and F. Precioso, “Semi-supervised consensus clustering based on closed patterns,” Knowledge-Based Systems, vol. 235, p. 107599, 2022. 238 Bibliography [38] P. P lo´nski and K. Zaremba, “Full and semi-supervised k-means clustering optimised by class membership hesitation,” in International Conference on Adaptive and Natural Computing Algorithms. Springer, 2013, pp. 218–225. [39] S. Basu, M. Bilenko, and R. J. Mooney, “A probabilistic framework for semi-supervised clustering,” in Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining, 2004, pp. 59–68. [40] I. Stojanovi´c, I. Brajevi´c, P. S. Stanimirovi´c, L. A. Kazakovtsev, and Z. Zdravev, “Application of heuristic and metaheuristic algorithms in solving constrained weber problem with feasible region bounded by arcs,” Mathematical Problems in Engineering, vol. 2017, 2017. [41] S. Basu, A. Banerjee, and R. Mooney, “Semi-supervised clustering by seeding,” in In Proceedings of 19th International Conference on Machine Learning (ICML-2002. Citeseer, 2002. [42] L. Lelis and J. Sander, “Semi-supervised density-based clustering,” in 2009 Ninth IEEE International Conference on Data Mining. IEEE, 2009, pp. 842–847. [43] L. Zheng and T. Li, “Semi-supervised hierarchical clustering,” in 2011 IEEE 11th International Conference on Data Mining. IEEE, 2011, pp. 982–991. [44] M. Blocho, “Heuristics, metaheuristics, and hyperheuristics for rich vehicle routing problems,” in Smart Delivery Systems. Elsevier, 2020, pp. 101–156. [45] M. Abdel-Basset, L. Abdel-Fatah, and A. K. Sangaiah, “Metaheuristic algorithms: A comprehensive review,” Computational intelligence for multimedia big data on the cloud with engineering applications, pp. 185–231, 2018. [46] T. Dokeroglu, E. Sevinc, T. Kucukyilmaz, and A. Cosar, “A survey on new generation metaheuristic algorithms,” Computers & Industrial Engineering, vol. 137, p. 106040, 2019. [47] A. A. Bara’a, A. D. Abbood, A. A. Hasan, C. Pizzuti, M. AlAni, S. ¨ Ozdemir, and R. D. Al-Dabbagh, “A review of heuristics 239 Bibliography and metaheuristics for community detection in complex networks: Current usage, emerging development and future directions,” Swarm and Evolutionary Computation, vol. 63, p. 100885, 2021. [48] J. Xiong, X. Liu, X. Zhu, H. Zhu, H. Li, and Q. Zhang, “Semisupervised fuzzy c-means clustering optimized by simulated annealing and genetic algorithm for fault diagnosis of bearings,” IEEE Access, vol. 8, pp. 181 976–181 987, 2020. [49] R. Kothari and V. Jain, “Learning from labeled and unlabeled data using a minimal number of queries,” IEEE Transactions on Neural Networks, vol. 14, no. 6, pp. 1496–1505, 2003. [50] A. K. Alok, S. Saha, and A. Ekbal, “A new semi-supervised clustering technique using multi-objective optimization,” Applied Intelligence, vol. 43, no. 3, pp. 633–661, 2015. [51] A. Navajas-Guerrero, D. Manjarres, E. Portillo, and I. LandaTorres, “A hyper-heuristic inspired approach for automatic failure prediction in the context of industry 4.0,” Computers & Industrial Engineering, p. 108381, 2022. [52] E. Burke, G. Kendall, J. Newall, E. Hart, P. Ross, and S. Schulenburg, “Hyper-heuristics: An emerging direction in modern search technology,” in Handbook of metaheuristics. Springer, 2003, pp. 457–474. [53] A. Swiercz, “Hyper-heuristics and metaheuristics for selected bio-inspired combinatorial optimization problems,” Heuristics and Hyper-Heuristics-Principles and Applications, 2017. [54] E. K. Burke, M. Gendreau, M. Hyde, G. Kendall, G. Ochoa, E. ¨ Ozcan, and R. Qu, “Hyper-heuristics: A survey of the state of the art,” Journal of the Operational Research Society, vol. 64, no. 12, pp. 1695–1724, 2013. [55] E. K. Burke, M. Hyde, G. Kendall, G. Ochoa, E. ¨ Ozcan, and J. R. Woodward, “A classification of hyper-heuristic approaches,” in Handbook of metaheuristics. Springer, 2010, pp. 449–468. [56] J. H. Drake, A. Kheiri, E. ¨ Ozcan, and E. K. Burke, “Recent advances in selection hyper-heuristics,” European Journal of Operational Research, vol. 285, no. 2, pp. 405–428, 2020. 240 Bibliography [57] K. Chakhlevitch and P. Cowling, Hyperheuristics: Recent Developments, 06 2008, vol. 136, pp. 3–29. [58] M. Leng, X. Chen, L. Li, M. Leng et al., “K-means clustering algorithm based on semi-supervised learning,” 2008. [59] X. Wang, C. Wang, and J. Shen, “Semi–supervised k-means clustering by optimizing initial cluster centers,” in International conference on web information systems and mining. Springer, 2011, pp. 178–187. [60] J. Li, J. Sander, R. Campello, and A. Zimek, “Active learning strategies for semi-supervised dbscan,” in Canadian Conference on Artificial Intelligence. Springer, 2014, pp. 179–190. [61] V. Macario and F. d. A. de Carvalho, “An adaptive semisupervised fuzzy clustering algorithm based on objective function optimization,” in 2012 IEEE International Conference on Fuzzy Systems. IEEE, 2012, pp. 1–8. [62] H. Gan, Y. Fan, Z. Luo, R. Huang, and Z. Yang, “Confidenceweighted safe semi-supervised clustering,” Engineering Applications of Artificial Intelligence, vol. 81, pp. 107–116, 2019. [63] C. Chrysouli and A. Tefas, “Spectral clustering and semisupervised learning using evolving similarity graphs,” Applied Soft Computing, vol. 34, pp. 625–637, 2015. [64] A. Demiriz, K. Bennett, and M. Embrechts, “Semi-supervised clustering using genetic algorithms,” Artif. Neural Netw. Eng, 09 1999. [65] S. Saha, A. Ekbal, and A. K. Alok, “Semi-supervised clustering using multiobjective optimization,” in 2012 12th International Conference on Hybrid Intelligent Systems (HIS). IEEE, 2012, pp. 360–365. [66] C. Ruiz, M. Spiliopoulou, and E. Menasalvas, “C-dbscan: Density-based clustering with constraints,” in International workshop on rough sets, fuzzy sets, data mining, and granularsoft computing. Springer, 2007, pp. 216–223. [67] K. Wagstaff, C. Cardie, S. Rogers, S. Schr¨odl et al., “Constrained k-means clustering with background knowledge,” in Icml, vol. 1, 2001, pp. 577–584. 241 Bibliography [68] M. Okabe and S. Yamada, “Clustering using boosted constrained k-means algorithm,” Frontiers in Robotics and AI, vol. 5, p. 18, 2018. [69] P. Rathore, J. C. Bezdek, P. Santi, and C. Ratti, “Conivat: Cluster tendency assessment and clustering with partial background knowledge,” arXiv preprint arXiv:2008.09570, 2020. [70] V. Melnykov, I. Melnykov, and S. Michael, “Semi-supervised model-based clustering with positive and negative constraints,” Advances in data analysis and classification, vol. 10, no. 3, pp. 327–349, 2016. [71] A. Vouros and E. Vasilaki, “A semi-supervised sparse k-means algorithm,” Pattern Recognition Letters, vol. 142, pp. 65–71, 2021. [72] S. Basu, A. Banerjee, and R. J. Mooney, “Active semisupervision for pairwise constrained clustering,” in Proceedings of the 2004 SIAM international conference on data mining. SIAM, 2004, pp. 333–344. [73] H. Guo, J. Ma, and Z. Li, “Active semi-supervised k-means clustering based on silhouette coefficient,” in International Conference on Intelligent and Interactive Systems and Applications. Springer, 2018, pp. 202–209. [74] D. Tiano, A. Bonifati, and R. Ng, “Featts: Feature-based time series clustering,” in Proceedings of the 2021 International Conference on Management of Data, 2021, pp. 2784–2788. [75] ——, “Feature-driven time series clustering.” in 24th International Conference on Extending Database Technology, EDBT 2021, 2021, pp. 349–354. [76] G. He, Y. Pan, X. Xia, J. He, R. Peng, and N. N. Xiong, “A fast semi-supervised clustering framework for large-scale time series data,” IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 51, no. 7, pp. 4201–4216, 2019. [77] J. Zhou, S.-F. Zhu, X. Huang, and Y. Zhang, “Enhancing time series clustering by incorporating multiple distance measures with semi-supervised learning,” Journal of Computer Science and Technology, vol. 30, no. 4, pp. 859–873, 2015. 242 Bibliography [78] H. A. Dau, N. Begum, and E. Keogh, “Semi-supervision dramatically improves time series clustering under dynamic time warping,” in Proceedings of the 25th ACM International on Conference on Information and Knowledge Management, 2016, pp. 999–1008. [79] J. Tan, W. Fu, K. Wang, X. Xue, W. Hu, and Y. Shan, “Fault diagnosis for rolling bearing based on semi-supervised clustering and support vector data description with adaptive parameter optimization and improved decision strategy,” Applied Sciences, vol. 9, no. 8, p. 1676, 2019. [80] A. Arriandiaga, E. Portillo, J. A. S´anchez, I. Cabanes, and I. Pombo, “Virtual sensors for on-line wheel wear and part roughness measurement in the grinding process,” Sensors, vol. 14, no. 5, pp. 8756–8778, 2014. [81] R.-M. Hage, I. Hage, C. Ghnatios, I. Jawahir, and R. Hamade, “Optimized tabu search estimation of wear characteristics and cutting forces in compact core drilling of basalt rock using pcd tool inserts,” Computers & Industrial Engineering, vol. 136, pp. 477–493, 2019. [82] A. Hatami-Marbini, S. M. Sajadi, and H. Malekpour, “Optimal control and simulation for production planning of network failure-prone manufacturing systems with perishable goods,” Computers & Industrial Engineering, vol. 146, p. 106614, 2020. [83] W. J. Lee, G. P. Mendis, M. J. Triebe, and J. W. Sutherland, “Monitoring of a machining process using kernel principal component analysis and kernel density estimation,” Journal of Intelligent Manufacturing, vol. 31, no. 5, pp. 1175–1189, 2020. [84] E. Portillo, M. Marcos, I. Cabanes, and D. Orive, “Real-time monitoring and diagnosing in wire-electro discharge machining,” The International Journal of Advanced Manufacturing Technology, vol. 44, no. 3-4, pp. 273–282, 2009. [85] M. Hu, Z. Ji, K. Yan, Y. Guo, X. Feng, J. Gong, X. Zhao, and L. Dong, “Detecting anomalies in time series data via a metafeature based approach,” IEEE Access, vol. 6, pp. 27 760–27 776, 2018. 243 Bibliography [86] A. H. Gandomi, X.-S. Yang, S. Talatahari, and A. H. Alavi, Metaheuristic applications in structures and infrastructures. Newnes, 2013. [87] M. A. D´ıaz-Cort´es, E. Cuevas, J. G´alvez, and O. Camarena, “A new metaheuristic optimization methodology based on fuzzy logic,” Applied Soft Computing, vol. 61, pp. 549–569, 2017. [88] E. ¨ Ozcan, B. Bilgin, and E. E. Korkmaz, “A comprehensive analysis of hyper-heuristics,” Intelligent data analysis, vol. 12, no. 1, pp. 3–23, 2008. [89] N. Laptev, S. Amizadeh, and I. Flint, “Generic and scalable framework for automated time-series anomaly detection,” in Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2015, pp. 1939– 1947. [90] W. Lu, Y. Li, Y. Cheng, D. Meng, B. Liang, and P. Zhou, “Early fault detection approach with deep architectures,” IEEE Transactions on Instrumentation and Measurement, vol. 67, no. 7, pp. 1679–1689, 2018. [91] H. Cai, J. Feng, J. Moyne, J. Iskandar, M. Armacost, F. Li, and J. Lee, “A framework for semi-automated fault detection configuration with automated feature extraction and limits setting,” in 2020 31st Annual SEMI Advanced Semiconductor Manufacturing Conference (ASMC). IEEE, 2020, pp. 1–6. [92] F. Calabrese, A. Regattieri, L. Botti, C. Mora, and F. G. Galizia, “Unsupervised fault detection and prediction of remaining useful life for online prognostic health management of mechanical systems,” Applied Sciences, vol. 10, no. 12, p. 4120, 2020. [93] J. A. Carino, M. Delgado-Prieto, J. A. Iglesias, A. Sanchis, D. Zurita, M. Millan, J. A. O. Redondo, and R. RomeroTroncoso, “Fault detection and identification methodology under an incremental learning framework applied to industrial machinery,” IEEE access, vol. 6, pp. 49 755–49 766, 2018. [94] H. Izakian and W. Pedrycz, “Anomaly detection in time series data using a fuzzy c-means clustering,” in 2013 Joint IFSA World Congress and NAFIPS Annual Meeting (IFSA/NAFIPS). IEEE, 2013, pp. 1513–1518. 244 Bibliography [95] L. Xie, S. H˚abrekke, Y. Liu, and M. A. Lundteigen, “Operational data-driven prediction for failure rates of equipment in safety instrumented systems: A case study from the oil and gas industry,” Journal of Loss Prevention in the Process Industries, vol. 60, pp. 96–105, 2019. [96] J. Yu, “Fault detection using principal components-based gaussian mixture model for semiconductor manufacturing processes,” IEEE Transactions on Semiconductor Manufacturing, vol. 24, no. 3, pp. 432–444, 2011. [97] M. A. Atoui, S. Verron, and A. Kobi, “Fault detection with conditional gaussian network,” Engineering Applications of Artificial Intelligence, vol. 45, pp. 473–481, 2015. [98] H. Chen, B. Jiang, W. Chen, and Z. Li, “Edge computingaided framework of fault detection for traction control systems in high-speed trains,” IEEE Transactions on Vehicular Technology, vol. 69, no. 2, pp. 1309–1318, 2019. [99] X. Li, X. Yang, Y. Yang, I. Bennett, and D. Mba, “A novel diagnostic and prognostic framework for incipient fault detection and remaining service life prediction with application to industrial rotating machines,” Applied Soft Computing, vol. 82, p. 105564, 2019. [100] J. Liu, Y.-F. Li, and E. Zio, “A svm framework for fault detection of the braking system in a high speed train,” Mechanical Systems and Signal Processing, vol. 87, pp. 401–409, 2017. [101] B. Luo, H. Wang, H. Liu, B. Li, and F. Peng, “Early fault detection of machine tools based on deep learning and dynamic identification,” IEEE Transactions on Industrial Electronics, vol. 66, no. 1, pp. 509–518, 2018. [102] D. Zang, J. Liu, and H. Wang, “Markov chain-based feature extraction for anomaly detection in time series and its industrial application,” in 2018 Chinese Control And Decision Conference (CCDC). IEEE, 2018, pp. 1059–1063. [103] J. Chen, W. Hu, D. Cao, B. Zhang, Q. Huang, Z. Chen, and F. Blaabjerg, “An imbalance fault detection algorithm for variable-speed wind turbines: A deep learning approach,” Energies, vol. 12, no. 14, p. 2764, 2019. 245