scieee AI-readable full text Open interactive document viewer

Contribution to reliable control of dynamic systems

Salazar Cortés, Jean Carlo

Abstract

Institutional repository that preserves and disseminates the academic and scientific output of the institution.

Full text

UNIVERSITAT POLITÈCNICA DE CATALUNYA Programa de Doctorat: AUTOMÀTICA, ROBÒTICA I VISIÓ Tesi Doctoral CONTRIBUTION TO RELIABLE CONTROL OF DYNAMIC SYSTEMS Jean Carlo Salazar Cortes Directors: Dr. Fatiha Nejjari Akhi-Elarab i Dr. Ramon Sarrate Estruch Març 2018 Abstract This thesis presents some contributions to the field of Health-Aware Control (HAC) of dynamic systems. In the first part of this thesis, a review of the concepts and methodologies related to reliability versus degradation and fault tolerant control versus health-aware control is presented. Firstly, in an attempt to unify concepts, an overview of HAC, degradation, and reliability modeling including some of the most relevant theoretical and applied contributions is given. Moreover, reliability modeling is formalized and exemplified using the structure function, Bayesian networks (BNs) and Dynamic Bayesian networks (DBNs) as modeling tools in reliability analysis. In addition, some Reliability Importance Measures (RIMs) are presented. In particular, this thesis develops BNs models for overall system reliability analysis by using Bayesian inference techniques. Bayesian networks are powerful tools in system reliability assessment due to their flexibility in modeling the reliability structure of complex systems. For the HAC scheme implementation, this thesis presents and discusses the integration of actuators health information by means of RIMs and degradation in Model Predictive Control (MPC) and Linear Quadratic Regulator algorithms. In the proposed strategies, the cost function parameters are tuned using RIMs. The methodology is able to avoid the occurrence of catastrophic and incipient faults by monitoring the overall system reliability. The proposed HAC strategies are applied to a Drinking Water Network (DWN) and a multirotor UAV system. Moreover, a third approach, which uses MPC and restricts the degradation of the system components is applied to a twin rotor system. i ii Finally, this thesis presents and discusses two reliability interpretations. These interpretations, namely instantaneous and expected, differ in the manner how reliability is evaluated and how its evolution along time is considered. This comparison is made within a HAC framework and studies the system reliability under both approaches. Keywords: prognostics and health-management, health-aware control, reliability analysis, reliability importance measures, Bayesian networks, Dynamic Bayesian Networks, model predictive control, linear quadratic regulator, drinking water networks, octorotor Resum Aquesta tesi presenta algunes contribucions al camp del control basat en la salut dels components "Health-Aware Control"(HAC) de sistemes dinàmics. A la primera part d’aquesta tesi, es presenta una revisió dels conceptes i meto-dologies relacionats amb la fiabilitat versus degradació, el control tolerant a fallades versus el HAC. En primer lloc, i per unificar els conceptes, s’introdueixen els conceptes de degradació i fiabilitat, models de fiabilitat i de HAC incloent algunes de les contribucions teòriques i aplicades més rellevants. La tesi, a més, el modelatge de la fiabilitat es formalitza i exemplifica utilitzant la funció d’estructura del sistema, xarxes bayesianes (BN) i xarxes bayesianes dinàmiques (DBN) com a eines de modelat i anàlisi de la fiabilitat com també presenta algunes mesures d’importància de la fiabilitat (RIMs). En particular, aquesta tesi desenvolupa models de BNs per a l’anàlisi de la fiabilitat del sistema a través de l’ús de tècniques d’inferència bayesiana. Les xarxes bayesianes són eines poderoses en l’avaluació de la fiabilitat del sistema gràcies a la seva flexibilitat en el modelat de la fiabilitat de sistemes complexos. Per a la implementació de l’esquema de HAC, aquesta tesi presenta i discuteix la integració de la informació sobre la salut i degradació dels actuadors mitjançant les RIMs en algoritmes de control predictiu basat en models (MPC) i control lineal quadràtic (LQR). En les estratègies proposades, els paràmetres de la funció de cost s’ajusten uti-litzant els RIMs. Aquestes técniques de control fiable permetran millorar la disponibilitat i la seguretat dels sistemes evitant l’aparició de fallades a través de la incorporació d’aquesta informació de la salut dels components en l’algoritme de control. Les estratègies de HAC proposades s’apliquen a una xarxa d’aigua potable (DWN) i a un sistema UAV multirrotor. A més, un tercer enfocament fent servir la degradació dels iii iv actuadors com a restricció dins l’algoritme de control MPC s’aplica a un sistema aeri a dos graus de llibertat (TRMS). Finalment, aquesta tesi també presenta i discuteix dues interpretacions de la fiabilitat. Aquestes interpretacions, nomenades instantània iesperada, difereixen en la forma en què s’avalua la fiabilitat i com es considera la seva evolució al llarg del temps. Aquesta comparació es realitza en el marc del control HAC i estudia la fiabilitat del sistema en tots dos enfocaments. Resumen Esta tesis presenta algunas contribuciones en el campo del control basado en la salud de los componentes “Health-Aware Control” (HAC) de sistemas dinámicos. En la primera parte de esta tesis, se presenta una revisión de los conceptos y metodologías relacionados con la fiabilidad versus degradación, el control tolerante a fallos versus el HAC. En primer lugar, y para unificar los conceptos, se introducen los conceptos de degradación y fiabilidad, modelos de fiabilidad y de HAC incluyendo algunas de las contribuciones teóricas y aplicadas más relevantes. La tesis, demás formaliza y ejemplifica el modelado de fiabilidad utilizando la función de estructura del sistema, redes bayesianas (BN) y redes bayesianas diná-micas (DBN) como herramientas de modelado y análisis de fiabilidad como también presenta algunas medidas de importancia de la fiabilidad (RIMs). En particular, esta tesis desarrolla modelos de BNs para el análisis de la fiabilidad del sistema a través del uso de técnicas de inferencia bayesiana. Las redes bayesianas son herramientas poderosas en la evaluación de la fiabilidad del sistema gracias a su flexibilidad en el modelado de la fiabilidad de sistemas complejos. Para la implementación del esquema de HAC, esta tesis presenta y discute la integración de la información sobre la salud y degradación de los actuadores mediante las RIMs en algoritmos de control predictivo basado en modelos (MPC) y del control cuadrático lineal (LQR). En las estrategias propuestas, los parámetros de la función de coste se ajustan utilizando las RIMs. Estas técnicas de control fiable permitirán mejorar la disponibilidad y la seguridad de los sistemas evitando la aparición de fallos a través de la incorporación de la información de la salud de los componentes en el algoritmo de control. Las estrategias de HAC propuestas se aplican a una red de agua potable (DWN) y a v vi un sistema UAV multirotor. Además, un tercer enfoque que usa la degradación de los actuadores como restricción en el algoritmo de control MPC se aplica a un sistema aéreo con dos grados de libertad (TRMS). Finalmente, esta tesis también presenta y discute dos interpretaciones de la fiabilidad. Estas interpretaciones, llamadas instantánea yesperada, difieren en la forma en que se evalúa la fiabilidad y cómo se considera su evolución a lo largo del tiempo. Esta comparación se realiza en el marco del control HAC y estudia la fiabilidad del sistema en ambos enfoques. Acknowledgements This thesis was carried out at the Research Center for Supervision, Safety and Automatic Control (CS2AC) with the financial support of Ministerio de Economía y Competitividad from Spanish Government through the grant BES-2012-059571. Firstly, I want to express my sincere gratitude to my supervisors, Prof. Fatiha Nejjari and Prof. Ramon Sarrate, I have learned a lot from them, and also I had an excellent environment for working, learning and communicating my ideas. I offer them my absolute thankfulness for their support and the confidence they have put on me, for their ideas, suggestions, and words of encouragement whenever I needed them. I also want to thank Prof. Vicenç Puig and Prof. Joseba Quevedo for their support and for offering me the opportunity to carry out this thesis. I would like to thank Prof. Louise Travé and Prof. Christophe Simon for the valuable time spent reading and reviewing this dissertation. I would also like to thank them, together with Prof. Teresa Escobet for accepting to be part of the thesis examination panel. During the thesis, I have had the privilege of working at the Centre de Recherche en Automatique de Nancy, where I spent some months carrying out two research visits. My sincere thanks to Prof. Didier Theilliol and Prof. Philippe Weber for all the fruitful scientific discussions, their advice, and their constant good mood. A Big Thanks to my family, who despite the distance, has given me their support and words of encouragement. A special thanks to my mother Julia, my wife Sofía for their continuous encouragement and care, and my sister Marisol for given me her support when I needed it the most. I would also like to thank my fellows from CS2AC, who have contributed to making these years enjoyable. I especially estimate the break moments having a coffee or chocolate at the bar, the Xocomàtiques, and the excursions to the mountain. A special recognition to vii viii the administrative staff for their support, their professionalism helped me to make my development as a Ph.D. student much easier. Moltes gràcies! List of acronyms and notation Acronyms AFTC Active Fault-Tolerant Control AFTM Accelerated Failure Time Model AHM Additive Hazard model ANN Artificial Neural Network CBM Condition-Based Maintenance cdf Cumulative distribution function CIF Critical Importance Factor CPT Conditional Probability Table DBN Dynamic Bayesian Network DFR Decreasing-Failure-Rate DIF Diagnostic Importance Factor DWN Drinking Water Network FDI Fault Detection and Isolation FEM Finite Element Method FFNN Feedforward Neural Network FTC Fault-Tolerant Control FV Fussell-Vesely importance measure HAC Health-Aware Control HMM Hidden Markov Models IFR Increasing-Failure-Rate ISE Integrated Square Error JCAR Joint Cumulative Actuator Reliabilities index LMI Linear Matrix Inequalities LP Linear Programing LRM Logistic Regression Model xv List of acronyms and notation xvi MC Markov chain MIF Marginal Importance Factor MPC Model Predictive Control MTBF Mean Time Between Failures MTTF Mean Time To Failure MTTR Mean Time To Repair NLMPC Nonlinear Model Predictive Control pdf Probability density function PFTC Passive Fault-Tolerant Control PHM Prognostics and Health Management PrHM Proportional Hazard model PIM Proportional Intensity Model QP Quadratic Programing RAW Reliability Achievement Worth RIM Reliability Importance Measure RAW Reliability Reduction Worth RScum Cumulative system reliability for the instantaneous reliability interpretation Rkf Scum Cumulative system reliability for the expected reliability interpretation RUL Remaining Useful Life SIM Structural Importance Measure SOM Self-Organizing Map SSI Stress-Strength Interference model SVM Support Vector Machine TRMS Twin-Rotor MIMO System UAV Unmanned Aerial Vehicles Ucum Cumulative control effort WPHM Weibull Proportional Hazard Model Notation IβIdentity matrix of size β×β [Λ]βBlock column matrix composed by β×1blocks of matrices Λ List of acronyms and notation xvii TΛ βBlock lower triangular matrix composed by β×βblocks of matrices Λ βShaper parameter (Weibull and Gamma distribution) h0Baseline failure rate Bmr Viscous friction coefficient of the main propeller uiLower bound of the control input for the iactuator uiUpper bound of the control input for the iactuator Btr Viscous friction coefficient of the tail propeller Dn Component/system state is down or failed εWeighting parameter of cost function objectives ηScale parameter (Weibull distributions) kfvn Aerodynamic force coefficient of the main rotor for negative ωv ΓComplete Gamma function Γ(β)=(β−1)! (Gamma distribution) γColumn vector of regression coefficients HcControl horizon HpPrediction horizon [Λ]T βBlock row matrix composed by 1×βblocks of matrices Λ αDiagonal weighting matrix of elements αi δDiagonal weighting matrix of elements δi ρDiagonal weighting matrix of elements ρi ˜αDiagonal weight matrix of elements α ˜ δDiagonal weight matrix of elements δ ˜ρDiagonal weight matrix of elements ρ Yref Output set point IBVariable representing the Birnbaum measure IβIdentity matrix of size β×β TΛ βBlock lower triangular matrix composed by β×βblocks of matrices Λ ICIF Variable representing Critical Importance Factor IDIF Variable representing Diagnostic Importance Factor Ic DIF Variable representing DIF using minimal cut sets Ip DIF Variable representing DIF using minimal path sets IRAW Variable representing Reliability Achievement Worth IRRW Variable representing Reliability Reduction Worth List of acronyms and notation xviii Jmr Moment of inertia main DC motor Jtr Moment of inertia in tail motor k1Input constant of the tail motor k2Input constant of the main motor kah/vϕh/v Physical constant kchn Cable force coefficient for negative θh kchp Cable force coefficient for positive θv kfhn Aerodynamic force coefficient of the tail rotor for negative ωh kfhp Aerodynamic force coefficient of the tail rotor for positive ωh kfvp Aerodynamic force coefficient of the main rotor for positive ωv kgGyroscopic constant kmPositive constant koh Horizontal friction coefficient of the beam subsystem kov Vertical friction coefficient of the beam subsystem ktPositive constant kth Drag friction coefficient of the tail propeller ktv Drag friction coefficient λiFailure rate of the ith component Lah/av Armature inductance of tail / main motor λExponential and Gamma distribution parameter, failure rate and inverse scale respectively [Λ]βBlock column matrix composed by β×1blocks of matrices Λ [Λ]α×βBlock matrix composed by α×βblocks of matrices Λ lbLength of counter-weight beam lcb Distance between the counterweight and the joint lmLength of main part of the beam ltLength of tail part of the beam (·)−1Inverse of matrix (·)TTranspose of matrix mbMass of the counter-weight beam mcb Mass of the counter-weight Ckkth minimal cut set List of acronyms and notation xix TMMission time mmMass of main part of the beam mms Mass of the main shield Pssth minimal path set mmr Mass of the main DC motor mtMass of the tail part of the beam mts Mass of the tail shield µLocation parameter (Lognormal distribution) ΩhAngular velocity of the TRMS around the vertical axis ωhRotational velocity of the tail rotor ΩvAngular velocity of the TRMS around the horizontal axis ωvRotational velocity of the main rotor Φcdf for the standard normal distribution (Lognormal distribution) piProbability of being up qiProbability of being down Pr(A) Probability of event A ψ(·)Covariate function Rah/av Armature resistance of tail / main motor RiReliability of the ith component rms Radius of the main shield θhRevolutions per minute Rs(TM)System reliability at the end of the mission RsReliability of the system rts Radius of the tail shield TsSampling time σScale parameter (Lognormal distribution) Φ(·)Structure function SRandom variable representing the state of the system TLifetime of a device θ0 vEquilibrium pitch angle corresponding to uv= 0.2753V θhYaw angle of the beam θvPitch angle of the beam List of acronyms and notation xx tTime mtr Mass of the tail DC motor uhInput voltage of the tail motor Up Component/system state is up or functioning uvInput voltage of the main motor zRow vector of covariates Z(t)Degradation process Zth Degradation failure threshold Chapter 1 Introduction 1.1 Background and motivation 1.1.1 Degradation vs. Reliability The degradation of physical components in engineering systems is generally inevitable, in view of factors such as wear due to usage, aging of the materials and hostile environmental conditions. In particular, the degradation of actuators in a closed-loop control system can lead to poor performance and sometimes to a loss of controllability when the level of degradation increases and the actuator reduces its capabilities such as speed response, force, pressure, strength, etc; or becomes prone to faults when its degradation level reaches or goes beyond a certain safe level, known also as failure threshold or critical value [185]. There exist two types of degradation: natural and forced degradation. On the one hand, natural degradation is an ageor time-dependent internal process where components gradually degrade, leading to their failure or breakdown. On the other hand, the forced degradation is artificially induced by an external agent, where component loading gradually increases in response to an increased demand [16, p. 32, 37, 52]. Such degradation can be characterized into three categories: binary, degradation with a finite number of levels, and degradation with an infinite number of levels [105]. The health of system components is of primordial importance for the safety and reliability of the controlled system. Reliability prediction based on degradation modeling can be an efficient method to estimate the health for some highly reliable components or systems where observations of failures are uncommon. Thus, to avoid failures it is important to enhance system safety by taking into consideration the degradation of components health in the controller design [73, 124]. As fault and failure concepts are used in different fields such as reliability, safety, and 1 Chapter 1. Introduction 2 fault-tolerant systems, in different technological areas, their terminological use is not uniform. In this thesis, the definition given by [60] is used. Definition 1.1. Fault: A fault is an impermissible deviation of at least one characteristic property (feature) of the system from the acceptable, usual, standard condition. Definition 1.2. Failure: A failure is a permanent interruption of a system ability to perform a required function under specified operating conditions. Recently, the interest of the research on performance degradation in control system design has increased [22, 41, 126, 189]. If the design objective is still to maintain the original system performance, this may force the remaining components to work beyond their normal service level to compensate for the handicaps caused by the degraded ones. Therefore, the trade-off between achievable performance and available actuator capability and their importance for the reliability of the system should be carefully considered in all control designs [146]. System components can experience physical degradation during their operation, and the severity of such degradation is related to the total operating life of the component. In this context, it may be interesting to design controllers that can exploit the available knowledge of the degradation dynamics to maintain adequate performance and extend the useful life of the components. Some assumptions are usually taken into account in order to model the degradation of a component. For instance, in [56, 74, 75, 143–146] it is assumed that the degradation is proportional to the control effort of actuators and it modifies the failure rate of each actuator. More accurate assumptions can also be taken, for instance in [139, 141] the degradation is assumed to be dependent not only on the load but also on the time and environmental conditions. Constraints could be imposed to ensure that the cumulative degradation will be acceptable at the end of the maintenance horizon [124]. Higher levels of availability and reliability are important objectives for the design of most modern engineering systems and recently, a growing interest to model the reliability of complex industrial systems by means of Bayesian Networks (BN) has appeared [174] and some works on BN and system safety have been developed [178, 180, 181]. This can be performed by redistributing the control effort among the available components or actuators to alleviate the work load and the stress factors on the equipment with worst conditions to avoid their break down. For this purpose, an appropriate policy should be developed to redistribute this effort until maintenance actions can be taken. Such policy could be defined in terms of remaining useful life, reliability, degradation, structure importance, aging, among others [10, 74]. The application of BNs to reliability started at the end of 90’s. In [163] the authors present Chapter 1. Introduction 3 the advantages of BNs in comparison with Reliability Block Diagrams (RBD). In [15] the authors propose to model a fault tree using a BN. Reliability is the ability of a system to operate successfully long enough to complete its assigned mission under stated conditions. It can be modeled as an exponential function [43, 182], a Weibull function [9, 67] or a Gamma function [84, 95, 112], among others. Reliability can also be expressed as a stochastic process [117]. For example, it is common to use Markov Chains (MC) to model the reliability of components [118]. Unfortunately, in practice the complexity of the system leads to a combinatorial explosion of states resulting in a MC with a very large size. The reliability information obtained with the MC is propagated to the system using a Dynamic Bayesian Network (DBN) which includes temporal information to calculate the impact of the component reliability on the system reliability [10]. The research in this field was mainly focused on improving the maintenance methodologies. However, the growing importance of maintenance has generated an increasing interest in the development and implementation of optimal maintenance strategies for improving system reliability, preventing the occurrence of system failures, and reducing maintenance costs of deteriorated systems [171]. The maintenance has been done traditionally based on one of two conceptions; preventive maintenance or corrective maintenance. Preventative maintenance performs regularly scheduled maintenance actions to maintain system in good conditions and avoid failures during service. Corrective maintenance leaves the system in operation until it fails and then takes restorative maintenance actions. In contrast, both of them have drawbacks. On one hand, preventive maintenance is expensive and the life cycle of system and components is not maximized. On the other hand, the corrective maintenance maximizes the life cycle of components but it has risks of damage to other components when failures occur. Whichever of the approach that is taken, unanticipated failures result in downtime of the system, and therefore there will always be a reactive maintenance needed. As a result, the system downtime will be as long as the necessary spare parts and personnel be available and the time necessary to carry out the maintenance task. ConditionBased Maintenance (CBM) has appeared as a new maintenance methodology which involves the real-time analysis of system condition or system health state and based on this information the maintenance tasks are performed. By contrast to preventive and corrective maintenance approaches, CBM has the potential Chapter 1. Introduction 4 for minimizing system failures incidents, reducing scheduled maintenance tasks, maximization of the life cycle of components, and increment of system availability. The technical capabilities to infer system and components condition in real-time from measurements of the process are critical to the success of a CBM implementation. The use of new technologies in the production systems allows to improve products quality, reduce costs and increase productivity. However, new technologies usually involve a high level of complexity and this complexity may result in more failure-prone systems. Fault-tolerant control (FTC) has emerged as a response to this problem [23]. Therefore, it is also possible to implement fault tolerant control (FTC) techniques whose objective is to allow system functioning in spite of having faulty components such as actuators or sensors [190]. Nevertheless, it would be interesting not to wait until a failure occurs but to anticipate and prevent them from happening, especially to avoid incipient faults which are difficult to detect. The detection of incipient faults is an open field of research due to the fact that some detection methods such as observers or parity equations tend to track the system even when there is a fault [39]. To achieve such an objective a Prognostics and Health Management (PHM) strategy is commonly used. The problematic of fault tolerant in systems has been widely treated by several authors. Now days it is possible to talk about Fault Detection and Isolation (FDI) as a field of research addressed to the problem of detection and localization of faults. In FTC and FDI research, works such as [13, 23, 44, 60–62] are obligatory reference due to the progress they have made. In recent years, HAC has been emerging as a new technique to handle this problem. It consists in taking predictive actions to prevent a failure occurrence, instead of FTC that takes actions after fault occurrence. Hence, this problem can be addressed taking into account the system to improve the system safety. HAC can extend the life time of the entire system or components and avoid failures until the next maintenance task. 1.1.2 Prognostics and Health Management Prognostics and Health-Management (PHM) involves the application of three concepts: diagnostics, prognostics and health management. The first one identifies the state of the system during its functioning, providing an accurate fault detection and isolation capability with low false alarm rate [122]. PHM provides system health information based on the evaluation of its reliability and/or its remaining useful life (RUL) which allows to make prediction of the advent of future Chapter 1. Introduction 11 Safety 167 (2017). Special Section: Applications of Probabilistic Graphical Models in Dependability, Diagnosis and Prognosis, pages 663 –672. J. C. Salazar, A. Sanjuan, F. Nejjari, and R. Sarrate. “Health-Aware and Fault-Tolerant control of an octorotor UAV system based on actuator reliability”. In: To be submitted to International Journal of Applied Mathematics and Computer Science (2018). J. C. Salazar, R. Sarrate, and F. Nejjari. “Health-Aware Control: A selective review and survey of current development”. In: To be submitted to Annual Reviews in Control (2018). • Book chapter J. C. Salazar, P. Weber, F. Nejjari, D. Theilliol, and R. Sarrate. “MPC Framework for System Reliability Optimization”. In: Advanced and Intelligent Computations in Diagnosis and Control. Ed. by Z. Kowalczuk. Advances in Intelligent Systems and Computing 386. Springer International Publishing, 2015, pages 161–177. • International conferences J. C. Salazar, R. Sarrate, and F. Nejjari. “Upper bound system reliability assessment for health-aware control of complex systems”. In: Proceedings of the 4th European Conference of the Prognostics and Health Management Society (PHME 2018). Utrecht, The Netherlands, 2018. J. C. Salazar, R. Sarrate, F. Nejjari, P. Weber, and D. Theilliol. “Reliability computation within an MPC health-aware framework”. In: Proceedings of the 20th World Congress of the International Federation of Automatic Control (IFAC 2017). Toulouse, France, 2017, pages 12230–12235. J. C. Salazar, A. Sanjuan, F. Nejjari, and R. Sarrate. “Health-Aware Control of an octorotor UAV system based on actuator reliability”. In: Proceedings of the 4th International Conference on Control, Decision and Information Technologies (CoDIT’17). Barcelona, Spain, 2017. J. C. Salazar, F. Nejjari, R. Sarrate, P. Weber, and D. Theilliol. “Reliability importance measures for a health-aware control of drinking water networks”. In: Proceedings of the 3rd Conference on Control and Fault-Tolerant Systems (SysTol’16). Barcelona, Spain, 2016, pages 572–578. J. C. Salazar, P. Weber, R. Sarrate, D. Theilliol, and F. Nejjari. “MPC design based on a DBN reliability model: Application to drinking water networks”. In: Proceedings of the 9th IFAC Symposium on Fault Detection, Supervision and Safety for Technical Processes (SAFEPROCESS 2015). Vol. 48. 21. Paris, France: IFAC, 2015, pages 688– Chapter 1. Introduction 12 693. J. C. Salazar, P. Weber, F. Nejjari, D. Theilliol, and R. Sarrate. “MPC Framework for System Reliability Optimization”. In: Proceedings of the 12th Diagnosis of Processes and Systems (DPS 2015). Ustka, Poland, 2015, pages 386. J. C. Salazar, F. Nejjari, and R. Sarrate. “Reliable control of a twin rotor MIMO system using actuator health monitoring”. In: Proceedings of the 22nd Mediterranean Conference of Control and Automation (MED’14). Palermo, Italy, 2014, pages 481–486. • Collaboration A. Soldevila, J. Cayero, J. C. Salazar, D. Rotondo, and V. Puig. “Control of a quadruple tank process using a mixed economic and standard MPC”. in: Actas de las XXXV Jornadas de Automática. Valencia, Spain, 2014.First place on the CEA contest 2014. 1.5.1 Research stays During the development of this thesis, two research stays were made in the Centre de Recherche en Automatique de Nancy at the Université de Lorraine with a duration of 4 months each one. The first one, from May to July 2014 where the work was concerned to the topic of reliable control of complex systems, the study of the mathematical modeling of the degradation/reliability using Bayesian approaches, and the investigation of the integration of this modeling in the control algorithm. And the second one, from May to July 2015 where the work was concerned with the topic of reliable control of complex systems, the study of reliability importance measures, the study of the integration of degradation/reliability model in the control algorithm, and also some ideas on the integration of reliability with linear matrix inequalities to compute the controller feedback gain were defined. 13 Part I Fundamentals Chapter 2 Prognostics and Health Management This chapter presents the concept of Prognostics and Health Management and offers a review of PHM methodologies based on the control techniques used to implement an HAC approach. 2.1 Introduction In industrial processes or dynamical systems health status of its components such as actuators or sensors is of primary concern as its failure will lead to immediate system shutdown or loss of performance in terms of economical cost or productivity. A wellmanaged system that minimizes the risk of failure is therefore desirable in many applications. The capability to accurately predict the health of system components (and consequently the system health itself) is the key to ensuring their dependability, availability, reliability, safety, and security. A system is said to be dependable when it is trustworthy enough to have confidence on the service that it gives. For a system to be dependable, it must be available and ready for use when is needed; reliable, when it is able to provide continuity of service while it is in use; safe, when it does not have a catastrophic consequence on the environment; and secure, when it is able to preserve confidentiality [147]. Thus, the concern about preserving the health of complex system has led to the development of different techniques such as Fault-Tolerant Control (FTC) and Prognostics and Health Management (PHM). Briefly, those control techniques which have the capacity to maintain the overall system stability and a satisfactory performance in presence of faults are called FTC. This means that a closed-loop system is fault tolerant if it is able to tolerate component malfunctions while maintaining a desirable performance and stability properties. FTC techniques can be classified into three groups: 14 Chapter 2. Prognostics and Health Management 15 • Hardware redundancies techniques: The hardware redundancy techniques try to achieve fault tolerance by taking advantage of the hardware redundancy in the system. Their main advantage is its simplicity but it implies a cost of redundant hardware and maintenance. • Analytical redundancy techniques: Passive Fault Tolerant Control (PFTC): The passive FTC techniques are control laws that take into account the fault appearance as a system perturbation. Therefore, within defined margins, the control law has built-in fault-tolerant capabilities, allowing the system to face fault occurrence, thanks to its robustness against a class of faults. The advantage of this approach is that it needs neither fault diagnosis nor controller reconfiguration, but as it needs to take in consideration all possible faults of a system during the design stage it provides limited fault tolerance capabilities, thus it cannot be guaranteed that unconsidered faults can be handled. Moreover, it entails a loss of performance with respect to the nominal case. • Analytical redundancy techniques: Active Fault Tolerant Control (AFTC): The active FTC techniques readjust the control law based on the information provided by the faults diagnosis module. With this information, some automatic adjustments in the control law are done after the fault occurrence attempting to meet the control objectives with the minimum performance degradation. Discussions on FTC are beyond the scope of this thesis and interested readers are referred to [13, 190] and the references therein which review the developments made in this area. PHM is a methodology aimed at handling the “System Health” understood as system reliability, remaining useful life (RUL), system degradation, etc, and based on a real-time monitoring and incipient fault detection. The main difference between both methodologies is that while FTC is applied when the fault has occurred, PMH is applied during the whole system functioning and its aim is to avoid or at least to delay the fault occurrence. In other words, FTC techniques do not provide an active reconfiguration of the control law given the health component state. In so far, as the application of PHM and the development of on-line prognostics techniques have evolved, a new type of FTC called Proactive FTC, which has two primary objectives: damage avoidance while ensuring primary mission success [160]. Chapter 2. Prognostics and Health Management 16 Proactive Fault-Tolerant Control is also called Health-Aware Control (HAC), and its functioning is as follows: given proper on-line prognostic information of the system, the HAC modifies the controller actions or reschedules the mission profile in order to preserve a high level of system health. HAC technique evaluates the health system while performing control over the system in a non-faulty situation. Moreover, it avoids faulty scenarios by mitigating the health degradation via appropriate control actions considering health indicators in the control objectives. This is done by constantly evaluating the system health indicators and making corrections through the control actions based on those indicators. A review on methodologies which combine the use of reliability and control theories, their origins, and their applications, based on a bibliography search related to the topics involved in HAC will be presented. Remark in Figure 2.1 the persistent increase in the works related to the topics of reliability and control, including fault-tolerant and health-aware techniques from 2000 to 2016. FIGURE 2.1: Amount of Publications Evolution. Therefore, it is clear that the interest on this field has been increasing. Those works are classified in the following fields: Engineering, Computer Sciences and Mathematics, Material Sciences, Chemical Engineering and Multidisciplinary as it is shown in Figure 2.2. Furthermore, it is worth to highlight the work that some research communities are doing in this field encouraging the development of new work and promoting different conferences and journals. To mention just a few, there is the Prognostics and Health Management Society, which supports a conference each year and maintains a journal. Also, the IEEE Reliability Society, which supports a conference and a journal on these topics. Chapter 2. Prognostics and Health Management 17 FIGURE 2.2: Related works by fields. In addition, these topics are being addressed in reputed conferences, such as the Mediterranean Control Conference (MED), the IFAC Symposium on Fault Detection, Supervision and Safety on Technical Process (SAFEPROCESS), the International Conference on Control and Fault-Tolerant Systems (Systol), the International Science Conference Diagnostics of Processes and Systems (DPS), the International Conference on Control, Decision and Information technologies (CoDIT), and the World Congress of the International Federation of Automatic Control among others. 2.2 Prognostics and Health Management The prognostics information is useful because it supplies the decision maker with adequate information about the expected time to system or components failure and allows to take the suitable actions to deal with them. Assessing the health of a system provides information that can be used to meet several critical goals: (1) providing advance warning of failures; (2) minimizing unscheduled maintenance, extending maintenance cycles, and maintaining effectiveness through timely repair actions; (3) reducing the life-cycle cost of equipment by decreasing inspection costs, downtime, and inventory; and (4) improving qualification and assisting in the design and logistical support of fielded and future systems [121]. In this sense, PHM is an emerging engineering discipline that links studies of system failure to system life cycle management [36]. The main aim of PHM is to improve safety and reduce maintenance cost. To achieve this Chapter 2. Prognostics and Health Management 18 objective some tasks, such as system monitoring, failure prognostics, and RUL computation, can be involved. Based on that, some actions like logistics requirements, maintenance performance, components replacement or controller reconfiguration, among others, should be taken in order to manage the health of the system. The steps involved in a PHM strategy are observation, analysis, and decision making (see Figure 2.3). Observation is the step where the data is acquired and processed. The second step consists in analyzing the data and extracting information of it, such as health state, RUL, to perform diagnosis and prognostics. And, the third step consists in taking the convenient action based on the analyzed data to extend the useful life of the system or components. PHM                                  Observation(Data acquisition Data processing Analysis       Health assessment Diagnosis Prognostics Decision making       Condition-Based Maintenance Logistic actions Health-Aware Control FIGURE 2.3: Steps involved in a PHM strategy. PHM concept has its origins in system engineering and considers aspects such as quality, reliability, and maintenance which are used to provide indications of anomalies and make predictions of future failures [40]. Research on PHM methodologies motivated by the benefits it brings has been increased considerably in the last decade. For instance, a literature review on prognostics can be found in [122], and a particular review of data-driven methods for PHM can be found in [165], a PHM review in manufacturing process can be found in [170]. A review on machinery PHM implementing Condition-Based Maintenance (CBM) which summarizes the recent research with emphasis on models, algorithms and technologies for data processing and maintenance decision-making is presented in [64]. In [48] a diagnosis and prognostic approach for power electronic drives and electric machines (AC/DC, DC/DC and DC/AC systems) is presented. This approach incorporates a low cost monitoring of the power electronics such as power MOSFETs and IGBTs. The proposed HAC strategy consists in reducing the performance of the control accomplishing the mission with reduced performance. Prognostics strategies can be implemented using different techniques. A classification Chapter 2. Prognostics and Health Management 19 of them can be found in [4, 36]. The one given in [4] is presented in Figure 2.4, which considers four categories: physical-based, data-driven, hybrid and experimental-based techniques. These categories are explained in detail below primarily based in [4, 36, 64]. FIGURE 2.4: Prognostics approaches Physical-based methodologies In the physical-based approach, the first principles are used to obtain accurate theoretical models specific for a particular type of component. In this category, there are the following approaches: physical model, cumulative damage, hazard rate and proportional hazard rate, nonlinear dynamics. The physical models are used to describe the physics of the system and failure modes, such as crack propagation, wear corrosion among others. These models combine systemspecific mechanical knowledge, defect growth equations, and condition monitoring to provide better prognostics output. These methodologies are more accurate than datadriven ones due that they contain a functional mapping of the system parameters. Using an explicit model of the system and residual generation methods such as Kalman filter, parameter estimation (or system identification) and parity relations, it is possible to obtain signals, called residuals, which are indicative of fault presence in the system. The residuals are used to implement fault detection, isolation and identification techniques. Chapter 2. Prognostics and Health Management 20 Model-based approaches can be more effective than other model-free approaches. However, a correct and accurate model is needed, and an explicit mathematical model may not be feasible for complex systems [64]. There are some applications of physical models, for instance in [90], the authors used the Paris’ law to model spur gear crack growth from an analysis of the stress and strain fields based on gear tooth load, geometry and material properties. In [116], the authors presented RUL computation approach based on a crack growth model and implemented using observers and intensity stress measures. In [87], RUL estimation approach for gears based on fatigue teeth crack, gear dynamics, and fracture models was proposed. Physics-based prognostics have been applied to systems in which their degradation phenomenon can be mathematically modeled such as in a gearbox prognostic module [17]. Physical modeling and parametric identification techniques have been applied with fault detection an failure prediction algorithms in order to predict the system time-to-failure. The faults and failure modes are traced back to physically meaningful system parameters, providing valuable diagnostic and prognostic information. In [102], a residual-based failure prognostic technique applied to a hydraulic system was proposed. The remaining system useful life is estimated based on residual signals, a bond graph model of the system dynamics, and the degradation model, which allows the prediction of the future health state of the system. Physical-based prognostics approaches are very effective and descriptive because system degradation modeling is based on laws of nature. Remark that the accuracy and precision depend on model fidelity [166]. There are some disadvantages and limitations of this approach such as: developing a high fidelity model for RUL estimation is very costly, time-consuming, and computationally costly and sometimes it cannot be obtained. Also, it will be component/system specific which limits its use to other similar cases. Hence, sometimes the data-driven approach is preferably used [36]. Data-driven methodologies Data-driven methods are based on the fact that condition monitoring data and the extracted features vary with the development of either the initiation and propagation process or the degradation process. These methods are useful when a large quantity of noisy data needs to be transformed into a logical information to estimate the RUL whose accuracy depends on the quantity and quality of the data. The data-driven prognostics methodology is based on statistical and learning techniques, most of which come from the theory of pattern recognition. Such approaches incorporate Chapter 2. Prognostics and Health Management 27 Robust control approaches are used for instance in [58], where the objective is to maintain stability and performance of the system near to the desired performance in the presence of system component faults, and in certain circumstances reduce the performance requirements to achieve the objective. In [75] and [74], the authors propose the integration of reliability and reconfigurability analysis in an FTC system for a tank system and aircraft model, respectively. These works have in common the use of components reliability as indexes to perform the control reconfiguration after the occurrence of failures. In [73], the authors propose an FTC system based on a feedback controller which guarantees the highest system reliability. This controller is synthesized using linear matrix inequalities (LMIs) and incorporating a reliability indicator, this reliability indicator is the well known Birnbaum measure which indicates those system components whose reliability are critical for the reliability of the system. Note that in the mentioned works, the asset reliabilities are modeled using the exponential distribution function. In [72], the authors propose a reconfigurable control allocation problem applied to an over-actuated system in which the redistribution factor is defined in terms of the actuator reliabilities modeled by using the Weibull distribution function. Then, the control allocation problem consists in assigning more control effort to those actuators whose reliabilities are higher and to relieve those actuators whose reliabilities are lower. In [10] and [9], a control allocation problem is solved by incorporating reliability importance measures to redistribute the control effort among the available actuators of a hydraulic system. The reliability is modeled using a Weibull distribution function and it is computed by a Dynamic Bayesian Network (DBN). Table 2.2 present a list of application made in this topic. For instance, in [124] a PHM scheme using MPC is presented. An application using a tank level over-actuated system is presented. The main idea is to manage the actuators degradation over the maintenance horizon to achieve the control goals using MPC. In [113] an approach to design a reliable admissible model matching (AMM) fault tolerant control (FTC) for LPV systems is proposed. The main idea is to reconfigure the controller on-line taking into account changes due to the faults, maintaining a certain actuators reliability level in spite of the faults. In [91] the authors present a reliable robust tracking controller against actuator faults and control surface impairment for aircraft bases in on a mixed linear-quadratic (LQ)/ H∞performance indexes and multiobjective optimization using linear matrix inequalities (LMIs). In such work, the concept reliable is used to mean trustworthy due to the fact that authors proposed a control approach with fault tolerance capabilities under actuators outages. Chapter 2. Prognostics and Health Management 28 TABLE 2.2: Application examples of HAC Applications References Aircraft Chamseddine, Theilliol, et al. [20], Khelassi, Jiang, et al. [72], Khelassi, Theilliol, et al. [73, 74], LI, Zhao, et al. [88], Oca and Puig [113], Salazar, Nejjari, et al. [137], Theilliol, Weber, et al. [162], and Weber, Boussaid, et al. [175] Tank level systems Abdel-Geliel, Badreddin, et al. [1], Bicking, Weber, et al. [9], Dardinier-Maron, Hamelin, et al. [31], Grosso, Ocampo, et al. [54], Guenab, Weber, et al. [58], Khelassi, Theilliol, et al. [75], Nguyen, Dieulle, et al. [109–111], Pereira, Galvao, et al. [124, 125], Robles, Puig, et al. [133], Salazar, Nejjari, et al. [136, 137], Salazar, Weber, et al. [143, 144], and Weber, Simon, et al. [178] Electromechanics systems Escobet, Puig, et al. [38], Ginart, Barlas, et al. [48], Gokdere, Chiu, et al. [51], Lee, Kim, et al. [85], and Tang, Kacprzynski, et al. [159] In [75] the authors propose the integration of reliability evaluation in a fault-tolerant control system and illustrate with a flight control application. The controller is analyzed with respect to reliability requirements and its controllability defined through its Gramian. The admissible solution is proposed according to reliability evaluation based on energy consumption under degraded functional conditions. In [132] the authors present an overview of modeling and control strategies including fault-tolerant capabilities for wind turbines and wave energy devices. In these systems, the reliability improvement is achieved by a significant reduction of periods of null or very low power production. 2.5 Conclusions In this chapter, a literature review on Prognostics and Health Management, particularly in the historical development of Health-Aware Control methodologies has been presented. Besides the attempts to address the problem of HAC, a list of applications and control techniques used were given. Approaches to include the system health information in the controller design has been proposed, but it is still an open research field to explore. The problem of how to include Chapter 2. Prognostics and Health Management 29 the health information in a systematic way in the controller design should be further investigated. The interaction between prognostics and control establishes a new feedback loop. To understand the effects of this loop in the whole performance of the system, a mathematical formulation of the problem should be considered. Due to the combined discrete-event/ continuous nature of the reconfiguration/accommodation actions and the control loop, respectively, techniques coming from hybrid system theory could be applied. The appropriate health indicator to be used for reconfiguring/accommodating the controller is also a key issue. Some methodologies for the designer should be given in order to facilitate the design of the HAC reconfiguration/accommodation strategy. Furthermore, model-based prognostic approaches accuracy depends on the availability of an accurate model and data. Data-driven methods are preferred in the case of quick estimations with lower accuracy. Nevertheless, physics-based approaches seem to be the most adequate when prognostics accuracy is needed and the data is limited. Chapter 3 Background on reliability In this chapter, a review of degradation and reliability concepts and their modeling approaches is presented. The research in degradation and reliability has been a topic of relevant interest since they provide an estimation of span life for systems and components. Therefore, a review of the most relevant of both, theoretical and applied contributions found in the literature is given. In Section 3.1, an introduction to reliability and degradation is given, and the most used probability distributions are described. Then, in Section 3.2, an explanation about covariates and how they can be modeled and integrated into the probability distributions is presented (Section 3.2.1). After, in Section 3.2.2, some models to fit degradation data, such as fitting data to probability distribution or using regressors, are presented. 3.1 Reliability and degradation Reliability analysis of dynamical systems helps to prevent failures which could be costly and sometimes disastrous. In complex systems or in safety-critical systems, it is imperative to identify the key component of the system and prevent the system failure. Generally, this is done by implementing Conditioned-Based Maintenance (CBM) methods, where decisions are supported by the reliability analysis information. In the literature, authors refer indistinctly to the concepts of reliability, degradation, deterioration, etc. In general terms, reliability leads to the concept of dependability, successful operation or performance, and the absence of failures, whereas unreliability (lack of reliability) leads to the opposite [14]. Thus, it is convenient now to give a clear definition of these concepts. Definition 3.1. Reliability is the probability that an asset will perform its functioning correctly for a specified period of time and under specified operating conditions [14]. 30 Chapter 3. Background on reliability 31 Whereas, degradation can be defined as: Definition 3.2. Degradation is the the reduction in performance, reliability and lifespan of assets [52]. Degradation can be viewed as a damage that the system accumulates over time and eventually leads to a failure when the accumulated damage reaches a failure threshold. The condition of a component can be characterized according to the degree of detail given to the degradation process [105]. It could be characterized as binary condition (see Figure 3.1(a)), being equal to 1 if the component is in its working state, i.e., the component performance is satisfactory or acceptable, and 0 for the opposite situation, where the component is in the faulty state. In this characterization, the component starts in the working state and changes to the failed one after a period of time (failure time). This is represented as a random variable because the time instant of change from working to failed is uncertain. An example of characterization is an electric bulb where its state changes from working to failed in a very short time which can be assumed to be instantaneous. (a) (b) (c) FIGURE 3.1: Component condition characterization: (a) binary states, (b) multi-state with finite states, (c) multi-state with infinite states. Chapter 3. Background on reliability 32 Also, it could be characterized with a finite number of states (see Figure 3.1(b)) where the condition of the component can assume any value from the set {1,2, . . . , K}with: • 1 corresponding to component performance being fully acceptable, i.e., the component is in the good working state. •i,1< i < K corresponding to component performance being partially acceptable, i.e., component is in a working state with a higher value of iimplying a higher level of degradation and, •Kcorresponding to component performance being unacceptable, i.e, the component is in a faulty state. The time to failure of the component is given by tF= inf{t:Z(t) = K}. An example of this characterization consider the wear in a tire, where no wear corresponds to state 1 and complete wear corresponds to state K. And finally, it could be characterized with an infinite number of levels (see Figure 3.1(c)) being an extension of the above case with K=∞. Here a higher value implies a higher degradation, and the component failure time is given by tF= inf{t:Z(t) = Zth}. Remark that, when the level of degradation reaches the threshold Zth the asset failure occurrence increases. As a consequence, a failure is often a result of the effect of degradation. From these two definitions, remark that reliability declines when assets degrade or deteriorate. The failure threshold provides a link between degradation and assets failure, therefore, it is possible to use the degradation signals to estimate the failure time distribution, the RUL, etc. The degradation signals are obtained by a proper degradation model, which consists in developing a good probability model that is capable of describing the degradation process. More specifically, reliability is the probability that a system or component will operate properly for a specific period of time under design operating conditions (such as temperature, volts, etc.) without failure. In other words, reliability can be used as a measure of the success of the system to provide its function adequately. Mathematically, reliability R(t)is the probability that a system will be successful in the interval from time 0to time t: R(t) = Pr(T > t)t≥0,(3.1) Chapter 3. Background on reliability 33 where Tis a random variable denoting the time-to-failure or failure time. Which is the time until the system first enter the down (failure) state: T= inf{t:system state =down},(3.2) And the unreliability F(t), which is a measure of failure, is defined as the probability that a system will fail by time t: F(t) = Pr(T≤t)∀t≥0,(3.3) in other words, F(t)is the cumulative distribution function (cdf), also called failure distribution function. If the time-to-failure random variable Thas a density function f(t), then the reliability function R(t)is defined as [127, p. 10]: R(t) = Zt 0 f(x)dx, (3.4) and the function: h(t) = f(t) R(t) =f(t) 1−F(t), (3.5) is called the failure or hazard rate [47, p. 25]. The term h(t)dt is the probability that a device at the age of twill fail in the time interval tto (t+dt). The importance of the hazard function lies in that it indicates the change in the failure rate over the life of a population of components by plotting their hazard functions on a single axis. In reliability theory, several types of probability distributions are used; for example, exponential, Weibull, gamma, lognormal, among others. A brief description of some of these distribution functions will be given below. The exponential distribution is one of the most widely used in reliability engineering because it is relatively easy to handle in performing reliability analysis, and many engineering items exhibit constant hazard rate during their useful life [32]. Its probability density function (pdf) is defined by f(t) = λe−λt t≥0, λ > 0,(3.6) where λis the distribution parameter which is also known as the constant failure rate. Chapter 3. Background on reliability 34 And the reliability function given by (3.4) and the exponential distribution (3.6), is defined as: R(t) = Z∞ t λe−λtdt =e−λt. (3.7) Reliability (3.7) refers to the probability that a device’s lifetime is larger than t, the probability that the device will survive beyond time t, or the probability that the device fail after time t. Remark that R(0) = 1 and R(∞) = 0 and that reliability function is a non increasing function of t. From (3.5) and using (3.6) and (3.7) it is evident that, for the exponential distribution, the failure rate function is the failure rate: h(t) = λe−λt e−λt =λ. (3.8) The exponential distribution has the memoryless property, which means that the current reliability status does not depend on the previous one. This situation does not represents the phenomenon of aging which is very important for reliability theory. Intuitively, aging represents an increase of failure risk as a function of time in use. To introduce this dependency, the following definition of reliability can be used [47]: Pr(T > t) = R(t) = e−Rt 0λ(v)dv.(3.9) In the case of the Weibull distribution which is used to represent several physical phenomena, its probability density function is defined by [179]: f(t) = βtβ−1 ηβexp −t ηβ ,(3.10) where βand ηare the shape and scale parameters, respectively. The popularity of this distribution stands on the fact that, depending on the parameters, it may describe both increasing and decreasing failure rates. The gamma distribution is especially useful for reliability modeling of those asset lifetimes which degradation can be explained by the shock accumulation [47]. The lognormal distribution is a continuous probability distribution of a random variable whose logarithm is normally distributed. The lognormal distribution is applied to the description of the dispersion of the component failure rate data. Chapter 3. Background on reliability 35 Table 3.1 presents some of the most widely distribution functions used in reliability theory and summarizes their pdf,cdf, reliability functions, and hazard rates. TABLE 3.1: Common used continuous distributions. Distribution Formulae References Exponential f(t) =λe−λt, t ≥0 F(t) =1 −e−λt (3.11) R(t) =e−λt, t ≥0 h(t) =λ Deloux, Castanier, et al. [33], Finkelstein [43], Guenab, Weber, et al. [58], Khelassi, Theilliol, et al. [73–75], Oliveira and Yoneyama [115], and Wu, Wang, et al. [182] Weibull f(t) =βtβ−1 ηβexp −t ηβ F(t) =1 −exp −t ηβ(3.12) R(t) = exp −t ηβ(3.13) h(t) =β ηt ηβ−1(3.14) Bicking, Weber, et al. [9], Grosso, Ocampo, et al. [54], Jiang and Jardine [67], Khelassi, Jiang, et al. [72], and Toscano and Lyonnet [164] Gamma f(t) = λβ Γ (β)tβ−1e−λt (3.15) F(t) =λβ ΓZt 0 xβ−1e−λxdx (3.16) R(t) = λβ Γ(β)Z∞ t xβ−1e−λxdx (3.17) h(t) = tβ−1e−λt R∞ 0xβ−1e−λxdx (3.18) Çinlar [19], Langeron, Grall, et al. [82], Lawless and Crowder [84], Lu and Meeker [95], and Noortwijk [112] Lognormal f(t) = 1 σt√2πexp "−(ln t−µ)2 2σ2#(3.19) F(t) =Φ ln t−µ σ(3.20) R(t) =1 −Φ(ln t−µ σ)(3.21) h(t) = f(t) 1−Φ [(ln t−µ)/σ](3.22) Chen and Zheng [24], Langeron, Grall, et al. [82], and Meeker, Escobar, et al. [103] From the reliability theory, another useful concepts like the Mean Time To Failure (MTTF) Chapter 3. Background on reliability 36 and Availability will be explained. The MTTF is the expected life of the device and it can be evaluated through the following standard equation [34]: MTTF =E(T) = Z∞ 0 tf(t)dt =Z∞ 0 R(t)dt . (3.23) For repairable devices, the MTTF represents the mean time to the first failure. After it is repaired and put into operation again, the average time to the next failure is indicated by the mean time between failures (MTBF). Under perfect repairs, MTBF is equal to MTTF. Since there is usually an aging effect in most devices, MTBF decreases as more failures are experienced by the device. The time needed to perform a repair is called mean time to repair (MTTR). Then, it is possible to define the availability as: Definition 3.3. Availability: The availability is defined as the probability that the device is available when is needed and is often used as a measure of its performance. It is expressed as: A=MTTF MTTF +MTTR.(3.24) The failure rate of many devices exhibits the bathtub curve shown in Figure 3.2 which is divided into three sections [78]: FIGURE 3.2: Bathtub curve of hazard rate. • In the interval (0, t1), which is usually short, a decreasing-failure-rate (DFR) is observed. This is often referred to as the early failure period and the failures that occur in this interval are called early failures, burn-in failures, or infant mortality failures. They are mainly due to manufacturing defects and can be eliminated using burn-in techniques. Chapter 3. Background on reliability 43 be described as: R(t) = Pr(T≥t) = Pr(X(t)> D) = exp −Dβ b e−at !,(3.39) where, aand bare constants, X(t)is degradation level at time t, and Dis the threshold level. Mixture model for hard and soft failures: this method consists in obtaining two observation samples: catastrophic failures (hard) and degradation (soft), and then builds a mixture model with both data [191]. This can be modeled as: F(t) = p(z)Fc(t) + [1 −p(z)]Fd(t),(3.40) where, F(t)is the cumulative density function (cdf) of t,p(z)is the proportion of components failed catastrophically within specified observed degradation value z,Fc(t)and Fd(t)are respectively the catastrophic failure cdf of tand degradation failure cdf of t(lifetime). Time series model: it is a technique used to predict individual system performance reliability in real-time considering multiple failure modes. It includes an on-line multivariate monitoring and anticipation (using a Kalman filter) of selected performance measures and conditional performance reliability estimates [96]. Other commonly used degradation models: Brownian motion (or Wiener process) and Gamma processes are continuous-time models that are appropriate for modeling continuous degradation processes. Nowadays, these models as well as Markov models are widely applied as degradation modeling techniques [14, 71, 89, 98, 112, 128, 153]. ii) Accelerated degradation models: accelerated degradation models provide information about reliability at normal conditions using degradation data obtained at an accelerated time with or without stress conditions. Degradation process can be very slow in an industrial application at normal stress level and have a high MTBF, which makes difficult to estimate the failure time distribution of components with high-reliability [103, 161, 185]. Therefore, to obtain data quickly from a degradation test, an accelerated life test is performed. This test consists in applying an increasing level of acceleration variables, such as vibration amplitude, temperature, corrosive media, load, voltage, pressure [150]. However, this kind of tests is a costly approach. Chapter 3. Background on reliability 44 General assumptions of accelerated degradation models are: (a) Degradation is not a reversible process. (b) A model applies to a single degradation process, mechanism, or failure mode. (c) Degradation of a test asset’s performance before the test starts is negligible. (d) The failure processes at higher stress levels are the same as at the design or use stress levels. Accelerated degradation models can be physics-based or statistical based. In the following, a quick review on those categories is presented. • Physics-based models: these models are used for accelerated life tests when the deterioration is caused by thermal and non-thermal parameters, speed, load, corrosive environment, vibration amplitude, etc., in order to estimate their service lives. Applications can include dielectrics, semiconductors, battery cells, lubricant, plastic, insulating fluids, capacitors, bearings, and spindles. Arrhenius model: this model is used when the damage is caused by temperature, especially for: dielectrics [49], semi-conductors [103, 123], battery cells, lubricant, and plastic.. In general, it is used to describe many products that fail as a result of degradation due to chemical reactions or metal diffusion. The nominal time τto failure is [108]: τ=A e E (k T),(3.41) where Eis the activation energy of the reaction in eV, kis the Boltzmann’s constant, 8.3171 ×105eV/K, Tis the absolute Kelvin temperature, Ais a constant that depends on product geometry, specimen size and fabrication, test method and others factors. Eyring model: this model is used for accelerated life test with respect to the thermal and non-thermal variable. The Eyring relationship for nominal life τas a function of absolute temperature Tis: τ=A Te B (k T),(3.42) here Aand Bare constant parameters of the product and test method, and k is Boltzmann’s constant. For the small range of absolute temperature in most Chapter 3. Background on reliability 45 applications, (A/T is essentially constant, and (3.42) is close to the Arrhenius relationship (3.41)). Inverse power model: this model is used to analyze accelerated life test data of many electronic and mechanical components such as insulating fluids, capacitors, bearings, and spindles in order to estimate their service lives when the acceleration operating parameters are non-thermal e.g. speed, load, corrosive medium, and vibration amplitude [69, 148]. It is based on the inverse power law relationship between nominal life τof a product and the accelerating variable V, and it is expressed as: τ(V) = A Vγ1,(3.43) here Aand γ1are parameters characteristics of the product, specimen geometry and fabrication, the test method, etc. • Statistics-based models: these models have been developed to estimate the hazard of assets with covariates in both the reliability and biomedical fields. A review on statistics-based models can be found in [53]. In these methods, techniques as standard regressions can be applied to most aging degradation data, as such data are usually complete. However, these models are usually nonlinear in the parameters. In such cases, nonlinear regression methods must be used, which introduces considerable complexity [7]. 3.3 Conclusions In this chapter, the link between reliability and degradation has been explained, followed by a literature review on reliability modeling and degradation models for reliability estimation. The degradation modeling consists in fitting degradation data to a model or a probability distribution. The degradation data can be obtained at the normal operating condition or at accelerated ones. The reliability modeling approaches are related to probability distributions. In this sense, the choice of the probability distribution depends on the application, asset and, system nature. Chapter 3. Background on reliability 46 The hazard model with covariates is one of the most common statistical models in reliability and survival analysis. Some of them have been presented and grouped into nonparametric and semi-parametric models. In such a way, an explanation about covariates and how they explain the failure rate and its integration into probability distributions has been addressed. Some models to specific physical system have been reviewed, those models can be used to fit degradation data by adjusting parameters. In this thesis, the exponential distribution will be used to model the reliability of the assets and their aging phenomenon as described in (3.9). Moreover, the hazard models with covariates will be used to explain the aging process, particularly, the PrHM proposed by [30] because it can explain the degradation process of a large variety of physical systems in a simplified way. Chapter 4 Reliability Assessment This chapter addresses the reliability modeling and assessment of dynamic system using structure function, Bayesian and Dynamic Bayesian networks. This chapter also presents the complexity of reliability modeling and how the inference algorithms are able to handle it. Reliability Importance Measures as indicators of components importance into the system reliability are also presented and explained. All these concepts are then illustrated through an example consisting of a Drinking Water Network system. 4.1 System reliability Generally, systems are composed by subsystems or components. If the state of the subsystems or components can be known, the state of the system can also be known. System structure can be explained through two fundamental relations: series and parallel [60]. Only binary components will be considered, i.e. components having only two states: operational (up) and failed (down). Let xidenote the state of component i, where i= 1, . . . , n and, xi=(1 if component iis up 0 if component iis down .(4.1) Also, it will be assumed that the system can only have two states: up or down. The dependence of a system state on the states of its components will be determined by means the so-called structure function. The structure function allows to determine the dependency of the state of the system regarding the state of its components. This function, denoted as Φ(x), indicates the status of the system (success or failure) given the state of each component. Now, let x= (x1, x2, ..., xn)denotes the state of the ncomponents. It can have one of 2n values which correspond to the possible combinations of the states (working or failed) 47 Chapter 4. Reliability Assessment 48 for the ncomponents. Therefore, the state of the system is characterized by Φ(x)which is a binary value function where Φ(x) = (1 if the system is in working state 0 if the system is in failed state .(4.2) 4.1.1 Reliability of series and parallel systems A serial system is called to be up if and only if all its components are up (see Figure 4.1(a)). Formally, the reliability of a serial system is denoted as: Φ(x) = x1·x2···xn= n Y i=1 xi.(4.3) A parallel system is called to be up if and only if one of its components is up (see Figure 4.1(b)). Formally, the reliability of a parallel system is denoted as: Φ(x)=1−(1 −x1)(1 −x2)···(1 −xn)=1− n Y i=1 (1 −xi).(4.4) (a) (b) FIGURE 4.1: Representation of series (a) and parallel (b) systems. For instance, the structure function can be computed using either the minimal path sets (P1, P2, . . . , Ps) which are the minimal sets of elements of the system whose functioning (i.e., being up) ensures that the system is up, Φp(x)=1− s Y j=1  1−Y i∈Pj xi .(4.5) Chapter 4. Reliability Assessment 49 Or the minimal cut sets (C1, C2, . . . , Ck) which are the minimal sets of elements of the system whose failure (i.e. being down) causes the failure of the system, Φc(x) = k Y j=1  1−Y i∈Cj (1 −xi) .(4.6) Now, let us assume that the state of the ith component is described by a binary random variable Xi, defined by Pr(Xi= 1) = pi,Pr(Xi= 0) = qi= 1 −pi.(4.7) where 1 corresponds to the operational (up) state and 0 corresponds to the failure (down) state. It will be assumed that all components are mutually independent. This implies a considerable simplification, due that for independent components, the joint distribution of X1, X2, . . . , Xn,is completely determined by components reliabilities p1, p2, . . . , pn. Let X= (X1, X2, . . . , Xn)be the system state vector, which is a random vector. Consequently, the system structure function Φ(X) = Φ(X1, X2, . . . , Xn)becomes a binary random variable, i.e. Φ(X) = 1 corresponds to the system up state and Φ(X) = 0 corresponds to the system down state. Hence, the system reliability (r0) is the probability that the system structure function equals 1: r0=Pr(Φ(X) = 1).(4.8) Since, Φ(·)is a binary random variable, it can be written as: r0=E[Φ(X)] .(4.9) Therefore, in a serial system, its structure function is given by Φ(X) = Qn i=1 Xi. Thus: r0=E[Φ(X)] = n Y i=1 pi.(4.10) And a parallel system, its structure function is given by Φ(X)=1−Qn i=1(1 −Xi). Thus: r0=E[Φ(X)] = 1 − n Y i=1 (1 −pi).(4.11) Chapter 4. Reliability Assessment 50 A series system does not have any redundancy because there is only one way for the system to work properly; that is, all components have to work properly. In a parallel system, there are 2n−1different ways for the system to work properly in which each component constitutes a different way [78]. 4.1.2 Reliability of series-parallel systems There are more complex structures that can be explained as combinations of series and parallel structures. Nevertheless, in the cases where such structure reduction cannot be performed, e.g. the case of bridge structure (Figure 4.2), the pivotal decomposition approach is used to compute the system reliability. FIGURE 4.2: Bridge structure. Let (αi;p)denote the vector pwith its ith component replaced by αi. Hence, (1i;p) = (p1, ..., pi−l,1, pi+l, . . . , pn). The pivotal decomposition method consists in computing the reliability of the system pivoting around a component by taking it as fully reliable (up) or completely unreliable (down). In this way, the problem is reduced to compute the system reliability of serial and parallel systems. Example 4.1.1. Consider the bridge structure system shown in Figure 4.2. The best choice is to pivot around element 3. Suppose that component 3 is up. Then the bridge becomes a series connection of two parallel subsystems consisting of elements 1, 2 and 4, 5, respectively. Its reliability is r(13;p) = [1 −(1 −p1)(1 −p2)][1 −(1 −p4)(1 −p5)].(4.12) Now, consider that component 3 is down, then the bridge becomes a parallel connection of two series systems: one with components 1, 4 and the second with components 2, 5. Its reliability is r(03;p) = 1 −(1 −p1p4)(1 −p2p5). Therefore, the system reliability is ro=p3r(13;p) + (1 −p3)r(03;p). Chapter 4. Reliability Assessment 51 The final result is: ro=E[Φ(X)] = p1p3p5+p2p3p4+p2p5+p1p4−p1p2p3p5(4.13) −p1p2p4p5−p1p3p4p5−p1p2p3p4−p2p3p4p5+ 2p1p2p3p4p5.(4.14) The same results can be obtained by using minimal cuts and path sets. For example, the bridge has four minimal path sets: {1,3,5},{2,3,4},{1,4}, and {2,5}. Thus, the random structure function is: Φ(X)=1−(1 −X1X3X5)(1 −X2X3X4)(1 −X2X5)(1 −X1X4).(4.15) Finally, the terms in parentheses are expanded, the expression is simplified using the fact that Xk i=Xiand replacing the reliability of each component. 4.2 Reliability assessment using BNs 4.2.1 Bayesian Networks Basically, a Bayesian Networks (BN) computes the probability distribution in a set of variables according to the prior knowledge of some variables and the observation of others [65]. The BNs are also called as Direct Acyclic Graphs (DAGs). Let Aand Bbe two nodes with two possible states (S1and S2, see Figure 4.3). A probability is associated to each state of the node. This probability is defined a priori for root nodes and computed by inference for the others. The a priori probabilities of node Aare Pr(A=SA1)and Pr(A=SA2). A Conditional Probability Table (CPT) is associated to node Band defines the conditional probability of the state of Bgiven the state of A(Pr(B|A)). Thus, the BN inference computes the marginal distribution Pr(B=SB1): Pr(B=SB1) =Pr(B=SB1|A=SA1)Pr(A=SA1) +Pr(B=SB1|A=SA2)Pr(A=SA2).(4.16) In the Bayesian network approach the probabilistic interactions of the components of a system are represented using a DAG which nodes represent the variables and the arcs between nodes represent the causal relationships between variables [35]. Basically, BNs Chapter 4. Reliability Assessment 52 compute the probability distribution in a set of variables according to the prior knowledge of some variables and the observation of others [66]. FIGURE 4.3: Basic Bayesian Network. A Conditional Probability Table (CPT) is associated with node Band defines the conditional probability of the state of Bgiven the state of A(Pr(B|A)) (Table 4.1). The conditional probability is a measure of the probability of an event given that another event has occurred. TABLE 4.1: CPT of the BN shown in Figure 4.3 AB SB1SB2 SA1Pr(B=SB1|A=SA1) Pr(B=SB2|A=SB1) SA2Pr(B=SB1|A=SA2) Pr(B=SB2|A=SB2) 4.2.2 Inference mechanism Bayesian networks are easy to use thanks to their graphical interpretation. But, the probabilistic inference mechanism constitutes their real strength. The inference of a BN is able to compute the marginal probability distribution of any variable according to: • Observation or measurements of variables (evidence). • The likelihood regarding the state of certain variables. • The conditional probability distribution between variables. In this thesis, the BN and DBN models have been programed using the Bayes Net toolbox [104], which supports many different inference algorithms, such as: • Exact inference for static BNs: junction tree variable elimination brute force enumeration (for discrete nets) Chapter 4. Reliability Assessment 59 C100CFE d100CFE d10COR Legend: Components Sources Reservoir Demand Sector Pumping station FIGURE 4.9: Drinking water network diagram. This network consists of 5 sources and 1 sink. It is assumed that the demand forecast at the sink (dm(k)) is known (Figure 4.10), and that any single source can satisfy this required water demand. It is also assumed that the volume of the tanks should be between a minimum and maximum safety levels. 0 50 100 150 200 250 300 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Time [h] Flow [m3/s] FIGURE 4.10: Drinking water demand. Regarding the reliability of a DWN, in the literature, it is classified into two main categories. The first one, named hydraulic reliability, is related to the probability that a DWN can supply the consumer demands over a specified time interval under specified environmental conditions, i.e., the transport of desired quantities and qualities of water at required Chapter 4. Reliability Assessment 60 pressures to desired appropriate locations at desired appropriate times. The second one, named topological reliability, refers to the probability that a given network is physically connected given the mechanical reliabilities of its components [119]. This thesis and the methodology proposed here are focused only on the topological reliability. Moreover, the reliability modeling illustrated here concerns only to the active components which can be directly commanded. The DWN reliability is modeled using a DBN as follows: first, system components must be identified. In this case there are 10 pumps, 5 sources, 4 tanks and several pipes. Secondly, the minimal path sets should be determined. A minimal path set is composed by those components which allow a flow path between sources and sinks, such as pipes, tanks and pumps. A list of the components that correspond to each minimal path set is presented in Table 4.4. There are nine minimal path sets in the system of Figure 4.9. Each minimal path set is available depending on the reliability of its components. TABLE 4.4: Components and minimal path sets relationship. A1A2A3A4A5e1e2e3e4e5e6e7e8e9e10 P1× × × P2× × × P3× × × P4× × × × P5× × × × P6× × × × × P7× × × × P8× × × × P9× × × × × Note that pipes and tanks are considered perfectly reliable so they do not provide significant information to the network. Nevertheless, sources are included in the minimal path sets merely for illustrating the procedure. Provided the information of Table 4.4, the DBN presented in Figure 4.11 is built as follows: nodes eiand Aiare drawn for each component. Note that nodes eihave two time slices in time kand k+ 1 following the approach of Section 4.2.4. Then, these nodes are interconnected to their minimal path set nodes Piusing arcs. Finally, each minimal path set node is interconnected to the system reliability node S[178]. Initially, at instant k= 0, the pumps and the system are assumed to be fully reliable, i.e. their reliability is 1. Then, the probability of each node is computed using their CPT. Chapter 4. Reliability Assessment 61 FIGURE 4.11: Dynamic Bayesian network model of the DWN. At each sampling time, the reliability Riof each pump is computed according to its failure rate using a MC (Figure 4.11). Its behavior follows an exponential distribution as stated in (3.9). Note that it is independent of the previous states of the component. It only depends on its present state. In the DBN, this corresponds to the CPT shown in Table 4.5. TABLE 4.5: Inter-time slices CPT for node ei(k+ 1). ei(k)ei(k+ 1) Up Dn Up 1-λ0 i·Tsλ0 i·Ts Dn 0 1 The CPT of node P1is shown in Table 4.6. This CPT depends on the states of the source 1 (A1) and pumps 1 and 5 (e1, e5). Its behavior corresponds to an AND gate. It is assumed that with one source it is possible to satisfy the water demand. Thus, the availability of the system can be assured as long as at least one of paths Piis available, which corresponds to the CPT of node Sshown in Table 4.7. It depends on the state of nodes P1to P9and has the behavior of an OR gate. Chapter 4. Reliability Assessment 62 TABLE 4.6: CPTs for nodes P1. A1e1(k+ 1) e6(k+ 1) P1 Up Dn Up Up Up 1 0 Up Up Dn 0 1 Up Dn Up 0 1 Up Dn Dn 0 1 Dn Up Up 0 1 Dn Up Dn 0 1 Dn Dn Up 0 1 Dn Dn Dn 0 1 TABLE 4.7: CPT for node S. P1P2P3. . . P9 S Up Dn Up Up Up . . . Up 1 0 Up Up Up . . . Dn 1 0 ---. . . Up 1 0 . . .. . .. . ..... . .. . .. . . Dn Dn Dn . . . Dn 0 1 4.4 Reliability Importance Measures There are indices that can be used to take actions according to the analysis made of the system reliability and availability. These measures are frequently of significant value in performing trade-off analysis in system design or suggesting the most efficient way to operate and maintain a system or prioritizing improvement efforts. Reliability importance measures were first introduced by [12], and they are classified into two groups: Reliability Importance Measures (RIMs) and Structural Importance Measures (SIMs). The RIMs evaluate the relative importance of a component taking into account its contribution to the overall system reliability while the SIMs provide the relative importance of a component taking into account its position into the system structure. These metrics can be defined either according to their functional aspect, taking into account the minimal path sets, or according to their dysfunctional aspect, considering the minimal cut sets. As both are equivalent, in this thesis only the functional aspect is used. The aim from the system reliability analysis point of view, is to use the RIMs to identify Chapter 4. Reliability Assessment 63 the weakness or strengths in the system and to quantify the impact of component failures over system functioning. 4.4.1 Birnbaum’s Importance Measure The Birnbaum importance measure [12] also known as Marginal Importance Factor (MIF) is related to the probability of a component to be critical for the system functioning. It is defined as: Definition 4.1. The B-reliability importance of component ifor the functioning of the system, denoted as IB(i;p), for a coherent system with independent components is defined as: IB(i;p) =Pr(Φ(X)=1|Xi= 1) −Pr(Φ(X)=1|Xi= 0) =∂R(p) ∂pi =R(1i;p)−R(0i;p). (4.24) The notation R(1i;p)denotes the reliability of the system in which the ith components is replaced by an absolutely reliable one, while R(0i;p)denotes the reliability of the system in which the ith component is failed. The Birnbaum’s measure is the probability that the failure or functioning of the ith component coincides with system failure or functioning. This approach is well known from classical sensitivity analysis. Moreover, it can be interpreted as the maximum lost in system reliability when the ith component changes from the condition of perfect functioning to a failed condition. Note that Birnbaum’s importance measure (IB(i;p)) of the ith component depends only on the structure of the system and the reliabilities of the other components. IB(i;p)is independent of the actual reliability of the ith component (pi). This may be regarded as a weakness of Birnbaum’s measure. 4.4.2 Critical Reliability Importance Measure The Critical Reliability Importance, also known as Critical Importance Factor (CIF), was introduced by Lambert [81] and it is defined as: Definition 4.2. The critical reliability importance of component ifor system functioning, denoted by ICIF (i;p), is defined as the probability that ith component works and is Chapter 4. Reliability Assessment 64 critical for the system functioning given that the system is functioning. ICIF (i;p) =piPr(Φ(X)=1|Xi= 1) −Pr(Φ(X)=1|Xi= 0) Pr(Φ(X) = 1) =pi R(p)IB(i;p). (4.25) Moreover, this can be interpreted as the probability that the ith component has caused a system failure when it is known that the system is failed. 4.4.3 Fussell-Vesely Reliability Importance Measure The Fussell-Veselly (FV) importance measure also known as the Diagnostic Importance Factor (DIF) was proposed initially in the context of fault tree [45, 168]. It is classified in c-type and p-type. The c-type FV importance, takes into account the contribution of component to system failure, and its definition is based on minimal cuts. The p-type FV importance, takes into account the contribution of a component to system functioning, and it is derived from the minimal path sets. It represents the probability that at least one minimal path containing the ith component works, given that the system is functioning. Definition 4.3. The Fussell-Veselly, p-FV (c-FV) reliability importance measure of component i, denoted by Ip DIF (i;p)(Ic DIF (i;p)), is defined as the probability that a minimal path (cut) containing the ith component exists and causes system function (failure). Ip DIF (i;p) =Pr{∃P∈Pis.t.Xj= 1 ∀j∈P|Φ(X)=1} =piPr{(1i,X) : ∃P∈Pis.t. Xj= 1 ∀j∈P} R(p) =Pr(Xi= 1|φ(X) = 1) (4.26) Ic DIF (i;p) =Pr{∃C∈Cis.t.Xj= 0 ∀j∈C|φ(X)=0} =qiPr{(0i,X) : ∃C∈Xj= 0 ∀j∈C} 1−R(p) =Pr(0i|φ(X) = 0) = Pr(Xi= 0|φ(X) = 0) (4.27) where P∈Pi(C∈Ci) denotes the minimal path (cut) containing the ith component. As both are equivalent, the notation IDIF will be used. IDIF can be interpreted as the probability that the functioning of component icontributes to the functioning of the system given that the system is not failed. Chapter 4. Reliability Assessment 65 4.4.4 Reliability Achievement Worth The Reliability Achievement Worth (RAW) describes the increase of the system reliability if the ith component is replaced by a perfect reliable one. It is defined as: Definition 4.4. The RAW, denoted by IRAW (i;p)qualifies the maximum possible percentage of system reliability increase generated by the ith component. It is expressed as: IRAW (i;p) =Pr(Φ(1i,X) = 1) Pr(Φ(X) = 1) =Pr(Φ(X)=1|Xi= 1) Pr(Φ(X) = 1) =1 + qi R(p)IB(i;p).(4.28) 4.4.5 Reliability Reduction Worth, RRW The Reliability Reduction Worth (RRW) measure [86] reflects the reduction of system reliability if the ith component is failed. It is defined as: Definition 4.5. The RRW, denoted as IRRW (i;p)expresses the potential damage produced to the system by the failure of the ith component. IRRW (i;p) = Pr(Φ(X) = 1) Pr(Φ(0i,X) = 1) =Pr(Φ(X) = 1) Pr(Φ(X)=1|Xi= 0) =1 1−pi R(p)IB(i;p).(4.29) 4.5 Example For the computation of the RIMs consider the DWN system described in Example 4.3. it is supposed that sources, tanks ans pipelines are perfectly reliable and only actuators are affected by a loss of reliability according to (3.9). The aim of computing the RIMs in this example is to know, from different points of view, the importance of the pumps reliability to the overall system reliability. In this Section it is done in a static approach, since the failure rate is constant. Later, in Section 6.3.5, the analysis is done dynamically taking into account the evolution of the failure rate. Chapter 4. Reliability Assessment 66 First of all, a static RIM analysis is performed in order to get better knowledge on them. Component reliability is assumed to follow (3.9) with λi(Table 4.8) and the mission time is t=TM(2000 hours). The corresponding results are presented in Table 4.9 and Table 4.10. TABLE 4.8: Pump failure rates λ[h−1·10−4] p1p2p3p4p5p6p7p8p9p10 9.85 10.70 10.50 1.40 0.85 0.80 11.70 0.60 0.74 0.78 TABLE 4.9: Pumps Reliability Importance Measures at TM=2000 Pump λ[·10−4]Ri[%] IB[%] IDIF [%] ICIF [%] IRAW [%] IRRW [%] p19.85 13.94 4.54 14.61 0.77 0.10 0.10 p210.70 11.76 4.43 12.32 0.63 0.10 0.10 p310.50 12.24 4.45 12.83 0.66 0.10 0.10 p41.40 75.57 8.78 77.54 8.05 0.10 0.11 p50.85 84.37 13.72 86.56 14.04 0.10 0.11 p60.80 85.21 87.48 98.58 90.39 0.11 1.04 p711.70 9.63 12.89 10.99 1.50 0.11 0.10 p80.60 88.69 5.81 89.40 6.25 0.10 0.11 p90.74 86.24 12.66 88.06 13.24 0.10 0.11 p10 0.78 85.55 8.11 86.77 8.42 0.10 0.11 In Table 4.10 pumps are sorted according to their reliability importance measures. Remark that pump 6 is the most critical according to all RIMs. It is also interesting to highlight that some RIMs give a similar pump criticality order: CIF and RRW are equivalent, and MIF provides a close result. This a priori knowledge can be used to decide how to distribute the control efforts through the control algorithm. A dynamical RIM analysis will be presented in Section 6.3.5. 4.6 Conclusions In this chapter, the system reliability computation from the reliability of its components or subsystems has been presented. Basically, this can be done by using the system configuration structure, i.e. serial and parallel systems, and its corresponding expression or in the case of complex system configurations by using the pivotal decomposition method. Chapter 4. Reliability Assessment 67 TABLE 4.10: A priori classification of the pumps λiRiIBIDIF ICIF IRAW IRRW p7p8p6p6p6p6p6 p2p9p5p8p5p7p5 p3p10 p7p9p9p1p9 p1p6p9p10 p10 p2p10 p4p5p4p5p4p3p4 p5p4p10 p4p8p4p8 p6p1p8p1p7p5p7 p10 p3p1p3p1p9p1 p9p2p3p2p3p10 p3 p8p7p2p7p2p8p2 It has also addressed the concepts of MC, BN, and DBN. These concepts have been applied to model the reliability of a DWN system. In this chapter, a review of the available RIMs has been performed. These reliability index measures have been evaluated for a DWN system as a tool to identify the importance of each pump to the system functioning. The RIMs have shown to be an objective criterion that can be used to redistribute the control effort among the available system actuators to avoid, for instance, their degradation. Chapter 5 Control System This chapter addresses the concepts involved in control systems and presents the control approaches used in this thesis. Those control methodologies include the Model Predictive Control (MPC) and the Linear-Quadratic Regulator (LQR), which can be used to implement a Health-Aware Control scheme as it will be demonstrated later. 5.1 Model Predictive Control Model Predictive Control (MPC) was developed in the 70’s and has evolved considerably providing a wide range of control methods that have in common the use of an explicit model of the process to calculate the future control input by the minimization of an objective function. These controller design methods lead to schemes which basically have the following ideas [18]: • The use of an explicit model to predict the process output at future time instants (prediction horizon). • The computation of a control sequence that involves the minimization of an objective function. • A receding strategy which slides the prediction horizon towards the future at each time instant, and the application at each step of the first element of the computed control sequence. This control method has the following advantages: • It involves very intuitive concepts which make it relatively easy to implement and tune. • It can be used to control a wide variety of processes, from those with simple dynamics to those with long delay times or nonminimum phase or unstable. 68 Chapter 5. Control System 75 5.2 Linear-Quadratic Regulator The Linear-Quadratic Regulator (LQR) is a well known control design technique that provides feedback gains in a practical manner [114]. Consider the discrete-time, linear time-invariant (LTI) system given in (5.1), and given a cost function defined in a quadratic form as: JLQR =1 2 ∞ X k=0 xT(k)QLQRx(k) + uT(k)RLQRu(k)(5.34) where QLRQ ∈Rnx×nxand RLRQ ∈Rnu×nuare Hermitian positive-definite matrices. If the system is controllable and observable, a feedback control law can be defined as: u(k) = −Kx(k)(5.35) where the optimal feedback gain (K) is the solution of the cost function (5.34): K=RLQR +BTPB−1BTP A. (5.36) The positive-definite symmetric matrix Pis the solution of the discrete-time algebraic Riccati equation: P=QLQR +ATPA −ATPB(RLQR +BTPB)−1BTPA. (5.37) The aim of the cost function in this thesis is to use matrix RLQR as a weighting matrix for the control effort and to redistribute the control input over the available actuators based on their reliability information, while matrix QLQR is used to weight the system states for trajectory tracking. 5.3 Conclusions This chapter presented the control methodologies used in the development of the thesis. In the case of MPC methodology, the optimization problem has been formulated using both, a common quadratic and linear cost functions which include terms for minimize the control action and its variations, and depending the control problem it can include a term for the minimization of the tracking error. At the first glance, both, MPC and LQR are optimal control techniques, but they differ Chapter 5. Control System 76 from the fact that MPC solves the optimization problem in a finite time horizon and uses the receding horizon approach, while LQR solves an optimization problem over an infinity time horizon. An important feature of MPC is the possibility of explicitly including input and state constraints in the control law. In the MPC technique, its weights play an important role in solving the problem. The tune of these parameters can lead to aggressive control responses when the ratio between tracking error and control effort weights is larger. Also, larger horizon prediction produces a more “optimal” controller but increases it complexity. Nevertheless, the control effort weights of both, MPC and LQR, can be selected in order to assign more or less relative importance to each actuator producing lower or higher relative actuator use. This fact can be used to assign weights according to actuators reliabilities or RIMs and in this manner achieve better levels of system reliability. An open issue that requires further research is the procedure to obtain a cost function that can include other terms for their minimization but maintain its simplicity. However, it can lead to a nonlinear or non-convex optimization process whose solution is computationally heavy and could lead to non-implementable control schemes. 77 Part II Contributions Chapter 6 Health-Aware Control This chapter presents the integration of reliability (as a measure of the health of the system or its components) in the control objectives in order to avoid the occurrence of catastrophic and incipient faults. Reliability analysis and failures concern the actuators of the system. This integration is formulated using the reliability importance measures in the parameters of the optimization function. The sensitivity of the system reliability to the degradation of its actuators due to their working load produced by the control action is key to redistribute the control efforts. MPC and LQR techniques are investigated to implement such HAC approach. MPC will be applied to a DWN, and to a multirotot UAV. Moreover, an approach to reduce the degradation of system components is applied to a Twin Rotor system using MPC. 6.1 Reliability assessment for HAC Definition 3.1 and (3.9), the reliability of the ith component of the system will be modeled using the exponential function as: Ri(t) = e−Rt 0λi(v)dv ∀i= 1, . . . , p, where λi(t)is the failure rate depending on time. In this thesis, the overall system reliability will be computed from the reliability of its components using the Bayesian networks (see Section 4.2). Moreover, it is assumed that the overall system reliability is computed by the reliability of its actuators since they are key in achieving the system controllability. Nevertheless, this methodology could be extended to other components of the system. For instance, in a DWN, the reliability of the pipelines could be modeled as a function of the water flow and pressure [5, 80] 78 Chapter 6. Health-Aware Control 79 and include it into the overall system reliability computation, or the reliability of the reservoirs could be modeled as a function of their operational time and the volume of water they store. In any case, to achieve the goal of boosting system reliability, reliability computation should depend on controlled variables, such as voltage, current, flow, pressure, etc. 6.1.1 Failure rate The failure rate of the actuator varies with time and the actuator usage. The failure rate in the useful life period of the actuator is assumed to depend on the impact of the load (usage) and its age. In this thesis, and adaptation of the PrHM, proposed by Cox [30], and defined as (3.25) is used: λi(t) = λ0 i·g(u)∀i= 1, . . . , p, (6.1) where λ0 iis the baseline failure rate (nominal failure rate) for the ith actuator, and g(u) represents the effect of the covariates depending on the applied load u. In here, is it assumed that the parameter λ0 iis equivalent to a constant h0, and g(u)is equivalent to ψ(γz), given in (3.25). Different definitions of function g(u)exists in the literature. However, the exponential form is the most commonly used. In [74, 75] the authors propose a load function based on the root-mean-square of the applied control input (ui) up to the end of the mission (TM), and an actuator parameter defined from the upper and lower bounds of control input. In Guenab [57] a load function is proposed as: gi(u) = eα uprom i∀i= 1, ..., p, (6.2) where uprom iis the average level of the applied control. In [55], the covariate function is defined as: gi(u) = eβi||ui,o:k||2 2∀i= 1, ..., p, (6.3) where βi= (tM(ui,max −ui,min))−1is a shape parameter of the actuator failure for an expected life tM. Chapter 6. Health-Aware Control 80 In [74], the authors propose: gi(u) = eβiRtM 0u2 i(t)dt ∀i= 1, ..., p, (6.4) where the load function is defined according to the root-mean-square of the applied control input until the end of the mission (tM), and βiis an actuator parameter defined as: βi= (tM( ¯ui−ui))−1, and ¯uiand uiare the upper and lower saturation bound of ui, respectively. Another expression proposed in [75] is defined as: gi(u) = eα ui nom ∀i= 1, ..., p, (6.5) where αis a fixed factor depending on the actuator property, ui nom is the nominal control law delivered by the ith actuator to achieve the control objective. In this thesis, different definitions of the covariate function have been studied. Initially, in [143, 144, 146], the covariate function used represents the amount of load on the actuator corresponding to the normalized instantaneous actuator usage at each sampling time: gi(u) = ui(k)−ui ui−ui∀i= 1, ..., p, (6.6) where ui(k)is the control effort at time k,uiand uiare the minimum and maximum control efforts allowed for the ith actuator. In this case, the highest actuator load corresponds to ui(k) = ui, which leads to the highest failure rate λi=λ0 i. Expression (6.6) takes into account the impact of the load at each time instant but not the previous time. In order to include the historical use, equivalent to the age of the actuator, the following covariate function was proposed in [139, 141]: gi(u(t)) = 1 + βiZt 0|ui(v)|dv ∀i= 1, . . . , p, (6.7) where gi(ui(t)) is defined as the cumulative applied control effort of the ith actuator from the beginning of the mission up to time instant tfand βiis a constant parameter. Using (6.1) in (6.7) it yields, λi(t) = λ0 i1 + βiZt 0|ui(v)|dv∀i= 1, . . . , p (6.8) this definition implies that actuators are under a constant reliability decay due to the baseline failure rate which is increased when the actuators are used. Chapter 6. Health-Aware Control 81 Expressions (6.6) to (6.3) only take into account the asset load as a factor in the covariate, whereas the definition (6.8) represents a more realistic situation because it takes into account not only the load of the asset but also the aging produced by the pass of the time. 6.2 Health-Aware Control approaches In the proposed HAC approach, the controller is enriched with system health information provided by a monitoring module and, an accommodation process is performed to tackled the actuators degradation by adjusting the controller parameters. Figure 6.1 presents a generic block diagram of the proposed approach, where the cost function of an optimal control strategy is tuned up on the basis of health information provided by the monitoring module. FIGURE 6.1: Generic block diagram of the proposed HAC In the following sections, three approaches to implement the proposed HAC methodology will be presented: 1. The first is based on a MPC control and includes information about the components and system reliabilities. An MPC controller is set up based on system and component reliability. 2. An LQR controller is set up based on system and component reliability. 3. An MPC scheme is set up based on actuator usage information. Chapter 6. Health-Aware Control 82 To evaluate the benefit of these approaches, some performance indicators will next proposed. 6.2.1 Performance evaluation Different indexes are proposed to evaluate the performance, in control and reliability aspects, of the proposed HAC methodologies. Definition 6.1. The Cumulative Control Effort (Ucum) index indicates the amount of energy spent controlling the system, and is given by: Ucum =Ts TM/Ts X k=0 u(k)Tu(k)(6.9) where Tsand TMare the sampling time and the mission time, respectively. Definition 6.2. The Joint Actuator Reliability (JAR) index measures the remaining overall actuator reliability and is defined as: JAR = p Y i=1 Ri(TM).(6.10) And, the Integral Square Error defined as: Definition 6.3. The ISE measures how well the controller follows the tracking reference. ISE = TM/Ts X k=0 q X i=1 (ˆyi(k)−yref,i(k))2.(6.11) Additionally, the system reliability at the end of the mission time (Rs(TM)) will also be used as a performance measure. 6.3 MPC framework for system reliability optimization 6.3.1 MPC cost function As presented in Section 5.1, the MPC algorithm allows including as many objectives as needed in the cost function. The importance of these objectives in the optimization problem is handled by the weighting parameters, such as αi(k),ρi(k), and δi(k)(5.4). Chapter 6. Health-Aware Control 83 In this scenario, the control objectives study is done by selecting and δi(k) = 1. This means that the MPC performs a smooth control but this effect is not part of the study, and that the reference tracking is study by directly assigning the weight ε. In this section, the optimization problem in (5.4) is reformulated as follows: Consequently, the cost function used is: minimize (ˆu(k|k), ..., ˆu(k+Hc−1|k)) ∆ˆu(k|k), ..., ∆ˆu(k+Hc−1|k)) J(k) = ε  Hp−1 X j=0 q X i=1 (ˆyi(k+j|k)−yref,i(k))2  +(1 −ε)  Hc−1 X j=0 p X i=1 ρi(k)ˆui(k+j|k)2+ Hc−1 X j=0 p X i=1 ∆ˆui(k+j|k)2  subject to u≤ˆu(k+j|k)≤uj= 0, .., Hc−1 x≤ˆx(k+i|k)≤xi= 1, .., Hp (6.12) where αi= 1 and δi= 1 for all i. Note that a trade-off between reference tracking and energy consumption (control effort) arise. To manage this trade-off, a new weight is added to the formulation problem. Therefore, parameter εcan be used to find the appropriate equilibrium between both optimization objectives following the methodology presented later in Section 6.3.2. 6.3.2 MPC tuning methodology The MPC tuning consists in finding the appropriate values for the weighting parameters in the cost function. The values of ρiand εin (6.12) should be selected as follows: Step 1: Set ε= 0 and tune ρisuch that the reliability of the system at the end of the mission time is the highest. The comparison and selection of the best approach will be performed based on the JAR criteria (6.10), and the UCum index (6.9). Step 2: Tune εsuch that the highest system reliability is achieved at mission time Rs(TM) while ISE index (6.11) is the lowest. The goal is to study the effect produced by the variation of parameter εin the system reliability or in other words, the impact of the tracking error in the reliability. Chapter 6. Health-Aware Control 84 Thus, the optimal value for εwhich corresponds to the highest system reliability and the lowest ISE, balancing at the same time both objectives will be found. 6.3.3 Control effort redistribution In this thesis, two approaches to improve the system reliability are studied. On one hand, a local approach which focuses on the actuators reliability, and on the other hand, the global approach which focuses on the overall system reliability. To implement such approaches, the weight ρi(k)in the cost function (6.12) is used as a way to redistribute the control efforts among the actuators [144]. The greater the value given to component ρiis, the higher importance the ith actuator will have and the more penalization it will have in the optimization problem. The local approach attempts to preserve the component reliability: ρi(k)=1−Ri(k)∀i= 1,2, . . . , p. (6.13) This criteria aims at finding the optimal control actions and distributing them among the available actuators in such a way that actuators with lower reliability level are relieved. Hence, the use of highly reliable components is prioritized. The local approach assumes an equivalent contribution of component reliability to system reliability. However, this is hardly ever true because their contribution depends on the system structure and the interconnection of the actuators within such structure. The aim of the global approach is to determine the relative importance of the actuators with the objective of improving the overall system reliability. This study is based on the study of the RIMs (see Section 4.4), which provide different measures of actuator importance. For instance, by using the Birnbaum’s measure as follows: ρi(k) = IB,i(k)∀i= 1,2, . . . , p, (6.14) where IB,i(k)denotes the Birnbaum’s measure of the ith actuator at instant time k. In this case, it is expected that components with a greater contribution to the system reliability are used less than the others. The control strategy scheme is presented in Figure 6.2. The MPC computes the control inputs according to: the cost function, a set of bounds, the current system state and the weights ρi. Then, the control input is injected to the system and used to compute the component failure rates. Chapter 6. Health-Aware Control 91 FIGURE 6.8: Pumping inputs [m3/h]: blue line corresponds to ρi= 1 and red line corresponds to ρi=ICIF,i ·IRRW,i. be: ρi(k) = ICIF,i(k)·IRRW,i(k), and ε= 3.162 ×10−11 which result in a higher system reliability and lower ISE. Figure 6.11 presents the reference tracking response of the control algorithm. Although the system follows the given references for the four tanks, there is some ripples especially for volume of tank 1. This could be due to the less weight assigned to this objective and to the fact that the water demand sector disturbs the volume of tank 1 to a greater extent than the others. Chapter 6. Health-Aware Control 92 FIGURE 6.9: Overall system reliability evolution for the 4 reservoirs in semi-log scale FIGURE 6.10: Tracking error for the 4 reservoirs in semi-log scale 6.3.6 Application to a DWN with multiple demands Now, consider the DWN system presented in Figure 6.12 composed by multiple sources and demand sectors to illustrate the methodology to compute the overall system reliability of a more complex example. Chapter 6. Health-Aware Control 93 0 200 400 600 800 1000 1200 1400 1600 1800 2000 0 1 2 3 4 5 x 104 Time [h] Tanks volume [m3] tank 1 volume setpoint tank 1 tank 2 volume setpoint tank 2 tank 3 volume setpoint tank 3 tank 4 volume setpoint tank 4 FIGURE 6.11: Tracking references of tanks As stated in Section 4.3, this thesis and the methodology proposed here are focused only on the topological reliability. From the system structure, it is possible to obtain a description of the system which only takes into account those components which have a reliability degradation process, in this thesis they are the actuators of the system. Although, it is possible to make some reductions on the structure of complex systems by doing series and parallels equivalences, there are some structures that cannot be reduced and, in such cases other methods to obtain the overall system reliability expression should be used, e.g. the pivotal decomposition or the structure function computation from the minimal path/cut sets methods discussed in Section 4.1. Moreover, this system has more than one source where water can be taken and more than one demand sector where water must be supplied. Therefore, the computation of its overall reliability is different to the system with one water demand sector. In the following, a methodology to compute it is presented: Step 1: Assume a network with l demands. Find the hjminimal path sets of each demand i(mpsj k, corresponding to the kth minimal path sets of the jth demand). This involves all pumps which connect a demand with the sources. Chapter 6. Health-Aware Control 94 FIGURE 6.12: DWN with 3 sources and 4 demands sectors. For instance, the system of Figure 6.12 has 20 minimal path sets distributed as follows: 9 for demand 1, 4 for demand 2, 5 for demand 3, and 2 for demand 4: mps1 1={p1, p3, p7}(6.15) mps1 2={p1, p5, p10}(6.16) mps1 3={p2, p4, p7}(6.17) mps1 4={p2, p6, p10}(6.18) mps1 5={p14, p11, p12}(6.19) mps1 6={p1, p3, p12, p13}(6.20) mps1 7={p2, p4, p12, p13}(6.21) mps1 8={p1, p5, p15, p11, p12}(6.22) mps1 9={p2, p6, p15, p11, p12}(6.23) mps2 1={p1, p3, p9}(6.24) mps2 2={p1, p5, p8}(6.25) mps2 3={p2, p6, p8}(6.26) Chapter 6. Health-Aware Control 95 mps2 4={p2, p4, p9}(6.27) mps3 1={p14, p11}(6.28) mps3 2={p1, p3, p13}(6.29) mps3 3={p2, p4, p13}(6.30) mps3 4={p1, p5, p15, p13, p11}(6.31) mps3 5={p2, p6, p15, p13}(6.32) mps4 1={p1, p5, p16}(6.33) mps4 2={p2, p6, p16}(6.34) Step 2: Next, obtain the structure function of the system as: Φ(X) = l Y j=1   1− hj Y k=1   1−Y pi∈mpsj k Xi    (6.35) where Xiis a binary random variable representing the state of the ith component (as defined in (4.7)). Step 3: Now, the overall system reliability should be obtained by expanding the structure function expression (6.35) and then simplify it using the fact that Xα i=Xiand taking the expectation (and replacing P(Xi= 1) by Ri). Rs=E[Φ(X)] (6.36) For the system of this example, the time spent to compute the overall system reliability was 321.5709 seconds and the reliability expression has 610 terms, which could increase exponentially as the number of minimal path increases. Regarding the RIMs needed to apply the methodology proposed in Section 6.2, it is possible to compute them from the structure function as presented in Section 4.4. 6.4 Health-Aware LQR framework 6.4.1 LQR framework for HAC implementation As presented in Section 5.2, the LQR algorithm minimizes the cost function (5.34), that includes two weighting matrices, RLQR and QLQR. The RLQR matrix weights the control effort magnitude and the QLQR one weights the system states. Chapter 6. Health-Aware Control 96 Thus, matrix RLQR can be used to redistribute the control effort between the actuators based on reliability information. The matrix could be assigned to a combination of RIM indexes, following a similar study to that in the previous section. However, just the following assignment will be analyzed: RLQR(k) = diag(IB(k)),(6.37) with IB= [IB,1, IB,2, . . . , IB,p]. This methodology is illustrated through an example over an octorotor system. For comparison purposes, an additional scenario will be considered, where no actuator reliability is taken into account in the LQR cost function. In this second scenario RLQR will be set as RLQR(k) = I(Iis the identity matrix of proper dimensions). 6.4.2 Octorotor UAV model To describe the dynamics of a multirotor, it is necessary to define the two frames in which it will operate: Inertial frame and Body frame, which are related by the rotation matrix. The inertial frame {I}is static and represents the reference of the multirotor while the body frame {B}is defined by the orientation of the multirotor and is situated in its center of mass (see Figure 6.13). FIGURE 6.13: Octorotor PPNNPPNN structure. The dynamics of a multirotor is given by the following equations [99]: ˙xI=vx(6.38) Chapter 6. Health-Aware Control 97 ˙yI=vy(6.39) ˙zI=vz(6.40) ˙vx=1 m[cos(φI) sin(θI) cos(ψI) + sin(φI) sin(ψI)]T(6.41) ˙vy=1 m[cos(φI) sin(θI) sin(ψI)−sin(φI) cos(ψI)]T(6.42) ˙vz=1 m[cos(φI) cos(θI)]T−g(6.43) ˙ φI=p+ sin(φ) tan(θ)q+ cos(φ) tan(θ)r(6.44) ˙ θI= cos(φ)q−sin(φ)r(6.45) ˙ ψI=sin(φ) cos(θ)q+cos(φ) cos(θ)r(6.46) ˙p=1 Jxx [−(Jzz −Jyy)qr −JpqΩp+τx](6.47) ˙q=1 Jyy [(Jzz −Jxx)pr +JppΩp+τy](6.48) ˙r=1 Jzz [−(Jyy −Jxx)pq +τz](6.49) where Jpis the inertia moment of the motor (rotating parts) and the propeller around z axis, Tis the lift force, and J=diag(Jxx, Jyy, Jzz)is the inertia tensor of the octorotor body. Then, for the octorotor with the structure PPNNPPNN as the one presented in Figure 6.13, where P and N define a positive and negative reactive motor torque respectively (represented as arrows), Ωpis: Ωp=−|Ω1|−|Ω2|+|Ω3|+|Ω4|−|Ω5|−|Ω6|+|Ω7|+|Ω8|(6.50) where Ωiis the angular velocity of the ith motor. Four propellers can rotate in a clockwise direction, while the remaining can rotate anticlockwise. The octorotor is moved by changing the rotor speeds. For example, increasing or decreasing together the eight propellers speeds, vertical motion is achieved. Changing only the speeds of the propellers situated oppositely produces either roll or pitch rotation, coupled with the corresponding lateral motion. Finally, yaw rotation results from the difference in the counter-torque between each pair of propellers. Moreover, the octorotor has actuator redundancy and can function with at least four propellers forming a quadrotor structure. Notice that the arrows in Figure 6.13 represent the reactive torque of the motors. Furthermore, the system inputs Ωiproduce a lift force Tand torques τin x,y, and zaxis given Chapter 6. Health-Aware Control 98 by: uv=BstruΩ(6.51) which in its complete form is:        T τx τy τz        =       kbkbkbkbkbkbkbkb 0−kbls(45) −kbl−kbls(45) 0 kbls(45) kbl kbls(45) −kbl−kblc(45) 0 +kblc(45) kbl kblc(45) 0 −kblc(45) +kd+kd−kd−kd+kd+kd−kd−kd                          Ω2 1 Ω2 2 Ω2 3 Ω2 4 Ω2 5 Ω2 6 Ω2 7 Ω2 8                   , (6.52) where kband kdare coefficients of the motor and lis the distance between the center of mass and the center of the rotor. The parameters value which define the octorotor model are presented in Table 6.4. TABLE 6.4: Parameters value Parameter Symbol Value Body inertia Jxx =Jyy 25 ·10−3[kgm2] Body inertia Jzz 42 ·10−3[kgm2] Propeller inertia Jp104 ·10−6[kgm2] Mass m1.86 [kg] Arm length l0.4[m] Thrust factor kb54.2·10−6[Ns2] Drag factor kd1.1·10−6[Nms2] Equations (6.38)-(6.49) define the nonlinear state space model of an octorotor that should be linearized to apply a health aware linear-quadratic controller (LQR) for the UAV system. The state and inputs vectors considered are x= [x y z φ θ ψ vxvyvzp q r]Tand u= [Ω2 1Ω2 2Ω2 3Ω2 4Ω2 5Ω2 6Ω2 7Ω2 8]T, respectively, the Taylor series approximation at the hover position is applied. The hover position corresponds to the situation where the planes xy of both frames ({I} and {B}) are in parallel and the motors are generating a lifting force equal to the weight of the octorotor. Chapter 6. Health-Aware Control 99 6.4.3 Octorotor LQR controller In this example, the control of the UAV consists in a cascade structure (see Figure 6.14), where the outer loop controls the pose of the UAV in axis xand yand the inner loop controls the pose in axis zand the orientation in axis x,yand z. Note that the inner loop frequency is faster than the outer loop one. FIGURE 6.14: Control scheme. Therefore, the linear model will be divided into two subsystems. On one hand, the inner control loop, where eIis the inner state vector denoted as: eI=xrefi−xi = [ezeφeθeψevzepeqer]T,(6.53) and ∆uIis the inner input vector denoted as: ∆uI=h∆Ω2 1∆Ω2 2∆Ω2 3∆Ω2 4∆Ω2 5∆Ω2 6∆Ω2 7∆Ω2 8iT,(6.54) with ∆Ω2 i= Ω2 i−uff = Ω2 i−mg/(8kb), where uff stands for the input providing the equilibrium point. Therefore, the inner loop model is given by: ˙eI(t) = AIeI(t) + BIBstr∆uI(t)(6.55) where AI="04×4I4×4 04×404×4#, and BI="04×4 βI#, with βI=diag(1/m, 1/Jxx,1/Jyy,1/Jzz)is a diagonal matrix, I4×4is the identity matrix and Bstr is the structural matrix (6.52). Chapter 6. Health-Aware Control 100 On the other hand, the outer control loop, where eois the outer state vector denoted as: eo=xrefo−xo =exeyevxevyT,(6.56) and ∆uois the outer input vector denoted as: ∆uo= [∆φ∆θ]T = [φref θref ]T.(6.57) Therefore, the outer loop model is: ˙eo(t) = Aoeo(t) + Bo∆uo(t),(6.58) where Ao="02×2I2×2 02×202×2#, and Bo=       0 0 0 0 0g −g0        , with gis the gravitational acceleration equal to 9.81 [m/s2]. Regarding the HAC LQR-based, Figure 6.15 presents its scheme, where a module to compute the actuators and system reliabilities is attached to the inner control loop and is used to adapt the parameters of the controller, i.e. the RLQR matrix of the inner loop. 6.4.4 Octorotor reliability model In this application, the covariate is expressed as a function of the load and the age of the actuator. Therefore, the covariate function and the failure rate used are (6.7) and (6.8), respectively. Regarding the overall system reliability, it is computed using the system structure function as presented in Section 4.1. It is assumed that the overall system reliability is determined by the reliability of its actuators and the system controllability. Although the octorotor system has 8 actuators (ri∀i∈[1,8]) in terms of controllability, it can flight without any problem with at least 4 of them. In such scenario the system Chapter 6. Health-Aware Control 107 FIGURE 6.21: Difference between system reliabilities RIB Sand RI S. Assuming that the actuator wear is proportional to the exerted control effort u, a model for the cumulative degradation z(k)∈Rpcan be written as: z(k+ 1) = z(k) + Γ|u(k)|,(6.67) where Γ=diag(γ1, γ2, . . . , γp)is a diagonal matrix of degradation coefficients associated with the pelements of z(k). These coefficients are assumed to be constant and calibrated through experimental data and they may be related to measurable variables such as vibration, internal electrical resistance, or temperature, which are assumed to be known. Assume that a maintenance service is scheduled at sample k=kM. The goal consists in guaranteeing that the cumulative degradation remains below a safe threshold zth during the maintenance horizon, z(k)≤zth,∀k∈[0, kM].(6.68) 6.5.2 MPC formulation under actuator usage constraints A cumulative degradation model which assumes that the degradation of the actuators is proportional to the exerted control effort uand its variations ∆uis proposed in [124, 125]. In such an approach, Pereira, Galvao, et al. [124] propose to manage the actuators degradation through the control action, imposing a threshold in the cumulative degradation of the actuators, and adapting the used actuator to this threshold, by introducing the degradation as a constraint in the cost function. In this case, the system performance is not affected because of the redundancy of the actuators in the system. If an actuator reaches its maxim degradation, its usage is reduced and compensated by the redundant actuator. Chapter 6. Health-Aware Control 108 The actuators degradation model will be integrated in the cost function (5.12) which minimizes the tracking error, the control increments, and the magnitude of the control input. Constraint (6.68) should be included in the optimization problem. However, this would require extending the prediction horizon over the maintenance horizon, becoming an intractable problem. Instead the following constraint will be considered: z(k+Hp)≤zmax(k)(6.69) To determine zmax the uniform "rationing" heuristic proposed in [124] will be followed. This heuristic states that the degradation over a prediction horizon of length Hp< kMis allowed to increase by Hp zth −z(k) kM+Hp−k(6.70) Therefore, the maximum degradation at the end of prediction horizon must not exceed: zmax(k) = z(k) + Hp zth −z(k) kM+Hp−k(6.71) From (6.67), a prediction equation for the degradation index can be written as ˆ z(k+Hp|k) = z(k) + Ξ|ˆ u(k)|(6.72) where Ξ=              Γ 0p×p··· 0p×p 0p×pΓ··· 0p×p . . .. . ..... . . 0p×p0p×p··· Γ . . .. . ..... . . 0p×p0p×p··· (Hp−Hc+ 1)Γ              (6.73) Thus, using equations (6.72) and (5.18), the constraint (6.69) can be replaced with z(k) + ΞΘ≤zmax (6.74) Chapter 6. Health-Aware Control 109 Finally, (6.74) constraint should be added to the LP problem as follows: minimize a(k)cTa(k) s.t. "f1 f2#a(k)≤"b1(k) b2(k)#(6.75) where f2=h0p×pHc0p×qHpΞ 0p×pHc0pi(6.76) and b2(k) = hzmax(k)−z(k)i(6.77) The idea to include actuators usage constraints in the optimization problem was already presented in [124]. The contribution in this thesis consists in analyzing the role of the cost function weights in improving the system safety. Next, this methodology will be applied to a Twin Rotor MIMO system. 6.5.3 Twin Rotor MIMO system The Twin-Rotor MIMO System (TRMS) is a laboratory setup (Figure 6.22) developed by Feedback Instruments Limited for control experiments. The system is perceived as a challenging engineering problem due to its high non-linearity, cross-coupling between its two axes, and inaccessibility of some of its states measures. FIGURE 6.22: Components of the Twin Rotor MIMO System The TRMS mechanical unit has two rotors (the main and the tail, both driven by DC Chapter 6. Health-Aware Control 110 motors) placed on a beam together with a counterbalance whose arm with a weight at its end is fixed to the beam at the pivot and it determines a stable equilibrium position. The beam can rotate freely both in the horizontal and vertical planes. This application aims at showing the effect of the degradation on the system performance through the controller. This is part of the first results obtained in the research on this thesis. Accurate models for TRMS are proposed by [27, 46, 130] and [135]. Each of these models lead to a set of nonlinear differential equations where the TRMS is split into simpler subsystems: the DC-Motors, the propellers and the beam. The first two have independent dynamics, that is, the main motor does not affect the behavior of the tail motor, and vice versa. The same is true for the propellers. On the other hand, the dynamics of the beam are strongly nonlinear with the presence of interaction phenomenon among the horizontal and the vertical dynamics. The state of the beam is described by four process variables: horizontal and vertical angles measured by position sensors fitted at the pivot, and their two corresponding angular velocities. The tacho-generators are used to measure the angular velocities of the rotors. The constants of the nonlinear model are presented in Table 6.7. Parameter Symbol Value Aerodynamic force coeff. of the tail rotor for positive ωh kfhp 1.84 ·10−6[N/rpm2] Aerodynamic force coeff. of the tail rotor for negative ωh kfhn 2.2·10−6[N/rpm2] Aerodynamic force coeff. of the main rotor for positive ωv kfvp 1.62 ·10−5[N/rpm2] Aerodynamic force coeff. of the main rotor for negative ωv kfvn 1.08 ·10−5[N/rpm2] Angular velocity of the TRMS around the vertical axis Ωh[rpm] Angular velocity of the TRMS around the horizontal axis Ωv[rpm] Armature inductance of tail / main motor Lah/av 0.86 ×10−3[H] Armature resistance of tail / main motor Rah/av 8[Ω] Cable force coefficient for negative θhkchn kchp ∗0.9 Cable force coefficient for positive θvkchp 8.54 ·10−3 Horizontal friction coefficient of the beam subsystem koh 4.7·10−3 Vertical friction coefficient of the beam subsystem kov 1.31 ·10−3 Distance between the counterweight and the joint lcb 0.276[m] Drag friction coefficient of the tail propeller kth 5·10−8 Drag friction coefficient ktv 5.6·10−7 Chapter 6. Health-Aware Control 111 Equilibrium pitch angle (uv= 0.2753V)θ0 v0[◦] Gyroscopic constant kg0.2 Input constant of the main motor k28.5 Input voltage of the tail motor uh[V] Input voltage of the main motor uv[V] Input constant of the tail motor k16.5 Length of tail part of the beam lt0.282[m] Length of main part of the beam lm0.246[m] Length of counter-weight beam lb0.290[m] Mass of the counter-weight mcb 0.068[kg] Mass of the counter-weight beam mb0.022[kg] Mass of main part of the beam mm0.014[kg] Mass of the tail shield mts 0.119[kg] Mass of the main DC motor mmr 0.236[kg] Mass of the tail DC motor mtr 0.221[kg] Moment of inertia main DC motor Jmr 2.16 ·10−4[kgm2] Moment of inertia in tail motor Jtr 3.14 ·10−5[kgm2] Positive constant kt2.6·10−5 Mass of the main shield mms 0.219[kg] Mass of the tail part of the beam mt0.015[kg] Pitch angle of the beam θv[◦] Physical constant kah/vϕh/v 0.0202[Nm/A] Positive constant km2·10−4 Radius of the tail shield rts 0.1[m] Radius of the main shield rms 0.155[m] Rotational velocity of the tail rotor ωh[rpm] Rotational velocity of the main rotor ωv[rpm] Viscous friction coefficient of the tail propeller Btr 2.3·10−5[kg.m2/s] Viscous friction coefficient of the main propeller Bmr 4.5·10−5[kg.m2/s] Yaw angle of the beam θh[◦] The mathematical model of the TRMS is represented by the following set of nonlinear differential equations [107]: diah/v dt =−Rah/v Lah/v iah/v −kah/vϕh/v Lah/v ωh/v +k1/2 Lah/v uh/v (6.78) dωh/v dt =kah/vϕh/v Jtr/mr iah/v −Btr/mr Jtr/mr ωh/v −f1/4(ωh/v) Jtr/mr (6.79) dΩh dt =ltf2(ωh) cos θv−kohΩh−f3(θh) Dcos2θv+Esin2θv+F +kmωvsin θvΩvDcos2θv−Esin2θv−F−2Ecos2θv Dcos2θv+Esin2θv+F2 Chapter 6. Health-Aware Control 112 +kmcos θv(kavϕviav −Bmrωv−f4(ωv)) Dcos2θv+Esin2θv+FJmr (6.80) dθh dt =Ωh(6.81) dΩv dt =lmf5(ωv) + kgΩhf5(ωv) cos θv−kovΩv Jv +g((A−B) cos θv−Csin θv)−0.5Ωh2Hsin 2θv Jv +kt(kahϕhiah −Btrωh−f1(ωh)) JvJtr (6.82) dθv dt =Ωv(6.83) where uh/v is the input voltage of the tail/main motor, Ωh/v is the angular velocity around the vertical/horizontal axis, θh/v is the azimuth/pitch beam angle (horizontal/vertical plane), ωh/v is the rotational velocity of the tail/main rotor, Jtr/mr is the moment of inertia in DC-motor tail/main propeller subsystem, kah/vϕh/v is the torque constant of the tail/main motor, and Jvis the moment of inertia about the horizontal axis. Functions fi are defined as: f1(ωh) = sign(ωh)kthω2 h(6.84) f2(ωh) = (kfhpω2 hif ωh≥0 −kfhnω2 hif ωh<0(6.85) f3(θh) = (kchpθhif θh≥0 kchnθhif θh<0(6.86) f4(ωv) = sign(ωv)ktvω2 v(6.87) f5(ωv) = (kfvpω2 vif ωv≥0 −kfvnω2 vif ωv<0(6.88) Finally, the constants of the nonlinear model (6.78)-(6.83) are defined as: A=mt 2+mtr +mtsltB=mm 2+mmr +mmslm C=mb 2lb+mcblcb H=Alt+Blm+mb 2l2 b+mcbl2 cb D=mb 3l2 b+mcbl2 cb F=mmsr2 ms +mts 2r2 ts E=mm 3+mmr +mmsl2 m+mt 3+mtr +mtsl2 t where mms and mts are the masses of the main and tail shields, mmand mtare the masses of the main and the tail parts of the beam, mmr and mtr are the masses of the main and the tail DC-motor with main and tail rotor, mband lbare the mass and the length of Chapter 6. Health-Aware Control 113 the counter-weight beam, mcb and lcb represent the mass of the counter-weight and the distance between the counter-weight and the joint, and rms and rts are the radius of the main and tail shield. 6.5.4 MPC of TRMS system Using the nonlinear model of the TRMS (6.78)-(6.83) a linear model has been obtained by linearizing around an equilibrium point (¯uh= 0.9865,¯uv= 0.2753,¯ θh= 1.5700,¯ θv= 0) and discretizing with a sample time Ts= 0.015 s: A=                   1.0 0.015 4.1e−7 2.8e−8 2.3e−5−6.5e−6−1.9e−7 9.1e−9 −2.8e−3 1.0 5.4e−5 3.7e−6 3.1e−3−8.5e−4−2.5e−5 4.9e−7 0 0 0.96 0.066 0 0 0 0 0 0 −2.4e−3−1.7e−4 0 0 0 0 −1.2e−8 1.3e−5−7.6e−8 5.1e−9 1.0 0.015 8.2e−7 8.1e−9 −2.5e−6 1.8e−3−1.0e−5 5.4e−9−0.075 1.0 1.1e−4 1.1e−6 0 0 0 0 0 0 0.99 0.01 0 0 0 0 0 0 −2.5e−3−2.5e−5                   , (6.89) B=                   1.0e−6 7.2e−7 2.1e−4 9.0e−5 7.6 0 0.79 0 3.9e−7 4.0e−7 3.9e−5 8.0e−5 0 1.5 0 1.1                   , (6.90) and C="10000000 00001000# (6.91) The system input vector is u= [uhuv]Tand the system states are x= [θhΩhωhiah θvΩv ωviav]T. The control objective consists in maintaining the pitch angle θvand the azimuth angle θhin the desired angular positions under decoupling and actuator degradation effects. For the azimuth angle, a square reference signal has been defined, whereas the pitch angle reference signal has been set to 0◦. As not all state variables are measured, a reduced-order state observer is used to estimate nonmeasured variables, which are the angular momentums of the beams and the armature currents of the rotors (Ωh,Ωv,iav, iah). Regarding the HAC for the TRMS, Figure 6.23 presents the control scheme of the approach. Here, an MPC cost function is tuned based on the degradation coefficient of the Chapter 6. Health-Aware Control 114 actuators. Additionally, the data provided by a monitoring module is used to changes the constrains in the optimization problem. This module monitors system control inputs and compute the degradation of the actuators according to their use. FIGURE 6.23: HAC MPC-based for TRMS. Since actuator degradation depends on the time evolution of u, the design of a healthaware MPC law can be taken into account through two approaches: the inclusion of the degradation constraint (6.69) and the proper tuning of the weighting factors ρand α. To study this dependency, 4 case studies are simulated using the nonlinear model of the TRMS. The general MPC parameters are summarized in Table 6.8, and the specific parameters corresponding to each case are shown in Table 6.9. In both tables, subindex 1 (2) corresponds to input uh(uv), except for parameter µwhich is related to θh(θv). HMis the maintenance interval, and is related to the maintenance horizon as follows: HM=kMTs. Simulation results are shown in Figures 6.24-6.27, for azimuth angle θhand pitch angle θv. Cumulative degradation zis shown in Figure 6.28(a) for tail and Figure 6.28(b) for main motor. In Table 6.10 some performance indexes have been evaluated for a simulation interval of 1 hour (Tsim). Rows 2 and 3 indicate the integral of the square of the tracking error (ISE) corresponding to both output angles. The smaller the ISE indexes are, the better the control performance is. Rows 4 and 5 correspond to the cumulative degradation for both actuators (tail/main) at the end of simulation. Chapter 6. Health-Aware Control 115 TABLE 6.8: List of MPC parameters Parameter Symbol Value Prediction horizon Hp80 [s] Control horizon Hc4 [s] Maintenance horizon HM2 [years] Sampling time Ts0.015 [s] Simulation time Tsim 1 [h] Tracking weights µ1/µ21 Control effort weights α1/α2 Control effort variation weights ρ1/ρ2 Upper control efforts u1max/u2max 12 [V] Lower control efforts u1mim/u2mim -12 [V] Degradation threshold zth11 Degradation threshold zth21 Degradation rate γ11.1·10−11 Degradation rate γ22.7·10−10 TABLE 6.9: Case study parameters Case ρ1ρ2α1α2Deg. constraint 1 0 0 0 0 No 2 0 0 0 0 Yes 3 1/120 1/120 1/120 1/120 Yes 4 1/1000 1/120 1/1000 1/120 Yes TABLE 6.10: Performance evaluation Case 1 2 3 4 ISE Θh20.2395 20.2317 21.9480 21.5749 ISE Θv1.4016 1.4246 1.1740 1.1289 Zhcum. (·10−4)0.0305 0.0307 0.0261 0.0266 Zvcum. (·10−4)0.4434 0.4390 0.3109 0.3114 Case 1 corresponds to a neutral MPC control, because it does not take into account the degradation constraint, and only tracking errors are penalized in the cost function. As a result, control performance is good, but the main motor degradation at Tsim is higher in comparison to the other cases. Chapter 6. Health-Aware Control 116 3000 3010 3020 3030 3040 3050 3060 3070 3080 3090 3100 0 20 40 60 80 100 Time [s] yaw Θv [°] Θh Reference 3000 3010 3020 3030 3040 3050 3060 3070 3080 3090 3100 −10 0 10 20 Time [s] pitch angle Θv [°] Θv Reference FIGURE 6.24: Response of azimuth and pitch angles for Case 1. 3000 3010 3020 3030 3040 3050 3060 3070 3080 3090 3100 0 20 40 60 80 100 Time [s] yaw Θv [°] Θh Reference 3000 3010 3020 3030 3040 3050 3060 3070 3080 3090 3100 −10 0 10 20 Time [s] pitch angle Θv [°] Θv Reference FIGURE 6.25: Response of azimuth and pitch angles for Case 2. In Case 2, the degradation constraint is taken into account and the main motor degradation is not allowed to exceed the maximum threshold before the maintenance horizon. As a consequence the pitch control performance is degraded (i.e., ISE increases), but degradation is reduced (i.e., cumulative zdecreases). Chapter 7. Reliability Interpretations 123 FIGURE 7.2: Expected reliability of the asset Rτ i(TM). interval [τ, TM]. Hence, the corresponding reliability of the asset becomes: Rτ i(TM) = e−λτ i×(TM−τ)∀i= 1, . . . , p (7.4) where λi(τ)is the ith asset failure rate computed at time τ. Regarding system reliability, under the first interpretation it will be called the instantaneous system reliability, denoted as RS(t). Under the second interpretation it will be called the expected system reliability, denoted as Rτ S(TM). The Reliability Importance Measure (RIM) should be expressed accordingly with the expected reliability interpretation as it is presented below. 7.1.1 Importance reliability measures As presented in Section 4.4, to measure and quantify the impact of asset failures over the functioning of the system, several indicators concerning reliability importance have been proposed, each of them with a particular purpose [79]. For instance, the Birnbaum’s importance measure IBidefined in (4.24) quantifies the maximum decrease of system reliability due to reliability changes of the ith actuator. According to both reliability interpretations, two Birnbaum’s measures are proposed. On the one hand, under the instantaneous reliability interpretation, the asset Birnbaum’s importance measure will be computed as in (4.24), whereas under the expected reliability interpretation the Birnbaum’s importance measure will be determined as follows: Iτ Bi(TM) =∂Rτ S(TM) ∂Rτ i(TM)=Rτ S(1i, TM)−Rτ S(0i, TM)(7.5) Chapter 7. Reliability Interpretations 124 which denotes the asset Birnbaum’s importance measure at mission time instant TMcomputed at current time τ. 7.1.2 Redistribution policy The redistribution policy discussed in Section 6.3.3 which is based on the actuators RIMs is now used to compare both reliability interpretations (i.e., instantaneous and expected). The MPC technique facilitates the implementation of a Health-Aware Control and the exploration of different redistribution policies without significant changes in the control algorithm. Five scenarios are proposed setting a different weight (ρi) under the two reliability interpretations. In the first scenario actuator reliability is targeted. Accordingly, ρiis set under the instantaneous reliability interpretation as follows: ρi(k)=1−Ri(k).(7.6) In the first scenario, the instantaneous reliability interpretation is assumed. In the second scenario the overall system reliability is targeted using the Birnbaum’s importance measure. Accordingly, ρiis set under the instantaneous reliability interpretation as follows: ρi(k) = IBi(k).(7.7) Similarly, under the expected reliability interpretation, the third scenario corresponds to: ρi(k)=1−Rk i(kf)(7.8) and the fourth scenario to: ρi(k) = Ik Bi(kf)(7.9) where kfis the end of mission sample, corresponding to TM. In the fifth weight assignment, which is common for both reliability interpretations, no reliability feedback is taken into account, i.e., ρi(k)=1. Chapter 7. Reliability Interpretations 125 7.1.3 Performance evaluation In addition to the cumulative control effort index (Ucum) new indices are defined next, which will be used in assessing the health-aware control performance under both reliability interpretations. Definition 7.1. The Cumulative System Reliability indices indicate the aggregated system reliability over the mission time. Under the instantaneous reliability interpretation, it is denoted in discrete time as: RScum =Ts TM/Ts X k=0 Rs(k).(7.10) And under the expected reliability interpretation it is denoted as: Rkf Scum =Ts TM/Ts X k=0 Rk s(kf).(7.11) A higher value indicates a better management of the assets reliabilities with the objective of improving the overall system reliability. 7.2 Drinking Water Network example Both reliability interpretations will be compared on an application to the DWN presented in Section 6.3.4. The objective is to apply the MPC HAC methodology to maintain the pumps and tanks within their bounds and extend the reliability of the system. The cost function used in this example is (6.12) with ε= 0, and the failure rate of the pumps is computed using (6.8) in accordance with the reliability interpretation. The parameters of the simulation are presented in Table 7.1. Figure 7.3 shows the instantaneous system reliability in the scenario where no reliability feedback method is applied (ρi= 1) where the overall system reliability presents a gradual decreasing behavior, which is characteristic of the instantaneous reliability interpretation. This is the nominal scenario and will be compared with the results of the other ρisettings. To illustrate how reliability behaves under the expected reliability interpretation, the reliability computed at each sampling time were projected τ+TMhours into the future. Chapter 7. Reliability Interpretations 126 TABLE 7.1: Simulation parameters Parameter Symbol Value Prediction horizon Hp24 [h] Control horizon Hc8 [h] Sampling time Ts1 [h] Component parameter βi10−2∀i∈[1,10] Upper control bound ui 0.75 0.75 0.75 1.20 0.85 [m3/s] 1.60 1.70 0.85 1.70 1.60 Lower control bound ui0∀i∈[1,10] [m3/s] Baseline failure rate λ0 i 9.85 10.70 10.50 1.40 0.85 [h−1·10−4] 0.80 11.70 0.60 0.74 0.78 Upper state bound xi65200 3100 14450 11745 [m3] Lower state bound xi25000 2200 5200 3500 [m3] Initial states xi(0) 45100 2650 9825 7622 [m3] 0 200 400 600 800 1000 1200 1400 1600 1800 2000 0.4 0.5 0.6 0.7 0.8 0.9 1 FIGURE 7.3: Instantaneous overall system reliability profile evolution with ρi= 1. Some of those values are presented in Figure 7.4, where the overall system reliability computed at τ= 1,500,1000,1500, and 2000 hours is projected into the future. Next, Figure 7.5 presents the expected overall system reliability evolution at time instant TM, which consists of the values of each system reliability projection at TMcomputed at each sampling time. In this case, the reliability starts from a value between 0 and less than 1 and moves towards 1. Chapter 7. Reliability Interpretations 127 0 500 1000 1500 2000 2500 3000 3500 4000 0.4 0.5 0.6 0.7 0.8 0.9 1 FIGURE 7.4: Expected overall system reliability profile evolution at different time instants with ρi= 1. 0 200 400 600 800 1000 1200 1400 1600 1800 2000 0.65 0.7 0.75 0.8 0.85 0.9 0.95 1 FIGURE 7.5: Expected system reliability at mission time TM= 2000hevolution with ρi= 1. 7.2.1 Reliabilities comparison The five scenarios proposed in Section 7.1.2 have been considered in the MPC control of the DWN. All cases will be assessed under both cumulative reliability indices (7.10) and (7.11). Figure 7.6 shows the instantaneous system reliability evolution for the five scenarios. Remark that the most suitable policies that improve system reliability are those based on the Birnbaum’s measure. Chapter 7. Reliability Interpretations 128 0 200 400 600 800 1000 1200 1400 1600 1800 2000 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 FIGURE 7.6: Instantaneous system reliability. Figure 7.7 provides the expected system reliability at the end of the mission time TMfor the five scenarios. Again the best results correspond to those based on the Birnbaum’s measure. 0 200 400 600 800 1000 1200 1400 1600 1800 2000 0.6 0.65 0.7 0.75 0.8 0.85 0.9 0.95 1 FIGURE 7.7: Expected system reliability at mission time tf. The reliability performance indices were also computed for each scenario and are presented in Table 7.2. The performance indices confirm that the best reliability performance is attained when using a redistribution policy based on the Birnbaum’s importance measure, i.e., (7.7) and Chapter 7. Reliability Interpretations 129 TABLE 7.2: Reliability performance indexes. ρi(k)RScum [·106]Rkf Scum [·106]Ucum [·106] 1 5.6131 5.5583 1.5370 1−Ri(k) 5.4046 5.4054 1.9687 1−Rk i(kf) 5.4525 5.4340 1.9002 IBi(k) 6.1006 6.1653 3.2158 Ik Bi(kf) 6.0915 6.1447 3.5040 (7.9), with a small improvement when the instantaneous reliability interpretation is followed, i.e., (7.7). Focusing on actuator reliability (i.e., (7.6) and (7.8)) does not optimize system reliability. However, targeting system reliability leads to a greater actuator energy expenditure. To illustrate the performance of the control algorithm, tank volumes for the best redistribution policy, corresponding to (7.7), are presented in Figure 7.8. Note that, the DWN is able to supply the required water demand maintaining tank volumes within the specified bounds. 0 200 400 600 800 1000 1200 1400 1600 1800 2000 0 0.5 1 1.5 2 2.5 3 3.5 10 4 FIGURE 7.8: Tank volumes corresponding to ρi(k) = IBi(k). And the pump control actions corresponding to the policy given in (7.7) are presented in Figure 7.9. Note that the control actions evolve in time according to the importance of each actuator, which change dynamically as their reliabilities change. Remark that in this scenario, the cost function which computes the control actions does not take into account any tracking objective. Therefore, tank volumes freely evolve within Chapter 7. Reliability Interpretations 130 FIGURE 7.9: Pump commands corresponding to ρi(k) = IBi(k). their bounds. 7.3 Conclusions In this chapter, two reliability interpretations have been presented and illustrated using a DWN system. Both approaches have been applied to a Health-Aware Control scheme based on an MPC algorithm with the objective of improving system reliability. Chapter 7. Reliability Interpretations 131 The study of two different interpretations of reliability were presented. This study was done due to the need of clarify what is meant by reliability in both cases and what it is obtained by using each of those interpretations. The results of the redistribution policy provide similar results in terms of reliability enhancement independently of the reliability interpretation. Thus, both interpretations are virtually equivalent. This Chapter aims to illustrate another interpretation of the reliability. It is not intended to do the same experimentations done in the previous chapter where different weighting criteria were compared, and only one RIM was used in the experiment. Moreover, another RIMs or a mix of them can give better results in terms of reliability improvement, as it was demonstrated in the previous chapter. Moreover, by applying both of them results coincide in determining that the Birnbaum’s measure is the best approach to integrate the assets reliability into the control law. 132 Part III Conclusions and perspectives