scieee AI-readable full text Open interactive document viewer

System identification and fault reconstruction in solar plants via extended Kalman filter-based training of recurrent neural networks

Ruiz-Moreno, Sara; Bemporad, Alberto; Gallego Len, Antonio Javier; Camacho, Eduardo F.

Abstract

This article proposes using the extended Kalman filter (EKF) for recurrent neural network (RNN) training and fault estimation within a parabolic-trough solar plant. The initial step involves employing an RNN to model the system. Given the challenge of fault discernibility in the collectors, parallel EKFs are employed to reconstruct the parameters of the faults. The parameters are used independently to estimate the system output, and the type of fault is isolated based on the estimation errors using another feedforward neural network. To evaluate the effectiveness of the methodology, simulations are conducted on a loop of the ACUREX plant with irradiances from sunny and cloudy days. The results reveal a fault classification accuracy of approximately 90% and a fault reconstruction error below 3%, with even better accuracies in the cloudy dataset than in the sunny dataset.

Full text

Contents lists available at ScienceDirect ISA Transactions journal homepage: www.elsevier.com/locate/isatrans Research article System identification and fault reconstruction in solar plants via extended Kalman filter-based training of recurrent neural networks✩ Sara Ruiz-Morenoa,∗,1, Alberto Bemporada, Antonio Javier Gallegob, Eduardo Fernández Camachob aIMT School for Advanced Studies Lucca, 55100, Lucca, Italy bDept. de Ingeniería de Sistemas y Automática, University of Seville, Camino de los Descubrimientos, no number E-41092, Seville, Spain ARTICLE INFO Keywords: Extended Kalman filter Recurrent neural network Fault diagnosis Solar energy Deep learning ABSTRACT This article proposes using the extended Kalman filter (EKF) for recurrent neural network (RNN) training and fault estimation within a parabolic-trough solar plant. The initial step involves employing an RNN to model the system. Given the challenge of fault discernibility in the collectors, parallel EKFs are employed to reconstruct the parameters of the faults. The parameters are used independently to estimate the system output, and the type of fault is isolated based on the estimation errors using another feedforward neural network. To evaluate the effectiveness of the methodology, simulations are conducted on a loop of the ACUREX plant with irradiances from sunny and cloudy days. The results reveal a fault classification accuracy of approximately 90% and a fault reconstruction error below 3%, with even better accuracies in the cloudy dataset than in the sunny dataset. 1. Introduction In modern industries, many autonomous systems are integrated into processes, resulting in an increasing number of sensors and actuators that may fail unexpectedly. However, this progress also brings with it concerns about safety and reliability [1], a focus that is also crucial for today’s companies [2]. Therefore, there is a need to automatically detect faults and failures in the system [3]. To this end, the field of fault detection and diagnosis (FDD) aims to determine the occurrence of a fault and reveal some relevant information about it. Fault diagnosis is further categorized into fault isolation, which locates and assesses the type of fault, and fault identification, which determines its magnitude. FDD is a rapidly expanding field with diverse applications across various industrial sectors. For instance, Molinié et al. [4] propose an unsupervised clustering method to identify anomalies in industrial systems. Their approach involves recursively partitioning a space using an integrity criterion. In another study, Chen et al. [5] introduce an interpretable mechanism based on convolutional neural networks with score-weighted class activation to classify faults in climate control systems. Abdel Karim et al. [6] apply FDD to communication networks ✩This research was supported by the European Union’s Horizon 2020 research and innovation program under the ERC Advanced Grant agreement No 789051, and by the Spanish Ministry of Science and Innovation (Grant n. FPU20/01958). ∗Corresponding author. E-mail addresses: [email protected] (S. Ruiz-Moreno), [email protected] (A. Bemporad), [email protected] (A.J. Gallego), [email protected] (E.F. Camacho). 1IMT Permanent address: Dept. de Ingeniería de Sistemas y Automática, University of Seville, Camino de los Descubrimientos, no number E-41092, Seville, Spain. using power-line communications and residuals. In the field of photovoltaics, Hajji et al. [7] apply and compare different artificial neural networks (ANNs) to diagnose several typical faults. The work by Kumar and Devakumar [8] focuses on diagnosing sensor faults by estimating state variables with recurrent neural networks (RNNs) and obtaining residuals. In the event of a fault, a new ANN classifies it into different sensor conditions. Rodríguez et al. [9] use digital twins in a solar cooling plant, incorporating a neuro-fuzzy system to detect faults and an RNN to identify them. Additionally, a Kalman filter with the RauchTung-Striebel (RTS) smoother is used by Bidou et al. [10] to estimate failure and restart times in heat sources. As introduced earlier, the FDD field is enriched by the use of artificial intelligence (AI), with artificial neural networks being one of its well-known components. The use of AI is well developed in systems engineering, with applications such as intrusion monitoring [11], digital twins in Fresnel plants [12], reinforcement learning for adaptive control of solar collector plants [13], or Lagrange multipliers initialization for computational reduction in distributed model predictive control (MPC) [14]. The continuous evolution of technology has allowed the https://doi.org/10.1016/j.isatra.2025.01.002 Received 21 March 2024; Received in revised form 2 January 2025; Accepted 2 January 2025 ISA Transactions 158 (2025) 272–284 Available online 8 January 2025 0019-0578/© 2025 The Authors. Published by Elsevier Ltd on behalf of International Society of Automation. This is an open access article under the CC BY license ( http://creativecommons.org/licenses/by/4.0/ ). S. Ruiz-Moreno et al. Nomenclature Parameters and variables 𝛼(t) Fault multiplier (−) 𝛿sDeclination (◦) 𝜔s(t) Hourly angle (◦) 𝜌(𝑇)Density (kg/m3) 𝐴Pipe cross-sectional area (m2) 𝐶(𝑇)Specific heat capacity (J/ (kg ◦C)) 𝐺Collector aperture (m) 𝐻l(𝑇)Thermal loss coefficient (W/ (m2◦C)) 𝐻t(𝑇)Convective heat transfer coefficient (W/ (m2 ◦C)) 𝐾opt Optical efficiency (−) 𝐿Tube perimeter (m) 𝑛o(𝑡)Geometric efficiency (−) 𝑞(𝑡)Flow rate (m3/s) 𝑆Total area of the field (m2) 𝑡Time (s) 𝑇(𝑡, 𝑥)Temperature (◦C) 𝑥Space (m) Subscripts 𝑎Ambient f Fluid in Input m Metal mean Mean between input and output ref Reference development of many different neural network architectures. Notably, recurrent neural networks are models that can capture unknown dynamics well, although their peculiarity of incorporating hidden states makes them much more challenging to train. While stochastic gradient descent (SGD) [15] is the most commonly used algorithm, alternative approaches offer different advantages. In this study, an RNN is trained using the Extended Kalman Filter (EKF) due to its faster performance, as demonstrated by Trebatický and Pospíchal [16] and by Bemporad [17] for nonlinear systems within the framework of MPC. Solar energy has garnered substantial attention due to the growing interest in renewable energy sources [18]. This research delves explicitly into the realm of thermal solar plants, focusing on parabolic trough collectors (PTCs). These collectors reflect solar radiation to heat a fluid to produce thermal energy. The application of FDD techniques to thermal solar plants is a broad research topic to be explored, with most applications in the existing literature focusing on water systems and small-scale plants. Notably, most applications center around fault detection or isolation rather than identifying collector module parameters. For example, the work by Schmelzer et al. [19] detects faults in a combi system by analyzing fractional solar consumption, and Brenner et al. [20] estimate mirror soiling in a PTC plant using feedforward neural networks. A related application of machine learning is presented in [21], where information from the defocusing mechanism of the plant feeds into an ANN classifier that detects and categorizes faults in the collector area. Despite these efforts, there remains a need for focused exploration into fault identification of collector module parameters in thermal solar plants. Most of the research in FDD is focused on systems with many outputs and subsystems. In the case of PTCs, the only output is the temperature, and the faults are highly correlated, which makes them difficult to isolate. This work uses neural networks to distinguish faults when the plant models cannot by including them in a new fault detection and reconstruction methodology. Specifically, RNNs are helpful to model the complex internal dynamics of the system. An EKF was chosen to train the RNN and estimate faults with the RNN trained because of its fastness and ability to reject measurements with significant errors. The methodology proposed in this work consists of two steps: First, an RNN is trained using an EKF to model the dynamics of a PTC system, taking into account several faults introduced into the plant, benefiting from the fact that RNNs are time-aware and take into account the activations from previous data. In this step, the EKF helps capture the system dynamics and adapt the RNN to fault conditions. In this type of plants, the faults are highly coupled and different faults can produce the same changes in the output. For this reason, three EKFs governed by a feedforward neural network are applied in parallel, each dedicated to diagnosing a single fault. These filters utilize the previously trained RNN model to estimate the values of the faults present in the PTC system. The flowcharts in Fig. 1describe the complete process. To the best of the authors’ knowledge, this is the first time a combination of EKF and RNN is applied to PTC plants and fault reconstruction. This work proposes a new FDD methodology applied to solar plants using RNNs and EKFs and is not published elsewhere. The primary contributions of this research are outlined as follows: •Use of an EKF-trained RNN to model the outlet and intermediate temperatures in one collector loop of a PTC plant with batch learning and taking into account fault parameters as inputs to the model. •Fault estimation achieved through parameter reconstruction with EKF in the optical efficiency, flow rate, and thermal losses. •Fault classification with an ANN governing a set of EKFs and selection of inputs based on output estimation error. The remainder of this paper is as follows. Section 2presents an overview of the system. The training of the neural network and the FDD process are described in Sections 3and 4, respectively. Section 5 describes the evaluation metrics used in this paper. Some simulation results, focusing on the estimation and classification errors, are presented in Section 6. Finally, Section 7elaborates on the principal findings of the study and draws conclusions. 2. System description A PTC plant is a type of solar thermal system that comprises loops of parabolic mirrors designed to concentrate solar rays onto a focal line through which a heat transfer fluid (HTF) circulates. This fluid, typically water or oil, is heated to generate thermal energy. The heated HTF is then usually transported to a steam generator to drive a turbine, as illustrated in Fig. 2[22]. The loops of a PTC system are described by the distributed parameter model, which accounts for the energy balances in the tubes and the HTF as follows [23]: 𝜌m𝐶m𝐴m 𝜕 𝑇m 𝜕 𝑡=𝛼𝐾opt 𝐼 𝐾opt𝑛o𝐺−𝐿𝐻t(𝑇m−𝑇f) −𝛼𝐻l𝐻l𝐺(𝑇m−𝑇a)(1) 𝜌f𝐶f𝐴f 𝜕 𝑇f 𝜕 𝑡+𝛼𝑞𝜌f𝐶f𝑞𝜕 𝑇f 𝜕 𝑥=𝐿𝐻t(𝑇m−𝑇f)(2) The model has been adapted to incorporate the parameters 𝛼representing the faults introduced into the system. These faults include optical efficiency faults 𝛼𝐾opt , flow rate faults 𝛼𝑞, and thermal loss faults 𝛼𝐻l[24]. It is presupposed that the metal’s temperature remains constant radially, and the dimensions of the reflector and receiver remain the same along the loop, except for the passive parts (those not exposed to solar radiation), resulting in a uniform local concentration ratio. ISA Transactions 158 (2025) 272–284 273 S. Ruiz-Moreno et al. Fig. 1. Flowchart of the proposed methodology. Fig. 2. General scheme of a PTC plant. 2.1. Faults considered The three types of faults considered in this work, located in the collector area, are described as follows: •Faults in the optical efficiency, denoted as 𝛼𝐾opt <1. Since optical efficiency is influenced by tube absorptance, reflectivity, interception factor, and soiling of the reflectors, these faults encompass coating, breakage, tube deterioration, mirror defects, dirt, degradation, and corrosion. •Faults in the flow rate, expressed as 𝛼𝑞≠1, are related to flowmeter failures and loop imbalances. •Faults in the thermal loss, indicated by 𝛼𝐻l>1, are linked to pressure drop in the tubes resulting from wear, insulation, dirt accumulation, and pipe breakage. The parameters represent deviations from the nominal values and signify various issues that can affect the performance of the PTCs, although these faults only affect the model of the loop and not the rest of the subsystems. To adapt this methodology to a real plant, three paths ISA Transactions 158 (2025) 272–284 274 S. Ruiz-Moreno et al. Fig. 3. ACUREX collector loops. Table 1 Parameters of the ACUREX plant. Parameter Value 𝜌m7800 kg/m3 𝐶m550 J/kg◦C 𝐴m2.4806 ⋅10−4 m2 𝐺1.82 m 𝐿7.98 ⋅10−2 m 𝐴f5.0671 ⋅10−4 m2 𝑆2672 m2 could be taken: first, using a digital twin that accurately reproduces the system and from which to take the data for training the neural networks. Secondly, one can provoke these faults in the actual plant in an instrumented loop to gather the necessary data (by partially covering the mirrors, modifying the insulation of the pipes, or mismatching the flow rate). Finally, this information can be gathered from largescale plants where knowledge of failures is acquired retrospectively as analysis and maintenance tasks are performed periodically. In this case, the advantage of the proposed methodology would be to detect these failures with greater immediacy. 2.2. Case study The methodology was tested in simulation on one loop of ACUREX [25], an experimental, 1-MW PTC plant that was situated at the Plataforma Solar de Almería. Fig. 3depicts some of the collectors within ACUREX. The plant is comprised of 10 loops of single-axis collectors aligned in an east–west orientation, with each loop containing 12 modules grouped into 4 collectors. The loops are composed of an active part of 142 m and a passive part of 30 m. ACUREX serves as a practical and representative case for evaluating the proposed methodology. Table 1contains the parameters of the ACUREX plant, and the thermal loss coefficient 𝐻𝑙and convective thermal exchange within the inner tube 𝐻𝑡are given by, respectively [26]: 𝐻l= 0.00249 (𝑇f−𝑇a)− 0.06133 (3) 𝐻t=𝑞0.8(2.17 ⋅106− 5.01 ⋅104𝑇f+ 4.53 ⋅102𝑇2 f− 1.64𝑇3 f+ 2.1⋅10−3𝑇4 f)(4) The geometric efficiency [27,28] is derived from the correlation between the radiation beam’s direction and the perpendicular vector of the mirror. As the sun is tracked in elevation, the geometric efficiency is calculated using the following equation [29]: 𝑛o=(1 − cos2(𝛿s) sin2(𝜔s))1 2(5) The plant is equipped with a sun tracking system that precisely controls the mirror rotation around an axis aligned parallel to the pipe, enhancing the geometric efficiency for solar radiation capture and utilization [26]. The HTF used in the system is Therminol 55 thermal oil. The specific heat capacity 𝐶fand density 𝜌fof the HTF characterize its thermophysicial properties and are, respectively: 𝐶f= 3.478𝑇f+ 1820 (6) 𝜌f= −0.672𝑇f+ 903 (7) The loop is discretized longitudinally into 172 segments, each with a length of 1 m. The model is computed with an integration step of 0.25 s. Initially, the metal temperature is calculated. Subsequently, the fluid temperature is obtained under the assumption of steady-state conditions. Finally, the fluid temperature is corrected by considering the energy transfer between each fluid control volume. For further details, the reader is referred to [26]. 2.3. Flow-rate controller The flow rate is manipulated using valves to control the outlet temperature and track a reference. The concentrated parameter model of the plant, which describes the variation of the internal energy of the fluid in steady state, is used to implement the following feedforward controller: 𝑞=𝑛o𝐾opt𝑆 𝐼−𝐻l𝐴(𝑇mean −𝑇a) 𝑃cp(𝑇ref −𝑇in)(8) where 𝑃cp =𝜌m𝐶m. The flow rate is constrained within the range [0.2,1.2] l/s and the controller sampling time is 39 s. This control strategy effectively adjusts the flow rate per loop to attain the target outlet temperature, considering the objectives of this work. It is important to note that the primary focus of this study is not on optimizing control actions but rather on implementing an FDD strategy, so no other control technique is applied. In practice, an advanced control technique like MPC would be used. It is not relevant for this work because the ranges in which the control signal varies are similar for both controllers. 3. Recurrent neural network A recurrent neural network is a type of ANN that incorporates delayed feedback loops in its layers [30]. Similar to multilayer perceptrons, RNNs comprise an input layer, an output layer, and one or more hidden layers. Each hidden layer contains a varying number of neurons with an activation function that produces nonlinear effects. RNNs are time-aware and can capture the system dynamics by considering previous data in the sequence as inputs. They can be described by the state-space model: 𝑥(𝑘+ 1) =𝑓𝑥(𝑥(𝑘), 𝑢(𝑘), 𝜃𝑥) 𝑧(𝑘) =𝑓𝑧(𝑥(𝑘), 𝑢(𝑘), 𝜃𝑧)(9) where 𝑢∈R𝑛𝑢represents the input, 𝑧 ∈R𝑛𝑧is the predicted output, 𝑥∈R𝑛𝑥is the state vector and 𝜃𝑥∈R𝑛𝜃𝑥, and 𝜃𝑧∈R𝑛𝜃𝑧are the trainable model parameters. More explicity, as done in [17], Eq. (9) can be expressed by: ⎧ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎩ 𝑣𝑥 1(𝑘) =𝐴𝑥 1[𝑣𝑥 𝐿𝑥(𝑘− 1) 𝑢(𝑘)]+𝑏𝑥 1 𝑣𝑥 2(𝑘) =𝐴𝑥 2𝑓𝑥 1(𝑣𝑥 1(𝑘)) +𝑏𝑥 2 ⋮ 𝑣𝑥 𝐿𝑥(𝑘) =𝐴𝑥 𝐿𝑥𝑓𝑥 𝐿𝑥−1(𝑣𝑥 𝐿𝑥−1(𝑘)) +𝑏𝑥 𝐿𝑥 𝑥(𝑘+ 1) =𝑣𝑥 𝐿𝑥(𝑘) (10) ISA Transactions 158 (2025) 272–284 275 S. Ruiz-Moreno et al. ⎧ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎩ 𝑣𝑧 1(𝑘) =𝐴𝑧 1[𝑣𝑥 𝐿𝑥(𝑘) 𝑢(𝑘)]+𝑏𝑧 1 𝑣𝑧 2(𝑘) =𝐴𝑧 2𝑓𝑧 1(𝑣𝑧 1(𝑘)) +𝑏𝑧 2 ⋮ 𝑣𝑧 𝐿𝑧(𝑘) =𝐴𝑧 𝐿𝑧𝑓𝑧 𝐿𝑧−1(𝑣𝑧 𝐿𝑧−1(𝑘)) +𝑏𝑧 𝐿𝑧 𝑧(𝑘) =𝑣𝑧 𝐿𝑧(𝑘) (11) where 𝜃𝑥=(𝐴𝑥 1, 𝑏𝑥 1,…, 𝐴𝑥 𝐿𝑥, 𝑏𝑥 𝐿𝑥)and 𝜃𝑧=(𝐴𝑧 1, 𝑏𝑧 1,…, 𝐴𝑧 𝐿𝑧, 𝑏𝑧 𝐿𝑧),𝐿𝑥 and 𝐿𝑧are the number of hidden layers in the state-update and output functions, 𝑣𝑥 𝑖∈R𝑛𝑥 𝑖and 𝑣𝑧 𝑖∈R𝑛𝑧 𝑖are values associated with the neurons, 𝑓𝑥 𝑖∶R𝑛𝑥 𝑖→R𝑛𝑥 𝑖+1 and 𝑓𝑧 𝑖∶R𝑛𝑧 𝑖→R𝑛𝑧 𝑖+1 are the activation functions, 𝐴𝑥 𝑖∈R𝑛𝑥 𝑖×𝑛𝑥 𝑖−1 and 𝐴𝑧 𝑖∈R𝑛𝑧 𝑖×𝑛𝑧 𝑖−1 are the weight matrices, and 𝑏𝑥 𝑖∈R𝑛𝑥 𝑖and 𝑏𝑧 𝑖∈R𝑛𝑧 𝑖are the bias terms. 3.1. Training with EKF Typically, the parameters of ANNs are updated using gradient descent (GD) methods [31,32]. However, in this work, the supervised training of the RNN is performed using the EKF algorithm to achieve much faster and less expensive computations, as emphasized by Bemporad [17]. For the specific case of modeling a system with faults 𝛼, the input vector is extended to include them, treating them as inputs to the system. To obtain training and test data, these faults must be intentionally introduced and known to the training algorithm. The model can be rewritten in the form of: 𝑥(𝑘+ 1) =𝑓𝑥(𝑥(𝑘), 𝑢(𝑘), 𝜃𝑥(𝑘)) +𝜂𝑥(𝑘) 𝑧(𝑘) =𝑓𝑧(𝑥(𝑘), 𝑢, 𝜃𝑧) +𝜂𝑧(𝑘) 𝛼(𝑘+ 1) =𝛼(𝑘) +𝜂𝛼(𝑘) 𝜃(𝑘+ 1) =𝜃(𝑘) +𝜂𝜃(𝑘) (12) where 𝑢(𝑘) = [𝑢(𝑘), 𝛼(𝑘)]𝑇and 𝜃(𝑘) = [𝜃𝑥(𝑘), 𝜃𝑧(𝑘)]𝑇and 𝜂𝑥(𝑘) ∈R𝑛𝑥, 𝜂𝑧(𝑘) ∈R𝑛𝑧,𝜂𝛼(𝑘) ∈R𝑛𝛼and 𝜂𝜃(𝑘) ∈R𝑛𝜃are white noise vectors. Since noise estimation [33,34] is not part of the scope of this work, noises are assumed to be white. The covariance matrices are 𝑄𝑥(𝑘),𝑄𝑧(𝑘),𝑄𝑎(𝑘) and 𝑄𝜃(𝑘), respectively. The state vector 𝑥(𝑘)and the RNN coefficients  𝜃(𝑘)are estimated with the EKF updates [35] by augmenting the state vector with the parameter vector using the RNN model of Eqs. (10) and (11) as the transition and observation models 𝑓𝑥and 𝑓𝑧and computing the derivatives at each step, in accordance with the following: 𝐻(𝑘) =[𝜕 𝑓𝑧 𝜕 𝑥0𝜕 𝑓𝑧 𝜕 𝜃𝑧]|||| 𝜃(𝑘|𝑘−1), 𝑥(𝑘|𝑥−1), 𝑢(𝑘) (13a) 𝐾(𝑘) =𝑃(𝑘|𝑘− 1)𝐻(𝑘)𝑇[𝐻(𝑘)𝑃(𝑘|𝑘− 1)𝐻(𝑘)𝑇+𝑄𝑧(𝑘)]−1 (13b) 𝑒(𝑘) =𝑧(𝑘) −𝑓𝑧(𝑥(𝑘|𝑘− 1), 𝑢(𝑘), 𝜃𝑧(𝑘|𝑘− 1)) (13c) [𝑥(𝑘|𝑘)  𝜃(𝑘|𝑘)]=[𝑥(𝑘|𝑘− 1)  𝜃(𝑘|𝑘− 1)]+𝐾(𝑘)𝑒(𝑘)(13d) 𝑃(𝑘|𝑘) = (𝐼−𝐾(𝑘)𝐻(𝑘))𝑃(𝑘|𝑘− 1) (13e) [𝑥(𝑘+ 1|𝑘)  𝜃(𝑘+ 1|𝑘)]=[𝑓𝑥(𝑥(𝑘|𝑘), 𝑢(𝑘), 𝜃𝑥(𝑘|𝑘))  𝜃(𝑘|𝑘)](13f) 𝐴(𝑘) =⎡⎢⎢⎢⎣ 𝜕 𝑓𝑥 𝜕 𝑥 𝜕 𝑓𝑥 𝜕 𝜃𝑥 0 0𝐼0 0 0𝐼⎤⎥⎥⎥⎦|||||||| 𝜃(𝑘|𝑘), 𝑥(𝑘|𝑘), 𝑢(𝑘) (13g) 𝑃(𝑘+ 1|𝑘) =𝐴(𝑘)𝑃(𝑘|𝑘)𝐴(𝑘)𝑇+[𝑄𝑥(𝑘) 0 0𝑄𝜃(𝑘)](13h) The initial covariance matrix 𝑃(0|− 1) is chosen to incorporate 𝓁2-regularization by analogy with Newton’s method as stated by [17]: 𝑃(0|− 1) =[1 𝑁 𝜌𝑥 𝐼0 01 𝑁 𝜌𝜃](14) where 𝑁is the number of instances in the training set and 𝜌𝑥and 𝜌𝜃 are the penalizing terms on 1 2‖𝑥‖2 2and 1 2‖𝜃‖2 2. The EKF is applied by processing the data for several epochs, each one corresponding to a one-day simulation. This process involves comparing the estimation error of the RNN and the correct output on a test set. The stopping criteria are selected to ensure that the number of epochs falls within a specified range, and the mean squared error (MSE) on the validation set is neither constant nor increasing for a certain number of epochs. To recover the initial state and covariance matrix after each epoch, the extended Rauch-Tungg-Striebel smoother was applied [36]. This smoother first applies the EKF to the new data using the previous state and covariance matrix and then performs smoothing. For the present problem, the input vector contains the inputs to the concentrated parameter model, so 𝑢= (𝑇in, 𝑇𝑎, 𝐼⋅𝑛𝑜, 𝑞)𝑇, the outputs are the temperatures measured at the central point of every collector, the outlet temperature as 𝑧= (𝑇1, 𝑇2, 𝑇3, 𝑇4, 𝑇out)𝑇, and the fault vector contains the three aforementioned parameters 𝛼= (𝛼𝐾opt , 𝛼𝑞, 𝛼𝐻l)𝑇. 4. Fault detection and estimation Once the system has been modeled, a second EKF is applied online to identify the values of the faults using the RNN trained as described in Section 3. In this case, the parameters 𝜃of the model in Eq. (12) are fixed, while the inputs 𝛼are estimated. Since the three fault parameters are highly correlated, three parallel EKFs are applied, one for each type of fault, using the RNN model of Eqs. (10) and (11) as the transition and observation models 𝑓𝑥and 𝑓𝑧. For each fault 𝛼𝑖∈𝛼, the new EKF process is given by: 𝐻(𝑘) =[𝜕 𝑓𝑧 𝜕 𝑥 𝜕 𝑓𝑧 𝜕 𝛼]||||𝛼𝑖(𝑘|𝑘−1), 𝑥(𝑘|𝑥−1), 𝑢(𝑘) (15a) 𝐾(𝑘) =𝑃(𝑘|𝑘− 1)𝐻(𝑘)𝑇[𝐻(𝑘)𝑃(𝑘|𝑘− 1)𝐻(𝑘)𝑇+𝑄𝑧(𝑘)]−1 (15b) 𝑒(𝑘) =𝑧(𝑘) −𝑓𝑧(𝑥(𝑘|𝑘− 1), 𝑢(𝑘), 𝜃𝑧(𝑘|𝑘− 1)) (15c) [𝑥(𝑘|𝑘) 𝛼𝑖(𝑘|𝑘)]=[𝑥(𝑘|𝑘− 1) 𝛼𝑖(𝑘|𝑘− 1)]+𝐾(𝑘)𝑒(𝑘)(15d) 𝑃(𝑘|𝑘) = (𝐼−𝐾(𝑘)𝐻(𝑘))𝑃(𝑘|𝑘− 1) (15e) [𝑥(𝑘+ 1|𝑘) 𝛼𝑖(𝑘+ 1|𝑘)]=[𝑓𝑥(𝑥(𝑘|𝑘), 𝑢(𝑘), 𝜃𝑥(𝑘|𝑘)) 𝛼𝑖(𝑘|𝑘)](15f) 𝐴(𝑘) =[𝜕 𝑓𝑥 𝜕 𝑥 𝜕 𝑓𝑥 𝜕 𝛼𝑖 0𝐼]||||||𝜃(𝑘|𝑘), 𝑥(𝑘|𝑘), 𝑢(𝑘) (15g) 𝑃(𝑘+ 1|𝑘) =𝐴(𝑘)𝑃(𝑘|𝑘)𝐴(𝑘)𝑇+[𝑄𝑥(𝑘) 0 0𝑄𝛼(𝑘)](15h) During plant operation, three estimations are made, each assuming that the rest of the fault parameters are set to 1 (i.e., no more faults). In this context, to decide which is the correct one, each estimate 𝛼𝑖(𝑘) is used to feed the neural network, resulting in different output vector estimates 𝑧𝑖(𝑘)for every fault 𝑖∈ {𝐾opt, 𝑞 , 𝐻l}. For instance, 𝑥𝐾opt (𝑘) and 𝑧𝐾opt (𝑘)are the estimated state and output, assuming there is only a fault in the optical efficiency 𝐾opt. The entire process is illustrated in Fig. 4. This iterative estimation process occurs during the dynamic operation of the plant. 4.1. Fault selection with neural network In order to select which of the three estimates is the correct one, a fault detection phase is implemented based on the previous estimates using the EKF. For this purpose, we trained a second feedforward neural network classifier. It is constructed on the basic structure of the multilayer perceptron (MLP) with input-transformation blocks added at the input, as depicted in Fig. 5. The output of the ANN represents the likelihood of each fault and the case without fault in a vector 𝑦(𝑘) ∈R𝑛𝛼+1. In this case, 𝑦(𝑘) = (𝑦faultless(𝑘), 𝑦𝐾opt (𝑘), 𝑦𝑞(𝑘), 𝑦𝐻l(𝑘)). ISA Transactions 158 (2025) 272–284 276 S. Ruiz-Moreno et al. Fig. 4. Scheme of the fault parameters and output estimates. An EKF is applied three times to independently estimate the faults and feed the RNN model, assuming the remaining faults are 1. This is performed to obtain the different output estimates that will be used to isolate the correct fault and discard the rest. Fig. 5. Scheme of the fault detection phase. This fault detection phase aids in determining the most plausible fault scenario based on the estimates generated during plant operations. The first block obtains the errors 𝐸𝐾opt ,𝐸𝑞and 𝐸𝐻lover the actual outputs as 𝐸𝑖(𝑘) =𝑧𝑖(𝑘)∕𝑧(𝑘), where 𝐸𝑖(𝑘) ∈R𝑛𝑧. The second block performs a preliminary fault detection by selecting the type of fault corresponding to the minimum estimation error 𝐸𝑖(𝑘)in the vector 𝑦0∈ R𝑛𝛼+1. If the estimated fault is approximately 1 (by 5%), it is considered that there is no fault. The third block selects the system outputs that will be used as inputs to the MLP. For this problem, only the outlet temperature 𝑇out was considered. The ANN underwent training utilizing the scaled conjugate gradient backpropagation algorithm [37]. 5. Evaluation The methodology is evaluated according to (1) fault detection and isolation capabilities and (2) fault reconstruction capabilities. For the first purpose, the classification accuracy and F1-scores are computed as: 𝐴𝑐 𝑐=𝑇 𝑁+𝑇 𝑃 𝑇 𝑃+𝐹 𝑁+𝑇 𝑁+𝐹 𝑃(16) 𝐹1 = 2⋅𝑅𝑒𝑐 ⋅𝑃 𝑟𝑒 𝑅𝑒𝑐 +𝑃 𝑟𝑒 ,where {𝑃 𝑟𝑒 =𝑇 𝑃 𝐹 𝑃+𝑇 𝑃 𝑅𝑒𝑐 =𝑇 𝑃 𝐹 𝑁+𝑇 𝑃 (17) where TP is the correct identification of faults, TN is the correct identification of the absence of faults, FP is the incorrect identification of faults, and FN is the failure to identify actual faults. The fault reconstruction effectiveness is evaluated with the average estimation error 𝐸𝑖 𝑖(𝑘)obtained for the actual fault 𝑖with its corresponding EKF, as described in Section 4.1. 6. Simulation results The methodology was tested by simulating the ACUREX plant. For this purpose, two datasets were created: one obtained from synthetic irradiances representing sunny days and the other consisting of real irradiance profiles obtained from the Plataforma Solar de Almería data with different types of clouds. All computations were performed in MATLAB R2020b with Intel®Core™i7-9700F CPU at 3 GHz and 16 GB RAM using CasADi [38] for automatic differentiation. 6.1. Sunny dataset To create the dataset, 450 single-day simulations were run, collecting data every 39 s with random inputs, disturbances, and fault values. All the simulations utilized synthetic irradiance profiles of one day of various random shapes, assuming sunny days. The peak irradiances of these profiles ranged from 750 W/m2to 1000 W/m2. The reference temperatures varied from 200 ◦C to 300 ◦C, with steps up to 15 ◦C every hour. The fault values were fixed throughout the day and selected so that only one type of fault co-occurred. Specifically, faults in the optical efficiency ranged from 0.1 to 0.9, faults in the flow rate were from 0.5 to 1.5, and faults in the thermal losses ranged from 1.1 to 1.2. At the end of the day, the fault was restored to its nominal value and a new time series began with a new fault, a new reference temperature, and a new irradiance profile. With this, it is assumed that the system is not working during shutdowns of the plant, and an operator will fix the failure. The dataset was divided into training, validation, and test sets of 320, 80, and 50 instances, respectively. The learning process was carried out for the entire batch with a minimum of 15 epochs and a maximum of 5000 epochs, stopping training when the error remained constant or increased for four consecutive steps. The activation functions employed were hyperbolic tangent across all layers, excluding the final layer, which utilizes a linear function. The rest of the RNN parameters and hyperparameters were selected in a process of trial and error by comparing the mean squared error on the validation set, and the data were normalized. The selected RNN has 𝑛𝑥= 2,𝐿𝑥= 2with 20 and 10 neurons in each layer, and directly a linear connection for the output with 𝐿𝑦= 1and one neuron. Additionally, 𝑄𝑥(𝑘) = 0.1𝐼,𝑄𝑧(𝑘) = 100𝐼, and 𝜌𝑥=𝜌𝜃= 0.001. Fig. 6shows the results of four random, closed-loop experiments of the validation set with the selected RNN architecture, one corresponding to each type of fault. This neural network was trained with 22 epochs in 32.5890 min, achieving average MSEs of 2.0638 ⋅10−3 on the training set, 1.6604 ⋅10−3 on the validation set and 8.2602 ⋅10−4 on the test set. The graphs show the good adaptation of the neural network in its five outputs with minimal deviations, even when fluctuations occur. This confirms the behavior of this RNN as a model of the plant. Once the neural network was trained, the second EKF was applied with the components of 𝑃(0|− 1) =𝑑 𝑖𝑎𝑔(102,102,10−2,10−2,10−2), 𝑄𝑥= 10−8𝐼,𝑄𝛼=𝑑 𝑖𝑎𝑔(10−3,10−3,10−3)and 𝑄𝑧= 10−4. The estimates were smoothed with a low-pass filter of 15 min before being passed to the feedforward neural network. Fig. 7shows an example of the three estimates in a case with a fault of 𝛼𝑞= 0.625. Judging from the graph, the estimates only approached the real value for the flow rate estimator, as one would expect, with a small error. The estimates obtained considering only the optical efficiency and thermal losses are affected by the flow rate fault, producing wrong values and evidencing the importance of a decoupling strategy. Since the faults are strongly coupled, the next step needed is to isolate the faults and select the correct estimate, discarding the other two. It is also important to notice that the perturbations in the temperature do not affect the EKF estimations. Regarding time consumption, the mean time of each iteration of the EKF was 0.013 s. The second ANN was trained with the following hyperparameter values: 𝜆= 5⋅10−7 that regulates the indefiniteness of the Hessian matrix, 𝜎= 5⋅10−5 that determines changes in the weighting of the second derivative approximation, a maximum number of epochs of 4⋅103, a minimum gradient of 10−6and a maximum of 6 validation checks. All layers have hyperbolic tangent, except for a softmax function in the last layer. The ANN consists of three hidden layers, with ISA Transactions 158 (2025) 272–284 277 S. Ruiz-Moreno et al. Fig. 6. Modeling results of the RNN on four of the experiments of the sunny validation set. Outlet temperature and temperatures at the center of each collector obtained from the sensors and the RNN. Fig. 7. Simulation conditions and results of the parameter estimation with the three parallel EKFs in an experiment with 𝛼𝑞= 0.625 of the sunny validation set. 200, 100, and 50 neurons in each one. Fig. 8shows the outputs of the feedforward neural network in four different random experiments, each corresponding to a different faulty case. In these examples, the higher output corresponded to the correct type of fault most of the time, with ISA Transactions 158 (2025) 272–284 278 S. Ruiz-Moreno et al. Fig. 8. Outputs of the feedforward neural network on four experiments of the sunny validation set. Table 2 Evaluation metrics of the methodology on the training, validation and test sets. Without feedforward ANN With feedforward ANN Training Validation Test Training Validation Test Average estimation error 2.07 2.37 2.77 2.07 2.37 2.77 in the faulty parameter (%) F1 in the faultless case (%) 94.55 97.30 96.08 97.01 100.00 98.04 F1 in the 𝐾opt fault (%) 77.60 85.19 78.18 87.80 84.62 81.62 F1 in the 𝑞fault (%) 73.91 77.42 76.40 88.89 92.31 82.76 F1 in the 𝐻lfault (%) 89.61 94.74 92.93 94.87 83.87 91.09 Classification accuracy (%) 84.06 88.75 86.00 92.19 90.00 88.50 some errors at the beginning of the day. To avoid these errors, alarms were not triggered immediately; instead, they were activated after data had been recorded during a specific time window. In this case, this activation was performed at the end of the day. In the four cases, the faults were correctly classified. There was a small degree of doubt in the faultless day and the day with a fault in 𝐻l, but the neural network’s output was able to distinguish the correct class effectively. The initial 𝐾opt classification during the first hour of the faultless day indicates the need not to trigger alarms immediately. Table 2summarizes some average estimation and classification errors computed for the training, validation, and test subsets at the end of each day, both using only the preliminary fault selection without feedforward neural network 𝑦0and using the feedforward neural network 𝑦. The estimation errors are computed for the actual fault type instead of the estimated fault type to allow a separate analysis between estimation and classification behavior. The accuracies and F1scores improve with the feedforward ANN in all classes except for the thermal losses, although the accuracy is still higher with the ANN. The results show the good behavior of the FDD system, with accuracies and F1-scores around 90% and close to 100% in the faultless case. Both estimation errors are less than 3%, with even lower values in the test set than in the training set, which could be due to the randomness of the experiments. The results of the table also show the ability of the feedforward ANN to improve the classification accuracy around 2 points in validation and test subsets. To prove that the training is faster with EKF than with GD, Fig. 9 shows the evolution of the training MSE with the number of epochs and the training time. The EKF algorithm is compared with (i) GD with a learning rate of 10−4, (ii) RMSprop with a learning rate of 10−5.2 and 𝛽= 0.8with five initial steps of simple GD, (iii) RMSprop with a learning rate of 10−4and 𝛽= 0.9, (iv) Adam with a learning rate of 10−6,𝛽1= 0.87 and 𝛽2= 0.65, and (v) Adam with a learning rate of 10−3,𝛽1= 0.806 and 𝛽2= 0.999 with five initial steps of gradient descent, all of them with gradient clipping. A random search was performed in a coarse-to-fine scheme to select the hyperparameters. It is worth noting that GD performs better than RMSprop and Adam. In the only experiments where RMSprop and Adam were faster than EKF and GD, the weights ended up unstabilizing, as one can see in the figure. This can be due to the nature of this specific problem, which does not benefit from adaptive learning rates provided by Adam and RMSprop. Although each epoch of the EKF is slower than GD, the number of epochs needed to obtain a low MSE and the total training time are much lower. A similar experiment was carried out by Bemporad [17] with Adam. There are several aspects that justify the fast training with EKF. First, gradient descent relies on first-order gradients, but EKF uses second-order information through a covariance matrix. This way, EKF can adapt its updates more intelligently based on the local curvature, allowing for more precise parameter updates and fewer iterations. ISA Transactions 158 (2025) 272–284 279 S. Ruiz-Moreno et al. Fig. 9. Comparison of training with EKF and GD. Fig. 10. Modeling results of the RNN on four of the experiments of the cloudy validation set. Outlet temperature and temperatures at the center of each collector obtained from the sensors and the RNN. Moreover, the Kalman gain reduces the impact of noisy measurements, allowing a stable and rapid convergence. 6.2. Cloudy dataset As done with the sunny dataset, 450 single-day simulations were run, collecting data every 39 s with random values of inputs, disturbances, and faults. These simulations were performed using irradiances from actual days when different types of clouds passed by. The reference temperatures and fault values were introduced under the same conditions as those used for the sunny dataset. The dataset was divided into training, validation and test sets of 320, 80, and 50, instances, respectively. Each subset was obtained from different irradiance profiles. After trial and error, an RNN was trained using hyperbolic tangent and linear functions. The RNN has the same characteristics as the one used for the sunny dataset and was trained in 18 iterations and 27.1178 min. This neural network achieved average MSEs of 5.9542 ⋅10−3 on the training set, 1.7679 ⋅10−3 on the validation set and 3.6720 ⋅10−3 on the test set. Fig. 10 shows the adaptation of the RNN to four experiments of the validation set. Again, the RNN performs satisfactorily and adapts well to the system output, although in this case the error is slightly greater than with the sunny day due to the complex dynamics introduced by the clouds. To estimate the fault parameters, the second EKF was applied with the components of 𝑃(0|− 1) =𝑑 𝑖𝑎𝑔(102,102,10−2,10−2,10−2),𝑄𝑥= ISA Transactions 158 (2025) 272–284 280