Full text
Contents lists available at ScienceDirect Energy Conversion and Management journal homepage: www.elsevier.com/locate/enconman Research Paper Model design for photovoltaic facilities based on fuzzy neural network as core of its digital twin William D. Chicaizaa,b, Alex O. Topac, Adolfo J. Sánchezd, Juan M. Escañoa,∗, J.D. Álvarezc aDepartment of Systems Engineering and Automatics, University of Seville, Camino de los Descubrimientos, Sevilla, 41092, Spain bENGREEN - Laboratory of Engineering for Energy and Environmental Sustainability. University of Seville, Spain cDepartment of Informatics, CIESOL-ceiA3, University of Almería, Ctra. Sacramento s/n, La Cañada de San Urbano, Almería, 04120, Spain dDepartment of Mechanical, Biomedical, and Manufacturing Engineering, Munster Technological University, Bishopstown, T12 P928, Cork, Ireland A R T I C L E I N F O Keywords: Digital twin Fuzzy inference system Neurofuzzy Modeling Model predictive control Twinning A B S T R A C T This study presents the development of the core of a digital twin for a photovoltaic (PV) facility located at CIESOL-Almería. Two modeling approaches are proposed: a physics-based model using an equivalent electrical circuit, and a data-driven neurofuzzy model based on an adaptive neuro-fuzzy inference system (ANFIS). The neurofuzzy model, designed as a gray-box system, offers high interpretability and adaptability, and stands out for its rapid synchronization capability with the physical asset, enabling real-time behavior modeling essential to the digital twin framework. The ANFIS-based model accurately captures the dynamic power output of the PV system and is suitable for integration into energy management strategies based on predictive modeling. The model exhibits strong predictive performance, with a worst-case mean absolute error of only 16.37W, a standard deviation of 126.22W, a standard error of 0.73W, and a coefficient of determination of 0.99, indicating high consistency and accuracy. When compared to the equivalent electrical circuit model and a previously published artificial neural network applied to a lower-capacity PV system, the neurofuzzy model demonstrates superior accuracy. Specifically, the normalized mean absolute error and normalized root mean square error are 0.0036 and 0.0282 per watt, respectively, outperforming both the equivalent electrical circuit model (0.01982 and 0.0520 per watt) and the neural network approach. These differences represent relative improvements of 87.39 % and 61.55 % over the neural network benchmark. In addition, the neurofuzzy model requires significantly lower computational resources, making it suitable for real-time applications and implementation in industrial controllers. The results confirm the potential of gray-box neurofuzzy modeling as a core component of a digital twin, providing a reliable, efficient, and interpretable foundation for control, monitoring, and optimization of PV installations. 1. Introduction In recent years, the energy sector has been responsible for approximately 75% of global greenhouse gas emissions. In response, the Net Zero Emissions (NZE) by 2050 scenario was proposed in 2022, mainly supported by countries with advanced economies. The objective is to achieve net-zero energy-related CO2 emissions by 2050 [1]. A key strategy to meet this goal involves reducing reliance on Fossil Energy Sources (FESs) and accelerating the deployment of Renewable Energy Sources (RESs) [2]. According to [3], the Earth receives enough solar energy in just 90 min to meet the world’s total annual energy demand. Therefore, solar energy stands out as a clean, renewable, and virtually inexhaustible resource, accessible in nearly all geographical ∗Corresponding author. E-mail address: [email protected] (J.M. Escaño). locations. Among the various RES technologies, solar Photovoltaic (PV) is currently the only one on track to meet the NZE targets by 2050 [4]. Its proper deployment offers an effective solution to environmental and energy challenges [5]. Solar PV is one of the most promising RESs for achieving energy sustainability [6,7], especially due to its installation flexibility compared to other RES technologies such as wind, geothermal, tidal, or hydroelectric [8]. In 2023, solar PV accounted for nearly 75% of newly installed renewable capacity worldwide. Moreover, the total renewable capacity is expected to grow steadily over the next five years, with solar PV and wind forecast to comprise a record 96% of new installations (see Fig. 1). This trend is largely due to their decreasing generation costs, https://doi.org/10.1016/j.enconman.2025.120001 Received 20 January 2025; Received in revised form 19 May 2025; Accepted 25 May 2025 Energy Conversion and Management 342 (2025) 120001 Available online 21 June 2025 0196-8904/© 2025 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY-NC-ND license ( http://creativecommons.org/licenses/bync-nd/4.0/ ).
W.D. Chicaiza et al. Nomenclature Abreviations AC Alternating Current ANFIS Adaptive Neuro-Fuzzy Inference System ANN Artificial Neural Network BESS Battery Energy Storage System CIESOL Solar Energy Research Center (Centro de Investigaciones en Energía Solar) DC Direct Current DT Digital Twin DTI Digital Twin Instance EEC Equivalent Electrical Circuit EMS Energy Management System ESS Energy Storage System EV Electric Vehicle FESs Fossil Energy Sources FFNN Feedforward Neural Network FIS Fuzzy Inference System FL Fuzzy Logic MAE Mean Absolute Error MF Membership Function MG Microgrid MGDT Microgrid Digital Twin MLR Multiple Linear Regression MPC Model Predictive Control MPPT Maximum Power Point Tracking MSE Mean Squared Error NF NeuroFuzzy nRMSE Normalized Root Mean Square Error NZE Net Zero Emissions OPC-UA Open Platform Communications Unified Architecture PCA Principal Component Analysis PLC Programmable Logic Controller PV Photovoltaic RES Renewable Energy Source RMSD Root Mean Square Deviation RMSE Root Mean Square Error SCADA Supervisory Control and Data Acquisition SVM Support Vector Machine Variables 𝐸Mean estimation error [W m−2] 𝛥Sky brightness (dimensionless) 𝜖Sky clearness (dimensionless) 𝐼𝑇Estimated tilted irradiance [W m−2] 𝑃𝐷𝐶 Predicted DC power [W] 𝑥 Predicted output 𝜅Sky clearness adjustment constant (dimensionless) 𝐃Measurement matrix which are now lower than both fossil and non-fossil alternatives in most countries. By 2028, the installed capacity of solar PV and wind energy is expected to double compared to 2022, reaching nearly 710 GW [9]. The building sector — which includes residential, commercial, and public infrastructure — accounts for approximately 30% of total global 𝐓𝐞𝐬𝐭 Testing set of input variables 𝐓𝐫𝐧 Training set of input variables 𝐕𝐚𝐥 Validation set of input variables 𝐙Matrix of normalized input variables Number of rules in fuzzy inference system CCovariance matrix 𝜇Gaussian membership function ℜSet of real numbers 𝜌Pearson correlation coefficient (dimensionless) 𝜎Standard deviation 𝜎𝐸Standard deviation of the estimation error [W] 𝜎𝑖,𝑗 , 𝑐𝑖,𝑗 Parameters of Gaussian membership functions 𝑎Weighted solid angle (circumsolar) 𝑏Weighted solid angle (horizon) 𝐹1Circumsolar brightening coefficient 𝐹2Horizon/zenith brightening coefficient 𝐹𝑖𝑗 Fuzzy set 𝐼𝐷𝐶 DC output current [A] 𝐼ℎ,𝑏 Direct normal irradiance [W m−2] 𝐼ℎ,𝑑 Diffuse horizontal irradiance [W m−2] 𝐼ℎGlobal horizontal irradiance [W m−2] 𝐼𝑇Solar irradiance on tilted surface [W m−2] 𝑃𝐷𝐶 DC power generated by the PV system [W] 𝑅2Coefficient of determination (dimensionless) 𝑅𝑏Geometric factor: ratio of tilted to horizontal solar beam irradiance 𝑆𝐸 Standard error of the estimate [W] 𝑇𝑠Sampling time [s] 𝑇𝑎𝑚𝑏 Ambient temperature [◦C] 𝑡𝑡𝑤 Total twinning time [min] 𝑡𝑡𝑤∕day Twinning time per day [s] 𝑡𝑡𝑤∕sample Twinning time per sample [ms] 𝑉𝐷𝐶 DC output voltage [V] 𝑥Actual output energy consumption and 28% of CO2 emissions. Consequently, transitioning these buildings to RES-powered systems is critical for meeting NZE targets by 2050 [10,11]. In many cases, these systems are integrated with Energy Storage Systems (ESS) within a controlled energy environment known as a MicroGrid (MG). An MG is a small-scale energy system designed to meet local energy demands efficiently by combining RESs, FESs, ESS, and controllable and non-controllable loads. It can operate either in grid-connected mode or in islanded mode [12– 14]. However, the stochastic behavior of RESs introduces challenges in MG operation due to the variability in power estimation across different modeling approaches. To address this, digital twins have recently emerged as a valuable tool for improving the prediction and management of renewable generation [15–17]. Due to increasing energy prices and the need to reduce operational costs, many buildings are now incorporating RESs — especially PV systems — into their infrastructure. In this context, MGs are increasingly adopted in residential settings, allowing for optimal management of internal energy flows, ensuring the energy supply to local loads, and exporting excess energy to the public grid [18,19]. In residential MGs, PV systems typically contribute the largest share of energy supply, offering a clean and efficient solution to meet demand. This not only Energy Conversion and Management 342 (2025) 120001 2
W.D. Chicaiza et al. Fig. 1. Renewable electricity capacity additions by technology and segment from 2016 to 2028 [9]. optimizes energy resources but also reduces greenhouse gas emissions and supports the transition to sustainable energy systems [14,20]. To build modern MGs, these systems must operate as monitored and controlled entities in real-time, exhibiting flexibility, resilience, and efficiency. This implies the need for adaptability to integrate advanced digital technologies, autonomous operation in islanded mode, and intelligent decision-making to optimize various resources and operating conditions [21]. Emerging approaches to improve the energy efficiency of MGs are increasingly being developed through the evolution of traditional modeling methods into Digital Twin (DT) architectures. A DT is a digital information construct composed of a set of models whose fidelity or granularity aligns with its intended purpose. Its distinguishing feature is the bidirectional synchronization with its physical counterpart, enabling continuous data and information exchange [22]. This realtime interaction allows the virtual representation to reflect the actual operating conditions, providing reliable input to forecasting models and enabling accurate decision-making [23,24]. DT technology facilitates continuous improvement across both the design and operational phases of a system. Its greatest value lies in the design phase, where it serves as a comprehensive digital replica that supports lifecycle management. DTs have redefined how systems are designed and controlled—extending beyond static modeling and simulation to establish a symbiotic relationship between physical and virtual entities [22]. A core feature of DTs is the seamless integration between physical and digital spaces, grounded in modeling and simulation principles. Although model-based simulation has long supported earlystage design, validation, and planning [25], its use during real-time operation has historically been limited. By developing a DT of a physical asset, virtual support is enabled throughout the asset’s lifecycle, improving design, deployment, commissioning, and operation. The DT supports analysis under diverse conditions, real-time control, ongoing optimization, and informed decision-making—eliminating the need for costly physical experimentation. In this context, product data exists at every stage of its lifecycle, beginning with collaborative digital design. During this phase, digital representations are created — encompassing design models, virtual prototypes, and functional specifications — which together form the DT. This allows virtual testing and refinement of a product’s intended behavior prior to physical deployment. A recent study [22] reports that DT adoption is growing rapidly, particularly in sectors such as aerospace, automotive, energy, and technology. According to the 2022 Digital Twin Global Survey Report [26], 69% of organizations in these sectors currently use DT technology, with 71% of them having invested in the past year. Among non-users, 44% plan to adopt DTs within the next three years, and 11% within six months. Use cases vary: 42% simulate non-connected virtual prototypes, 50% monitor real-time asset health, 47% use DTs for predictive analytics, and 51% employ intelligent digital models for self-diagnosis and corrective action recommendations. The most cited benefits of DT adoption include real-time monitoring (38%), increased efficiency and safety (37%), cost reduction (33%), and improved sustainability in both products and processes. The report also highlights risk assessment as the top benefit, with 60% citing faster and more accurate evaluations. Other reported advantages include reduced product development time (36%), fewer physical prototypes (33%), and fewer simulations (28%), alongside gains in predictive maintenance (34%) and enhanced collaboration across teams (32%). While the advantages of DTs — such as real-time monitoring, improved efficiency, and cost reduction — are well recognized, the report also notes the initial cost considerations. Although upfront investments in DTs may exceed those of traditional methods, long-term benefits such as fewer physical prototypes, faster development cycles, and predictive maintenance often offset the initial expenditures. As a result, DTs represent a cost-effective solution for many organizations, balancing short-term investment with long-term operational gains. In the context of MGs, DT technology is a key enabler for breaking down data silos, allowing integrated data analysis, modeling, simulation, and improved decision-making through artificial intelligence (AI). DTs are applicable across all stages of an MG’s lifecycle—from design to operation and maintenance. Key applications include system design and development, real-time energy management, control and operation, MG protection, predictive maintenance, fault diagnostics, and real-time simulation and testing. These capabilities address major challenges in MG management using advanced technological solutions, thereby enhancing system efficiency and reliability [21,27]. Since the DT paradigm relies on modeling across multiple domains — geometric, physical, behavioral, and rule-based — the model granularity or fidelity must be carefully selected. In sectors such as aerospace and manufacturing, high-fidelity models are typically required. However, as DTs have expanded into a broader range of industries, model fidelity must be adapted to the specific objectives of each domain, as emphasized in the ISO 23247 standard referenced in [22]. Given that most MGs are composed of PV systems, developing a DT for each PV asset facilitates integration into a broader DT of the MG. In this context, each PV DT becomes a Digital Twin Instance (DTI) or unit twin, contributing to more efficient and accurate MG-wide management within a comprehensive DT framework. In general, the use of first-principles physical models in large-scale systems incurs high computational costs and can be time-intensive. To mitigate this, reduced-order models are proposed as an effective alternative. These models, derived from data-driven techniques, approximate the behavior of the physical system while maintaining sufficient accuracy. Reduced-order models can be incorporated into the DT framework, enabling real-time execution of the various applications supported by the DT, including control and optimization. Specifically, Energy Conversion and Management 342 (2025) 120001 3
W.D. Chicaiza et al. in the context of a MicroGrid Digital Twin (MGDT), such models support real-time energy management while reducing computational demands. MGs typically rely on an Energy Management System (EMS) to coordinate internal energy flows. Model Predictive Control (MPC) is a prominent control strategy in this context, offering optimal decisionmaking capabilities for MG operation [28]. Like conventional control methods, MPC compensates for disturbances via feedback. When disturbances are measurable or estimable, they can be explicitly included in the predictive model, allowing the controller to anticipate and counteract their effects, essentially incorporating a feedforward mechanism. This inherent predictive capability enhances MPC performance. However, MPC requires the solution of an optimization problem at each control interval, which can be computationally intensive— particularly when high sampling frequencies are necessary to track fast-changing electrical signals. Given that EMSs are commonly implemented in Programmable Logic Controllers (PLCs) with limited processing power, this computational burden imposes constraints on the achievable sampling time. Consequently, simplified linear models are often used within the MPC, which fail to capture the full dynamic behavior of the MG. MGs typically harness solar and wind energy — mainly through PV panels and wind turbines — to generate electricity. The climatic variables that interact directly with these systems, namely solar irradiance and wind speed, exhibit stochastic (and sometimes ergodic) behavior. Therefore, a nonlinear mathematical model based on differential equations is not ideally suited for systems of this nature [29–31]. Designing optimal control and energy management strategies using nonlinear models often imposes a high computational burden. This complexity becomes a major limitation when the optimization problem must be solved in real time, subject to strict sampling constraints. Consequently, as emphasized in [29], fast and accurate models are required. In the case of power generation from solar resources — specifically PV facilities, which are the subject of this study — several modeling approaches have been proposed. A widely used and straightforward method is the Equivalent Electrical Circuit (EEC) model. This model estimates the energy produced by the PV field based on a stationary representation [14] that does not account for the temporal variation of electrical parameters over the system’s lifetime. The EEC approach estimates the current and voltage at the maximum power point for each sampling instance by solving a system of equations using numerical methods, such as the Newton–Raphson algorithm or other solvers. Another approach for predicting generated power involves datadriven models such as Artificial Neural Networks (ANNs) and Multiple Linear Regression (MLR), as demonstrated in studies like [32,33]. These models typically use ambient meteorological variables as inputs, selected via correlation analysis with the output power. Results from these studies indicate that ANN-based models outperform MLR in terms of prediction accuracy. Similarly, works such as [34,35] have shown that ANN models achieve high accuracy under varying sky conditions. In particular, [35] compared ANN with the Support Vector Machine (SVM) technique and found nearly identical performance under different sky conditions. In this context, the model selected for this study is a type of ANN known as the Adaptive Neuro-Fuzzy Inference System (ANFIS), introduced by Jang [36]. ANFIS combines the learning capabilities of neural networks with the interpretability of fuzzy inference systems, allowing the incorporation of both data-driven learning and expert knowledge via rule-based reasoning. This hybrid approach broadens the spectrum of information used for model adaptation and facilitates a more accurate twinning process between the digital and physical counterparts of the PV system. In the domains of system modeling and control, fuzzy modeling has proven to be particularly effective for handling nonlinearities. Recent studies have demonstrated that both ANN and fuzzy-ANN models (like ANFIS) can accurately replicate the behavior of nonlinear systems within a DT framework [17,22,29,31,37,38]. One notable advantage of ANFIS, similar to traditional ANNs, is its capacity for rapid adaptation and execution. This modeling strategy aligns well with the behavioral and rulebased domain of DT architecture. Within this framework, new information technologies enable a continuous, bidirectional flow of data between the physical asset and its virtual representation. This synchronization or pairing process, known as twinning [22], ensures that the digital model is regularly updated to reflect changes in the physical system. The twinning rate, or frequency of synchronization, is a key characteristic of a DT and allows the digital entity to maintain an accurate representation of the physical system in real or near-real time. Modeling PV facilities within a DT framework requires a detailed understanding of system dynamics and environmental variability to ensure that the DT reflects operational fluctuations with high fidelity. To the best of the authors’ knowledge, no published dynamic model currently addresses the specific characteristics of PV facilities within the DT framework. Although various modeling approaches exist to prediction the power output of PV systems [31,33,39,40], these studies typically do not focus on the integration of such models into the development of a DT. Furthermore, they often overlook one of the core principles of the DT paradigm: the synchronization or twinning process, which enables continuous alignment between the physical and virtual entities throughout the system’s lifecycle. This limitation prevents the models from effectively supporting dynamic, real-time decision-making and operational adaptability, which are essential for modern DT applications. Recent literature on PV system modeling includes limited contributions that address reduced-order models suitable for DT implementations. One of the few relevant studies [32] models a PV installation with a capacity of 160[Wp], approximately 28 times smaller than the 4500[Wp] system analyzed in the present work. The authors propose a feedforward neural network (FFNN) with three hidden layers (6, 4, and 2 neurons, respectively), alongside two MLR models to predict power output. Model inputs are selected via correlation analysis and include four meteorological variables: global solar irradiation, ambient temperature, wind speed, and humidity. The ANN model achieves the best performance, with a mean absolute error (MAE) of 4.570[W], mean squared error (MSE) of 137.682[W], root mean squared error (RMSE) of 11.7733[W], and a coefficient of determination of 0.9159. These results are later compared with those obtained from the modeling strategies proposed in this study. Other DT-related studies, such as [41], focus on immersive monitoring platforms using Unreal Engine to replicate the geospatial and physical domains of PV plants. These approaches prioritize visualization and autonomous inspection but do not address reduced-order dynamic modeling for real-time control or energy management. Similarly, the review in [42] outlines the potential of DTs for predictive maintenance, real-time monitoring, and integration with technologies such as AI, IoT, and autonomous systems. However, the discussion remains conceptual and lacks validated implementations. In contrast, the present work develops and validates a DT for a 4500[Wp] PV installation, its applicability to smart EMS control and optimization in the context of an MG. This work focuses on the development of a DT for a PV facility that is part of an MG. The core of the proposed DT consists of two models: a mathematical model and a reduced-order model, both derived from real data collected at the bioclimatic building of CIESOL, located at the University of Almería. The proposed methodology accounts for the system’s evolution throughout the PV installation’s lifecycle, considering the impact of aging and environmental exposure on critical parameters and physical properties such as solar cells, structural components, and resistive elements. To enable real-time decision-making and synchronization with the physical system, the models are designed to support automation and control tasks—essential elements of the DT paradigm. These models provide a reliable and efficient foundation Energy Conversion and Management 342 (2025) 120001 4
W.D. Chicaiza et al. Fig. 2. CIESOL Research Center. The case study corresponds to the PV facility of the first row of solar panels, next to the field of flat-plate solar collectors. for implementing intelligent energy management strategies within the context of the MGDT. Traditional modeling approaches, such as the EEC model, rely on steady-state assumptions and fixed parameters, making them unsuitable for capturing the dynamic changes associated with system aging. While such models are simple to implement, they fail to reflect performance degradation over time. In contrast, the gray-box methodology proposed here offers both interpretability and adaptability. It adjusts to the real power output of the installation and implicitly accounts for shading effects without requiring complex geometric modeling. Moreover, the model can be rapidly updated with new data, allowing it to reflect the actual behavior of the PV system throughout its operational life. This study makes the following contributions: 1. It presents the development of the core of a DT for a PV facility, comprising both a first-principles mathematical model and a nonlinear NeuroFuzzy (NF) model based on an ANFIS. These models accurately capture the system’s behavior using data collected over 108 consecutive days of operation, including both daytime and nighttime periods, with a sampling interval of one minute. 2. The validated dynamic models constitute the core of the PV facility’s DT and serve as essential components for developing an EMS for the associated MG. They support performance analysis and enable the implementation of intelligent energy management and control strategies within the DT framework. 3. It proposes a methodology for developing a reduced-order, datadriven model — specifically, an NF model — which is benchmarked against a first-principles model and a recently published ANN model. The NF model stands out for its interpretability and transparency through rule-based inference, in addition to offering rapid synchronization (twinning) and execution capabilities. Furthermore, it can be seamlessly integrated into predictive control strategies for EMS and is suitable for implementation in PLCs compliant with the IEC 61131-7 standard. 4. Finally, the study evaluates the computational efficiency of both models during the twinning and simulation phases, establishing their runtime suitability for real-time control and optimization strategies. The rest of the article is organized as follows. Section 2 describes the case study developed in this work. Section 3 discusses data preprocessing and the calculation of irradiance on a tilted surface, which is a critical input for the proposed models. Section 4 presents the correlation analysis that guided the selection of input variables for the NF model and provides a description of the training and testing datasets. Section 5 details the models that comprise the Digital Twin of the PV facility. Section 6 presents and discusses the simulation results of both modeling approaches. Finally, Section 7 summarizes the conclusions and outlines directions for future research. 2. Case study This study is carried out at the Solar Energy Research Center (CIESOL). The center harnesses the potential of solar radiation to develop new solar energy applications in various areas, such as energy use, water treatment, and research into air conditioning and users’ comfort in buildings [43]. The research center, depicted in Fig. 2, is located on the campus of the University of Almería (Almería, Spain). It is a bioclimatic building designed according to energy efficiency criteria, integrating several solar energy technologies such as PV systems, solar thermal collectors, and solar cooling systems. Fig. 2(a) presents the frontal view of the CIESOL bioclimatic building, while Fig. 2(b) details the rooftop PV facility. It can be observed that the PV modules experience partial shading during specific periods of the day, mainly due to their proximity to the adjacent wall, impacting the energy production performance. The MG of the CIESOL building is designed to operate in gridconnected mode, enabling energy exchange with the main utility grid. It comprises a PV facility, a Battery Energy Storage System (BESS), and an Electric Vehicle (EV) charging point. The PV system is connected to a DC/DC converter equipped with an integrated Maximum Power Point Tracking (MPPT) device to optimize energy harvesting. The EV charging point operates as a dispatchable load or generator, depending on its availability. Both the BESS and the EV charging system are connected through dedicated DC/DC converters or chargers. The resulting DC power flows are directed to the input of a DC/AC converter, which, together with the main grid, supplies alternating current (AC) to the building’s electrical loads. Fig. 3 presents a schematic diagram of the MG, illustrating the components that comprise its energy system. The case study of this work focuses on the PV facility of the CIESOL building, which constitutes the physical asset for the development of its Digital Twin (DT). The system consists of 10 solar panels arranged in a series-parallel configuration, capable of delivering a maximum power of 4.5[kWp]. Each panel comprises 144 cells, with the overall configuration including 2 parallel strings of 5 modules connected in series. To characterize the PV facility, both its geographic location and the orientation of its solar panels are considered. Latitude (𝜙) and longitude (𝜓) define its position on Earth, while inclination (𝛽) and azimuthal angle (𝛾) describe the orientation of the panels, essential factors for maximizing solar energy capture. All these parameters are presented in Table 1. To characterize the solar resource at the site, the building is equipped with a 2AP two-axis solar tracker that supports two pyranometers and a pyrheliometer, all manufactured by Kipp & Zonen. These precision instruments measure the three main components of solar irradiance: global horizontal (𝐼ℎ), diffuse horizontal (𝐼ℎ,𝑑 ), and direct normal (𝐼ℎ,𝑏), each expressed in [W∕m2]. Global and diffuse Energy Conversion and Management 342 (2025) 120001 5
W.D. Chicaiza et al. Fig. 3. Conceptual scheme of the MG, which highlights its structure and the relationships between the different elements of the energy system. Fig. 4. Measured solar irradiance components at the CIESOL building: red solid line (global horizontal), blue dash-dot (diffuse horizontal), and yellow dashed (direct normal). Table 1 Angular parameters for the CIESOL PV facility. Parameters Symbol Value Latitude 𝜙36.83◦ Longitude 𝜓−2.41◦ Inclination 𝛽30◦ Orientation 𝛾−21◦ components are measured using CMP 11 pyranometers, the latter installed beneath shading balls to isolate the diffuse fraction, while direct irradiance is measured with a CH 1 pyrheliometer. Fig. 4 illustrates the three irradiance components, and Fig. 5 shows the solar tracking system. Table 2 summarizes the measurement parameters, sensor types, number of units, and specified accuracy. 3. Data pre-processing and analysis The dynamic behavior of the power generated by the PV facility is estimated through the development and validation of both an NF model and a mathematical model, using real data from the installation. The quality of any model—whether mathematical, NF, or based on ANNs— is directly associated with the quality of the data employed during the training stage. Inadequate data quality during training, validation, and testing stages may significantly compromise model accuracy and Fig. 5. SUN 2AP solar tracker installed at the CIESOL building, equipped with two pyranometers and one pyrheliometer. reliability, resulting in ineffective outcomes and suboptimal performance [29,44]. Consequently, ensuring high data quality is critical for optimal model development. Criteria such as accuracy, completeness, consistency, and timeliness are commonly applied to assess data quality [29,45]. To comply with these standards, several data pre-processing procedures are implemented: homogenization of sampling intervals, removal of outliers and inconsistent entries, noise filtering from instrumentation signals, and careful selection of variables involved in the power generation process. This study uses raw data acquired from the SCADA system of the CIESOL building, which records measurements from the installed instrumentation on a daily basis. Data are stored in plain text files, each containing 493 variables (columns). These files are imported into MATLAB as timetable, retaining their respective headers. Daily operation records are concatenated chronologically to form continuous datasets. During this concatenation, only nine variables associated with the PV system are retained. Data exploration reveals the presence of missing measurements, irregular sampling, and incomplete records. Preprocessing of the original timetable involves homogenizing measurement times for each variable. This is achieved by applying linear interpolation with a sampling interval of 1 min. This interval is chosen as it appears in 99.6% of the unprocessed data, making it the most frequently occurring interval in the SCADA data logs. Once daily data are homogenized, they are concatenated into a single timetable containing records from specific months. The resulting data structure consists of four uniform timetables covering the following periods: July 1st–18th and 24th–31st; September 1st–4th and 7th–28th; October 2nd–10th and 16th–18th; December 1st–8th and 13th–20th of 2023; January 8th–18th, 22nd–29th, and 31st; and March 5th–11th and 13th–17th of 2024. In total, this dataset represents 108 days of PV facility operation. The subsequent step in data treatment consists of filtering the information. A low-pass filter is applied to each variable in the resulting timetables. This procedure aims to eliminate noise from the sensor and instrumentation measurements. In addition to reducing noise, the filtering process smooths the data, removes outliers and inconsistent entries — particularly those caused by abrupt irradiance changes due to cloud cover — and ensures the overall quality and consistency of the dataset for further analysis. As an outcome of the preprocessing stage, a dataset containing 108 days of observations is prepared. This dataset comprises 147,640 samples across 9 variables and constitutes the foundation for selecting the variables involved in the power generation process. In studying the impact of solar radiation on PV power generation, it is important to understand the characteristics of incident solar irradiance. The solar irradiance incident on tilted surfaces (𝐼𝑇) comprises Energy Conversion and Management 342 (2025) 120001 6
W.D. Chicaiza et al. Table 2 Instruments used to measure solar irradiance at the CIESOL facility. Parameter Unit Sensor type Number of sensors Accuracy (manufacturer) Global horizontal [W∕m2]Pyranometer (Kipp & Zonen CMP 11) 1 ±2% Diffuse horizontal [W∕m2]Pyranometer (Kipp & Zonen CMP 11) 1 ±2% Direct normal [W∕m2]Pyrheliometer (Kipp & Zonen CH 1) 1 ±2% All measurements are expressed in [W∕m2]. Sensor models and accuracy are provided according to manufacturer specifications. the sum of several radiation components (see Eq. (1)), including direct (beam) radiation from the Sun, various components of diffuse radiation from the sky, and diffusely reflected radiation from surrounding surfaces (e.g., the ground) incident on the inclined surface. Several approaches for estimating these irradiance components under fluctuating weather conditions — characterized by highly variable atmospheric phenomena, such as sudden changes in solar radiation due to cloud cover that may affect sensor measurements — are presented in [46,47]. However, these approaches are not applicable to the instrumentation employed in this study, as the reliability of the collected data has been independently verified. 𝐼𝑇=𝐼𝑇 ,𝑏 +𝐼𝑇 ,𝑑 +𝐼𝑇 ,𝑟𝑒𝑓𝑙 (1a) 𝐼𝑇=𝐼𝑇 ,𝑏 +𝐼𝑇 ,𝑑,𝑖𝑠𝑜 +𝐼𝑇 ,𝑑,𝑐𝑠 +𝐼𝑇 ,𝑑,ℎ𝑏 ⏟⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏟⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏟ 𝐼𝑇 ,𝑑 +𝐼𝑇 ,𝑟𝑒𝑓𝑙 (1b) Diffuse radiation from the sky (Eq. (1b)) consists of three distinct components: an isotropic component (𝐼𝑇 ,𝑑,𝑖𝑠𝑜), which is received uniformly from all parts of the sky; diffuse circumsolar radiation (𝐼𝑇 ,𝑑,𝑐𝑠), resulting from the forward scattering of solar radiation and concentrated around the Sun; and the horizon brightening component (𝐼𝑇 ,𝑑,ℎ𝑏), which is concentrated near the horizon and becomes more pronounced under clear sky conditions, as described in [48]. The estimation of solar irradiance on inclined surfaces is typically performed using sky models, which may be either isotropic or anisotropic. The isotropic model offers a straightforward formulation but generally underestimates irradiance on tilted planes, as it considers only the isotropic diffuse component in its calculation. Its mathematical expression is provided in Eq. (2). 𝐼𝑇=𝐼ℎ,𝑏 ⋅𝑅𝑏+𝐼ℎ,𝑑 (1+ cos (𝛽) 2)+𝐼ℎ⋅𝜌𝑔(1− cos (𝛽) 2)(2a) 𝑅𝑏=cos (𝜃) cos (𝜃𝑧)(2b) where 𝜌𝑔, 𝑅𝑏, 𝜃, and 𝜃𝑧 are the reflectivity coefficient, the ratio of tilted and horizontal solar beam irradiance, the angle of incidence, and the zenith angle, respectively. Anisotropic models constitute enhanced formulations that account for additional components of diffuse radiation, such as circumsolar and horizon brightening, on inclined surfaces. Common examples include the models developed by Klucher, Hay-Davies, Reindl, Muneer, the Perez models from 1990 and 1998, and the Perez-Driesse formulation. These approaches differ in the way they incorporate the various diffuse components [48]. For instance, the Hay-Davies model estimates the fraction of diffuse radiation that is circumsolar and assumes it follows the same direction as direct radiation. The Reindl model extends this formulation by adding a horizon-brightening term, originally introduced in the Klucher model. The Perez 1990 transposition model is widely used to estimate global irradiance on tilted surfaces. It provides a detailed representation of isotropic diffuse, circumsolar, and horizon-brightening components by incorporating empirically derived coefficients (𝑓11, 𝑓12, 𝑓13, 𝑓21, 𝑓22) [49,50]. This model is formulated as shown in Eq. (3). 𝐼𝑇=𝐼ℎ,𝑏𝑅𝑏+𝐼ℎ,𝑑 [(1−𝐹1)(1+ cos 𝛽 2)⋯ +𝐹1𝑎 𝑏+𝐹2sin 𝛽]+𝐼ℎ⋅𝜌𝑔(1− cos 𝛽 2)(3a) 𝑎= max(0,cos 𝜃)(3b) 𝑏= max(cos 85◦, 𝜃𝑧)(3c) 𝐹1= max [0,(𝑓11 +𝑓22𝛥+𝜋𝜃𝑧 180𝑓33)] (3d) 𝐹2=𝑓21 +𝑓22𝛥+𝜋𝜃𝑧 180𝑓23 (3e) where 𝑎 and 𝑏 represent the cosine of the angles of circumsolar incidence on the tilted and horizontal planes, respectively. The anisotropic coefficients 𝐹1 and 𝐹2 depend on the sky condition parameters clearness (𝜖) and brightness (𝛥), as defined in Eq. (4): 𝜖= 𝐼ℎ,𝑑 +𝐼ℎ,𝑏 𝐼ℎ,𝑛 +𝜅𝜃𝑧3 1+𝜅𝜃𝑧3(4a) 𝛥=𝑚𝐼ℎ,𝑑 𝐼𝑜𝑛 (4b) where 𝐼𝑜𝑛 is the direct extraterrestrial normal irradiance, and 𝜅 is a constant equal to 1.041 when the solar zenith angle (𝜃𝑧) is in radians, as noted in [50]. A known limitation of the Perez 1990 model is its reliance on empirical coefficients defined over discrete intervals of the sky brightness parameter, which may introduce discontinuities in the irradiance estimation. To overcome this limitation, the Perez-Driesse model [50] replaces the empirical look-up table with continuous quadratic splines. In this formulation, the sky clarity index (𝜖) is transformed into a bounded equivalent parameter (𝜁), resulting in a smoother and more reliable estimation. Sky models are mathematical representations of diffuse radiation, differing primarily in the method used to compute this component. When combined with beam and reflected radiation, they enable the estimation of irradiance on inclined surfaces based on horizontal measurements. The mathematical formulations of isotropic and anisotropic models, along with detailed parameter descriptions, are provided in [48,51,52]. A comprehensive explanation of the Perez-Driesse model is available in [50]. The sky models described above are applied to estimate the incident solar irradiance on the inclined surface of an experimental module at the CIESOL facility, which is installed at an angle of 22◦. The resulting estimates are compared against measured data, as shown in Fig. 6. As previously mentioned, the actual irradiance (𝐼𝑇) is obtained using a Kipp & Zonen CMP 11 pyranometer, with a reading accuracy of ±2%. Fig. 6(a) illustrates the irradiance on the tilted surface as estimated by each sky model, while Fig. 6(b) presents a zoomed-in view focused on the morning and midday periods. These close-up views facilitate a clearer understanding of the differences among the 𝐼𝑇 estimates. The measured values (𝐼𝑇) are compared with the model estimates ( 𝐼𝑇), and the corresponding statistical metrics are presented in Table 3. A detailed description of these metrics is provided in Section 6. Based on the error metrics, the isotropic model exhibits the highest arithmetic mean error ( 𝐸) in estimated irradiance, with a value of 18.152[W∕m2], followed by the King model at 13.896[W∕m2]. In contrast, the Klucher model shows the lowest mean error at 6.023 [W∕m2], followed by the Perez and Perez-Driesse models. Regarding variability, quantified by the standard deviation (𝜎𝐸), the isotropic model again presents the highest value at 28.117[W∕m2], while the Klucher, Perez, and Perez-Driesse models exhibit the lowest deviations, indicating more consistent estimations. Doubling the Energy Conversion and Management 342 (2025) 120001 7
W.D. Chicaiza et al. Fig. 6. Comparison of irradiance on a tilted surface measured vs. irradiance calculated with various models. Table 3 Error indexes obtained between actual and calculated irradiance on a tilted surface. Indexes Models Isotropic Klucher Hay-Davies Reindl King Perez Perez-Driesse 𝐸[W∕m2]18.152 6.023 11.528 11.329 13.896 7.311 7.128 𝜎𝐸[W∕m2]28.117 13.659 19.483 19.268 23.637 11.994 11.947 2𝜎𝐸[W∕m2]56.235 27.317 38.966 38.537 47.273 23.988 23.894 𝑅𝑀𝑆𝐸 [W∕m2]7.435 8.264 5.875 5.894 8.130 3.839 3.866 𝑅2−0.9995 0.9994 0.9997 0.9997 0.9994 0.9999 0.9999 standard deviation (2𝜎𝐸) defines a 95% confidence interval for the error variability; the lowest intervals are observed for the Klucher, Perez, and Perez-Driesse models, suggesting reduced uncertainty in their predictions. In terms of precision, as measured by the root mean square error (RMSE), the Perez-Driesse model achieves the lowest value at 3.866[W∕m2], indicating the closest agreement with observed data. Conversely, the Klucher model reports the highest RMSE at 8.264[W∕m2], reflecting a greater discrepancy between estimated and measured irradiance. The coefficient of determination (𝑅2), which quantifies how well each model explains the observed variability, presents values close to 1 for all models, indicating strong overall agreement. Among them, the Perez and Perez-Driesse models achieve the highest 𝑅2 values and lowest variability, attributed to their consideration of all three diffuse components. In contrast, the isotropic, King, Hay-Davies, Reindl, and Klucher models present higher variability and comparatively lower fit quality. Considering all performance indicators, the Perez-Driesse model yields the most accurate and consistent results and is therefore adopted for irradiance estimation on tilted surfaces. After computing irradiance values for each day using the preprocessed data, the estimates are appended as a new variable to the final dataset, to be used in subsequent analysis and variable selection for modeling the power generation of the PV system, as described in the next section. 4. Selection of input variables In the case study, the MG is primarily supplied by a PV system, which provides most of the required energy. A reduced-order model is applied to predict the power output of the installation, as it captures its nonlinear dynamics effectively. Although all available variables may serve as model inputs, the computational complexity of datadriven approaches increases with the number of inputs. To address this, dimensionality reduction techniques can be employed to project the observations of an 𝑛-dimensional space onto a reduced space, while preserving the essential information contained in all variables. This facilitates the development of fast-processing models suitable for real-time control and optimization applications in the DT framework. Clustering based on correlation coefficient matrices and principal component analysis (PCA) are two commonly used techniques for this purpose. Correlation analysis quantifies the relationships between input variables and the model output, allows the discarding of uncorrelated variables and the grouping of highly correlated ones. PCA, in contrast, transforms the original — typically correlated — variables into a set of uncorrelated orthogonal components, providing an optimal representation in a new coordinate system known as the principal component subspace. Both techniques may be applied sequentially when the number of variables remains high after applying correlation-based selection, using a predefined correlation threshold, as demonstrated in [37]. In the present study, only correlation analysis is applied to reduce the input dimensionality. Specifically, the Pearson correlation coefficient 𝜌 is employed to quantify the linear association between each input variable and the output. The 𝜌 for a pair of variables (𝐱,𝐲) with 𝑛 samples 𝐱=[𝑥1,1,…, 𝑥𝑛,1] and 𝐲=[𝑦1,1,…, 𝑦𝑛,1] is calculated using Eq. (5): 𝜌(𝑥, 𝑦) = 1 𝑛−1 𝑛 ∑ 𝑖=1(𝑥𝑖,1−𝜇𝑥 𝜎𝑥)(𝑦𝑖,2−𝜇𝑦 𝜎𝑦)(5a) 𝜇𝑥=1 𝑛 𝑛 ∑ 𝑖=1 𝑥1,𝑖 (5b) 𝜎𝑥=√ √ √ √1 𝑛−1 𝑛 ∑ 𝑖=1 (𝑥1,𝑖 −𝜇𝑥)2(5c) where 𝜇𝑥, 𝜇𝑦 and 𝜎𝑥, 𝜎𝑦 are the mean and standard deviation of the samples 𝐱 and 𝐲, respectively. The values that 𝜌 can take are between [−1,1], where 𝜌= −1 represents a complete negative correlation, 𝜌=1 represents a complete positive correlation, and 𝜌=0 indicates no correlation between the variables (𝐱,𝐲). Prior to calculating the correlation coefficients, the pre-processed data set (𝑃 𝑉 𝐃∈ℜ𝑛×𝑚) must be normalized to mean zero and unit variance using Eq. (6) to ensure each variable contributes equally to Energy Conversion and Management 342 (2025) 120001 8
W.D. Chicaiza et al. Fig. 7. Matrix of correlation coefficients of the prepared data for the PV facility 𝑃 𝑉 C∈ℜ𝑚×𝑚. the analysis: 𝑧𝑛,𝑚 =𝑥𝑛,𝑚 −𝑥𝑚 𝜎𝑚 (6) where 𝑧𝑛,𝑚 is the normalized variable, 𝑥𝑛,𝑚 is the original variable, 𝑛=182,613 is the number of samples, 𝑚=9 is the number of variables, 𝑥𝑚 and 𝜎𝑚 is the mean and the standard deviation of the variable 𝑚. The correlation coefficient matrix C∈ℜ𝑚×𝑚 is then calculated, where each element C𝑖,𝑗 represents the correlation coefficient between the variables 𝑖 and 𝑗, as shown in Fig. 7. Inspection of the correlation coefficient matrix reveals that all variables exhibit positive correlations with the output. The coefficients are ranked from highest to lowest to assess the influence of each potential input variable on power generation (𝑃𝐷𝐶 ), which serves as the model output. A threshold of 𝜌≥0.9 is used for variable selection. Variables that do not exceed this threshold, such as wind speed (𝑤𝑠), diffuse horizontal irradiance (𝐼ℎ,𝑑 ), and direct normal irradiance (𝐼ℎ,𝑏) are excluded. Although 𝐼ℎ,𝑑 and 𝐼ℎ,𝑏 contribute to the computation of irradiance on tilted surfaces (𝐼𝑇), as described in the previous section, their individual correlation with the output does not justify their inclusion. Ambient temperature (𝑇𝑎𝑚𝑏) is retained based on its physical relevance, despite not exceeding the correlation threshold. Panel temperature, which would be even more informative, is unavailable in the SCADA data. Part of the solar energy absorbed by the PV module is converted into thermal energy, which must be dissipated through heat transfer mechanisms to maintain optimal cell operating temperatures [48]. Higher panel temperatures reduce energy conversion efficiency, resulting in lower energy output. Although current (𝐼𝐷𝐶 ) and voltage (𝑉𝐷𝐶 ) are strongly correlated with the output, they are excluded from the input set, as they represent electrical outputs of the system rather than external inputs [14,48]. The product of these variables directly yields the generated electrical power. Based on this analysis, the variables selected as inputs to the datadriven model presented in Section 5.2 are: global horizontal irradiance (𝐼ℎ), solar radiation on tilted surfaces (𝐼𝑇), and ambient temperature (𝑇𝑎𝑚𝑏). Since panel temperature is not available, it is approximated using (𝑇𝑎𝑚𝑏). The selected variables differ in magnitude and units, which may affect the training process. To mitigate this, normalization is applied to rescale the data into the range [0,1] using Eq. (7). 𝑧𝑛,𝑚 =𝑥𝑛,𝑚 −𝑥𝑚,min 𝑥𝑚,max −𝑥𝑚,min (7) The normalized variables are stored in the matrix 𝑃 𝑉 𝐙 and partitioned into three subsets: training (𝑃 𝑉 𝐓𝐫𝐧), validation (𝑃 𝑉 𝐕𝐚𝐥), and testing (𝑃 𝑉 𝐓𝐞𝐬𝐭), as illustrated in Fig. 8. The learning data account for 80% of the complete dataset, corresponding to 87 days, while the remaining 20% (21 days) are reserved for testing. The learning portion is further divided into a training set, comprising 70% of the learning data (61 days), and a validation set, comprising the remaining 30% (26 days). 5. Digital twin of a PV facility Based on the premise that a digital twin consists of a set of models sufficient to represent each domain in which it is developed and to mirror the behavior of the physical entity. In control and optimization applications, reduced-order or surrogate models are commonly employed to represent highly nonlinear dynamical systems while ensuring real-time computational feasibility. In industrial design contexts, high-fidelity models may be required to address specific performance questions. However, in automatic control and process optimization, simplified models are often sufficient, as high accuracy is not always necessary. Real-time applications, such as MPC, require that the optimization problem be solved within strict time constraints, making the use of complex models impractical in many cases [22]. For this reason, surrogate models are preferred when balancing model fidelity and computational efficiency. In this context, two types of models are considered to form the core of the DT for the PV facility: one based on mathematical equations, and another based on an NF approach, as depicted in Fig. 9. The development of both models is described below. 5.1. Physical model based on the equivalent electrical circuit of the PV panel A PV system is commonly represented by the steady-state model of a PV cell, based on its equivalent electrical circuit [14,48]. This circuit consists of a current source (𝐼𝐿), a diode (𝐷), a series resistor (𝑅𝑠), a load resistor (𝑅𝑙𝑜𝑎𝑑 ), and a shunt resistor (𝑅𝑠ℎ) connected in parallel, as illustrated in Fig. 10. The current–voltage (𝐼–𝑉) characteristic of this circuit1 is defined by Eq. (8a). This physical expression allows the calculation of the output current (𝐼) as a function of the output voltage (𝑉) of the PV system. 𝐼=𝐼𝐿−𝐼𝑜⎡⎢⎢⎢⎣ 𝑒(𝑉+𝐼⋅𝑅𝑠 𝑎)−1⎤⎥⎥⎥⎦ −𝑉+𝐼⋅𝑅𝑠 𝑅𝑠ℎ (8a) 𝑃=𝑉⋅𝐼(8b) Assuming operation at the maximum power point (MPP), the values of the current (𝐼𝑚𝑝) and voltage (𝑉𝑚𝑝) at MPP are substituted into Eq. (8), leading to Eq. (9), where 𝐼𝐿 denotes the solar-induced current, 𝐼𝑜 the diode saturation current, and 𝑎 the diode ideality factor. The electrical power of the cell is then calculated accordingly. 𝐼𝑚𝑝 =𝐼𝐿−𝐼𝑜⎡⎢⎢⎢⎢⎣ 𝑒⎛⎜⎜⎝ 𝑉𝑚𝑝 +𝐼𝑚𝑝𝑅𝑠 𝑎⎞⎟⎟⎠−1⎤⎥⎥⎥⎥⎦ −𝑉𝑚𝑝 +𝐼𝑚𝑝𝑅𝑠 𝑅𝑠ℎ (9a) 𝑃𝑚𝑝 =𝑉𝑚𝑝 ⋅𝐼𝑚𝑝 (9b) Taking the partial derivative of Eq. (9a) with respect to 𝑉𝑚𝑝 and setting it equal to zero yields the optimal values of 𝐼𝑚𝑝 and 𝑉𝑚𝑝, as shown in Eq. (10). 1A detailed explanation of the equivalent circuit of a PV module is provided in [48]. Energy Conversion and Management 342 (2025) 120001 9
W.D. Chicaiza et al. Fig. 17. Measured and predicted power using the NF and EEC models for October 9, 2023. Fig. 18. Measured and predicted power using the NF and EEC models for December 8, 2023. Fig. 19. Measured and predicted power using the NF and EEC models for January 31, 2024. considerably higher in the EEC model, with a difference of approximately 73[W], which represents a fivefold increase compared to the NF model. Although both models present errors in opposite directions, the overestimation bias of the EEC model is more pronounced. Such overestimation could negatively affect applications such as control design, where optimistic predictions may compromise system stability. On the other hand, the conservative nature of the NF model may be advantageous from an operational standpoint. The standard deviation of the error further emphasizes this contrast: 216.50[W] for the EEC model versus 126.22 [W] for the NF model. This 90.28 [W] difference in dispersion, which increases to 207.56[W] when considering the 95% confidence intervals (2𝜎), indicates that the EEC model produces more variable and less consistent predictions. The NF model, by contrast, demonstrates a narrower range of error values, resulting in greater reliability and lower uncertainty in its predictions. Given the large number of samples analyzed, the SE is small for both models: 1.26[W] for EEC and 0.73 [W] for NF. These values indicate that the mean error estimates are statistically robust. However, the lower SE in the NF model reinforces its consistency and the precision of its performance assessment. Energy Conversion and Management 342 (2025) 120001 16
W.D. Chicaiza et al. Fig. 20. Measured and predicted power using the NF and EEC models for March 13, 2024. Table 10 Overall performance indices for the EEC and NF models on the testing set. Testing indices Index Model output EEC NF 𝐸[W] −89.19 16.37 𝜎𝐸[W] 216.50 126.22 2𝜎𝐸[W] 432.99 252.44 SE [W] 1.26 0.73 RMSE [W] 234.14 127.27 𝑅2– 0.9718 0.9895 The RMSE and the 𝑅2 further support the superior performance of the NF model. The EEC model yields an RMSE of 234.14[W] and an 𝑅2 of 0.9718, while the NF model improves upon both metrics with an RMSE of 127.27[W] and an 𝑅2 of 0.9895. This reduction in RMSE by approximately 107[W] reflects a substantial improvement in accuracy. In addition, the NF model explains 1.77% more of the variance in the observed data, reinforcing its suitability for applications requiring high predictive precision. In summary, the statistical indicators used to evaluate both approaches confirm the superior performance of the NF model. Its lower mean error, reduced variance, and smaller RMSE demonstrate a consistent ability to produce accurate predictions. Furthermore, its higher 𝑅2 and lower SE suggest greater stability and robustness. In contrast, the EEC model shows a clear tendency to overestimate, coupled with higher variability and less reliability. These results position the NF model as a more suitable option for accurate and dependable DC power prediction. The performance indices presented in this work can be contrasted with those reported in a recent study [32], referenced in the introduction. That study proposes two prediction approaches for PV power generation using an ANN and two MLR models. For model fitting, the authors use a dataset consisting of 3 days of measurements sampled every 3 s, resulting in 86,400 samples (100%). Of these, 85% is allocated to the learning process and 15% to testing. The learning data is further split into 70% for training and 15% for validation. Additionally, two more days are used exclusively for testing and two additional days for comparing results with other approaches. In total, 4 days from November and December 2022 are used for performance evaluation. The ANN model in that study achieves an 𝑅2 of 0.9159, while the two MLR models achieve 𝑅2 values of 0.8972 and 0.8840. Given that the ANN model provides the best results, and its architecture is illustrated in Fig. 21, it has been selected for comparison with the approaches presented in this work. As shown in Table 10, both the EEC and NF models outperform the ANN approach, achieving 𝑅2 values that explain 6.10% and 8.04% more of the variance in the actual data, respectively. Table 11 Comparison of normalized performance indices and relative improvements over ANN. Metric ANN EEC NF Improvement (%) EEC NF Normalized MAE 0.02856 0.01982 0.0036 +30.61 +87.39 Normalized RMSE 0.07333 0.05200 0.0282 +29.10 +61.55 𝑅20.9159 0.9718 0.9895 +6.10 +8.04 The referenced article also evaluates performance using the mean absolute error (MAE) and RMSE. However, these metrics are not directly comparable due to the difference in system capacities: 160[Wp] in the compared study versus 4500[Wp] in the present work. For this reason, the metrics have been normalized with respect to the nominal capacity of each system. In the referenced work, the ANN model yields a normalized MAE of 0.02856 and a normalized RMSE of 0.07333. By contrast, the normalized error values obtained from the EEC model in this study are 0.01982 (MAE) and 0.0520 (RMSE), while for the NF model they are 0.0036 and 0.0282, respectively. These differences are significant: the EEC model reduces the normalized MAE and RMSE by 30.61% and 29.10%, respectively, compared to the ANN model. The NF model demonstrates even greater improvements, with reductions of 87.39% in MAE and 61.55% in RMSE. These percentages represent the relative improvement of each model over ANN, and highlight the higher accuracy and performance of the NF approach in capturing the dynamic behavior of the PV system. In addition to improved performance, it is worth noting that the NF model presented here uses only three input variables — one fewer than in the ANN approach — and is trained on a more extensive and diverse dataset. These aspects further highlight the robustness and efficiency of the proposed method. To provide a more equitable comparison, Table 11 presents the normalized performance metrics for the three models, along with their relative improvements over the ANN benchmark. Table 11 summarizes the normalized performance indices for the three models and includes the relative improvements of the EEC and NF models over the ANN approach. As shown, the NF model achieves a 87.39% reduction in normalized MAE and a 61.55% reduction in normalized RMSE compared to the ANN, along with an improvement of 8.04 percentage points in 𝑅2. The EEC model also outperforms the ANN, with reductions of 30.61% in MAE and 29.10% in RMSE, and an increase of 6.10 percentage points in 𝑅2. These results confirm that both proposed models offer improved predictive accuracy, with the NF model demonstrating the most significant overall gains in performance and reliability. The evaluation also includes detailed performance indices for specific groups of days across each month in the testing dataset, 𝑃 𝑉 𝐓𝐞𝐬𝐭, as shown in Table 12. Although both the NF and EEC models exhibit Energy Conversion and Management 342 (2025) 120001 17
W.D. Chicaiza et al. Fig. 21. Architecture of the ANN model from the study used as a baseline for comparison [32], consisting of three hidden layers with 6, 4, and 2 neurons, respectively. The input variables are ambient temperature (𝑇𝑎𝑚𝑏), solar irradiance (𝐼ℎ), wind speed (𝑤𝑠), and relative humidity (𝑅𝐻) to predict the power generation (𝑃𝑝𝑣) of photovoltaic module. Table 12 Monthly performance indices obtained in the testing of the EEC and NF models. Testing index Index Model Output JUL 2023 SEP 2023 OCT 2023 1st to 4th 17th to 20th 07th to 09th EEC NF EEC NF EEC NF 𝐸[W] −64.2257 −6.5471 −84.1177 27.2838 −206.8304 3.8690 𝜎𝐸[W] 217.9018 112.4418 217.9975 121.2105 281.7038 116.2790 2𝜎𝐸[W] 435.8036 224.8836 435.9950 242.4210 563.4076 232.558 SE [W] 2.8711 1.4816 2.8724 1.5971 4.2860 1.7691 RMSE [W] 227.1517 112.6225 233.6460 124.2330 349.4531 116.3299 𝑅2−0.9701 0.9920 0.9729 0.9910 0.9788 0.9914 Index DEC 2023 JAN 2024 MAR 2024 8th, 13th to 14th 16th to 18th, 31th 5th to 6th, 13th EEC NF EEC NF EEC NF 𝐸[W] −109.9025 14.0134 −62.9424 33.7137 −28.5575 23.8309 𝜎𝐸[W] 174.3890 93.4867 165.9755 142.1895 182.1221 153.1765 2𝜎𝐸[W] 348.7780 186.9734 331.9510 284.3790 364.2442 306.3530 SE [W] 2.8440 1.5246 2.1869 1.8735 2.7709 2.3305 RMSE [W] 206.1116 94.5189 177.4960 146.1197 184.3266 155.0017 𝑅2−0.9761 0.9913 0.9791 0.9864 0.9819 0.9875 errors in power estimation, the metrics — consistent with the global analysis — highlight that the NF model offers superior accuracy and better captures the non-linear behavior of the system. The first group of days, from July 2023, reveals that both models tend to overestimate power, as indicated by negative values of 𝐸. However, the NF model exhibits a significantly smaller mean error (−6.55 W) compared to the EEC (−64.23 W), suggesting a much lower bias. Similarly, the error variability (𝜎𝐸) is nearly halved in the NF model (112.44[W]) relative to the EEC (217.90 [W]). The 95% confidence interval (2𝜎𝐸) further confirms this, with the NF at 224.88 [W] versus 435.80[W] for the EEC. These values indicate that the NF model produces more consistent and less uncertain predictions. Regarding SE, the NF model achieves a lower value (1.48 [W]) than the EEC (2.87[W]), reinforcing its improved reliability. The RMSE of the NF is also significantly lower at 112.62[W], compared to 227.15[W] for the EEC. Given the PV system’s nominal power of 4.5k[W]peak, this difference is operationally relevant. The NF model’s coefficient of determination (𝑅2=0.9920) surpasses that of the EEC (𝑅2=0.9701), showing its higher explanatory power. In summary, across all metrics, the NF model outperforms the EEC in both predictive accuracy and consistency. This trend is observed throughout the remaining monthly groups (see Table 12). In general, the EEC model consistently overestimates power output, while the NF model slightly underestimates it, except in July. For each month, the absolute values of 𝐸, 𝜎𝐸, and SE are lower for the NF model, indicating reduced error magnitude and uncertainty. The MAE values reflect this clearly: the EEC model shows substantially higher deviations—by 57.68[W] in July, 56.83 [W] in September, 202.96[W] in October, and 95.89 [W] in December 2023. This trend continues in 2024 with differences of 29.23[W] in January and 4.73[W] in March. Likewise, the differences in 𝜎𝐸 between models range from 23.79[W] to over 165 [W], reinforcing the NF model’s superior consistency. RMSE comparisons also show a clear advantage for the NF model: reductions of over 100 [W] are observed in July, September, October, and December. In January and March 2024, RMSE is also significantly lower in the NF model, with improvements of 31.38 [W] and 29.32 [W], respectively. These findings underline the NF model’s efficiency in minimizing prediction error. Finally, the 𝑅2 values confirm the model’s robustness across all months. The NF model consistently explains a greater proportion of output variability — by 2.19 percentage points in July, 1.81 in September, 1.26 in October, 1.52 in December, 0.73 in January, and 0.56 in March — demonstrating its superior ability to capture seasonal and nonlinear system dynamics. In conclusion, the monthly evaluation metrics reinforce the global findings: the NF model provides more accurate, consistent, and robust predictions of PV power generation across varying seasonal conditions, Energy Conversion and Management 342 (2025) 120001 18
W.D. Chicaiza et al. making it better suited for both real-time and long-term forecasting scenarios. 6.2. Statistical significance analysis To statistically assess the observed performance differences between models under varying environmental conditions, a paired t-test is conducted using the error metrics presented in Table 12. Specifically, the analysis focuses on the MAE and RMSE indices of the EEC and NF models. For the MAE comparison, the test results in the rejection of the null hypothesis (ℎ=1), indicating a statistically significant difference between models. The 𝑝-value obtained is 0.0477. This means that, assuming the null hypothesis of no difference is true, there is a 4.77% probability of observing a difference in MAE as extreme as (or more than) the one obtained. The confidence interval for the mean difference ranges from 1.1335 to 147.9726, which does not include zero and therefore reinforces the conclusion of statistical significance. The calculated t-statistic is 𝑡=2.6103, with 5 degrees of freedom. Similarly, for the RMSE comparison, the 𝑝-value is 0.0182, indicating that, under the null hypothesis, the probability of observing such a large difference in RMSE is 1.82%. The corresponding confidence interval ranges from 26.7594 to 183.0271, again excluding zero. The associated t-statistic is 𝑡=3.4510, also with 5 degrees of freedom. These results provide statistical evidence that the NF model significantly outperforms the EEC model in terms of both RMSE and MAE. The confidence intervals further support this conclusion by quantifying the expected range of improvement, highlighting both the statistical and practical relevance of the NF model’s superior predictive performance for power forecasting in PV systems. 6.3. Graphical performance analysis An additional approach used to evaluate the performance of the EEC and NF models in estimating the active power generated by the PV system is the Taylor diagram, shown in Fig. 22. This diagram provides a visual representation of model performance in terms of correlation coefficient, standard deviation, and root mean square deviation (RMSD), comparing predictions with observed values [37]. In the diagram, the standard deviation and RMSD of each model’s predictions are plotted relative to the correlation coefficient, where the actual data are represented by the red point (R), the NF model by the blue point, and the EEC model by the green point. The proximity of a model point to the data point indicates better alignment with the observed behavior. The NF model yields a standard deviation of 1168.1[W], which lies between 1000[W] and 1200[W], while the EEC model shows a higher standard deviation of 1263.2[W], within the range of 1200[W] to 1400[W]. Both values are close to the standard deviation of the actual data (1202.5[W]). Regarding the RMSD and correlation coefficient, the NF model achieves an RMSD of 126.22[W] and a correlation coefficient of 0.9947, whereas the EEC model yields values of 216.50[W] and 0.9858, respectively. In addition, error frequency histograms are provided in Fig. 23 to complement the graphical performance evaluation. The corresponding probability density functions — computed from the error mean and standard deviation presented in Table 10 — are illustrated as green and blue lines for the EEC and NF models, respectively. The histogram for the EEC model approximates a normal distribution, EEC, but is skewed to the left of the zero-error line. This indicates a systematic tendency to overestimate the power output, likely due to the model not accounting for partial shading effects in the facility. In contrast, the NF model, which learns directly from real data, inherently captures such phenomena, offering a key advantage in realistic prediction. Fig. 22. Taylor diagrams for EEC and NF models. Fig. 23. Error frequency histograms for the EEC and NF models. The probability density functions, which reflect the mean and standard deviation of the errors (see Table 10), are shown as green and blue curves for the EEC and NF models, respectively. The histogram for the NF model, associated with NF, shows a slight skew to the right, suggesting a minor underestimation bias. This may result from the model’s sensitivity to lower irradiance levels, where it prioritizes conservative predictions. Nonetheless, the NF model benefits from its data-driven adaptability, providing a closer representation of the PV facility’s actual behavior. Consequently, it can be concluded that the NF model better approximates the observed data, with a standard deviation, RMSD, and correlation coefficient closer to those of the real measurements. This conclusion is supported by the Taylor diagram in Fig. 22, the error histograms in Fig. 23, and the numerical performance metrics in Tables 10 and 12. 6.4. Computational burden Some performance indices related to computational time are summarized in Table 13, reinforcing the superior efficiency of the NF approach. This model requires significantly less execution time compared to the EEC model, whose solution of the equations is based on a solver. Table 13 shows the simulation performance metrics for each model. The results include the average execution time per sample (𝑡𝑠∕sample), per day (𝑡𝑠∕day), and for the full test period of 21 days (𝑡𝑠∕total), computed over 29,680 samples. Energy Conversion and Management 342 (2025) 120001 19
W.D. Chicaiza et al. Table 13 Computational execution time. Avg. execution time Models NF EEC 𝑡𝑠∕sample [μs] 22.176 582.948 𝑡𝑠∕day [ms] 31.934 839.445 𝑡𝑠∕total [s] 0.658 17.302 The NF model exhibits substantially shorter execution times across all metrics. In particular, 𝑡𝑠∕sample is 22.176[μs] for the NF model, over 26 times faster than the 582.948[μs] required by the EEC model. Similarly, 𝑡𝑠∕day for the NF model is 31.934[ms], compared to 839.445 [ms] for the EEC. The complete test simulation for 21 days takes only 0.658s for the NF model, versus 17.302[s] for the EEC. Moreover, the EEC model does not account for partial shading caused by surrounding elements. To do so, a geometric model of shading would be required [48], increasing model complexity and computational cost. These results confirm that the NF model is well suited for computationally demanding applications such as MPC, where prediction speed is critical. MPC is commonly used in energy optimization processes to minimize the cost of electricity purchased from the grid, while indirectly maximizing the use of renewable energy. This is typically achieved by formulating the problem with an economic cost function, as shown in [20,28,67,68]. These strategies are often implemented at the second level of hierarchical control in microgrids. For instance, in an MPC-based energy management scenario with a 24-h prediction horizon and a 1 min sampling period, the NF model requires only 31.934[ms] of computation time per day on a standard PC, making its deployment in real-time applications feasible. In contrast, the EEC model’s reliance on numerical solvers makes it impractical for implementation on embedded systems such as PLCs, which often lack support for complex equation-solving capabilities. The NF model, on the other hand, can be readily embedded and supports rapid what-if analyses and optimization tasks due to its low computational burden and gray-box structure. In this context, the performance predictions provided by the NF model serve as a key component for evaluating and minimizing the associated cost function. Furthermore, in large-scale PV plants, where spatial variability in irradiance requires segmentation into multiple clusters, the computational advantage of the NF model scales proportionally. For 𝑛 panel clusters, the execution time of the NF model becomes 𝑛×31.934 [ms], while that of the EEC model scales as 𝑛×839.445[ms]. This confirms the suitability of the NF model for high-granularity scenarios and real-time control, where increased spatial resolution leads to greater computational demands. 7. Conclusions This work develops two predictive models to serve as the core of a DT for estimating the power output of a PV field located at the CIESOL research center of the University of Almería. These models are designed to simulate different operating scenarios, improve predictive control, and enable optimization through interaction between virtual and physical entities. The first model is based on mathematical equations, commonly referred to as the EEC model. The second is an NF model implemented via an ANFIS. Both models replicate the power generation behavior during operation and non-generation periods. According to the validation metrics, the NF model consistently outperforms the EEC model. The significantly lower values of 𝐸 and RMSE indicate that the NF model more accurately predicts PV power output under different seasonal conditions. These performance differences are crucial for selecting a model depending on the control or optimization strategy. Model training and validation are conducted using measurement data from a SCADA system. Twinning is applied only to the NF model, which adjusts its internal parameters using input–output data during the learning process to replicate the plant’s behavior. In contrast, the EEC model estimates current and voltage at the maximum power point using a system of equations solved numerically. The parameter tuning in the NF model imposes a significantly lower computational burden, making it suitable for implementation in PLCs. The developed models exhibit the following features: 1. Both models show adequate accuracy and precision, as indicated by low values of | 𝐸| and RMSE, and high 𝑅2, when compared to existing ANN-based models in the literature. 2. For a test data set comprising 21 days of operation (29680 samples), the total simulation time is 0.658 [s] for the NF model and 17.302 [s] for the EEC model, with time steps of 22.176 [μs] and 582.948[μs], respectively. The NF model is computationally more efficient and is suitable for implementation in MPC schemes. Its gray-box nature also enhances interpretability. 3. Both models accurately follow the trends in actual power measurements during generation and non-generation periods. The worst-case 𝐸 and 2𝜎 for the EEC model are 89.1885[W] and 432.99[W], respectively—values notably higher than those of the NF model. Moreover, the normalized MAE and RMSE of the EEC model presents a relative improvement of 30.61% and 29.10% over the ANN model from the literature, while those of the NF model show a relative improvement of 87.39% and 61.55%, respectively. 4. The NF model can be re-trained to reflect degradation or changes in the physical properties of the PV panels over time, enabling continuous adaptation. This dynamic update capability ensures sustained alignment between the virtual and physical systems, fulfilling the core twinning requirement of the DT paradigm. 5. While ANN models are often considered black boxes due to their lack of transparency, the ANFIS-based NF model maintains a degree of interpretability. This aspect is particularly important in DT applications involving control and automation, where traceability and understanding of decision-making are critical. 6. The developed models form the computational core of the Digital Twin of the PV facility, enabling accurate behavior replication and real-time synchronization with the physical system. Their integration supports decision-making, optimization, and the implementation of intelligent energy management strategies within the MGDT framework. In summary, this study contributes to the dynamic modeling of PV systems by proposing a fuzzy-neural network approach well suited for integration into a DT framework. Leveraging its gray-box nature, the proposed model offers robust predictive capabilities and efficient computation, facilitating its use in real-time control and optimization tasks. CRediT authorship contribution statement William D. Chicaiza: Writing – review & editing, Writing – original draft, Visualization, Validation, Supervision, Software, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Alex O. Topa: Writing – review & editing, Writing – original draft, Validation, Software, Investigation, Formal analysis. Adolfo J. Sánchez: Writing – review & editing, Writing – original draft, Supervision, Project administration. Juan M. Escaño: Writing – review & editing, Writing – original draft, Supervision, Resources, Project administration, Methodology, Investigation, Funding acquisition, Conceptualization. J.D. Álvarez: Writing – review & editing, Writing – original draft, Supervision, Resources, Project administration, Funding acquisition. Energy Conversion and Management 342 (2025) 120001 20
W.D. Chicaiza et al. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Acknowledgments The authors thanks to the Junta de Andalucía, Spain (Consejería de Transformación Económica, Industria, Conocimiento y Universidades) for the research grant ‘Laboratorio de Ingeniería para la Sostenibilidad Energética y Medioambiental, ENGREEN’, reference QUAL21 006 USE. This work also has been financed by COMMIT4.0EB (ref. PID2021126889OB-I00) funded by MCIN/AEI/10.13039/501100011033 and by ‘‘ERDF A way of making Europe’’ and NTech4Build (ref. TED2021131655B-I00) funded by State Research Agency of Spain (Agencia Estatal de Investigación) AEI/10.13039/501100011033 and by European Union ‘‘NextGenerationEU’’. The authors also acknowledge the support of the INTEREST project, which has received funding from the European Union’s Horizon Europe research and innovation programme under Grant Agreement No. 101160594. Appendix. Adaptive neuro-fuzzy inference system, ANFIS The ANFIS architecture is formed of five layers: fuzzification, product, normalization, defuzzification, and output. Layers 1 and 4 have adaptive parameters (square nodes), adjusted during the learning phase using input and output data to minimize the discrepancy between the target and the actual result. The parameters of these layers are referred to as premise or antecedent parameters and consequent parameters in the first and fourth layers, respectively. Whereas layers 2, 3, and 5 are not trainable (circle node), with parameters that are fixed in the system, as shown in Fig. 13. The fuzzification layer transforms the crisp inputs (𝑥𝑖) into linguistic labels, which is made up of the membership functions (MFs) for each fuzzy set [29] 𝐹𝑖𝑗 ⊂{𝐴𝑖𝑗 , 𝐵𝑖𝑗 }, defined by the membership degree. Gaussian functions are used to define these MFs, and the degree of membership 𝑢𝐹𝑖𝑗 (𝑥𝑖) of 𝑥𝑖 represents the output of each node in this layer. 𝜇𝐹𝑖𝑗 ∶𝑥𝑖∈R⟼𝜇𝐹𝑖𝑗 (𝑥𝑖) ∈ R(A.1a) 𝜇𝐹𝑖𝑗 (𝑥𝑖) = 𝑒 −(𝑥𝑖−𝑐𝑖𝑗 )2 2𝜎2 𝑖𝑗 (A.1b) The second layer features nodes labeled ∏ that implement fuzzy engine inference, which involves the implication of each of the inputs. 𝜔𝑗(𝑥) = 𝜇𝐴𝑖𝑗 (𝑥1)⋅𝜇𝐵𝑖𝑗 (𝑥2)⋅…𝜇𝐹𝑖𝑗 (𝑥𝑖)(A.2) The third layer normalizes the inference engine, which calculates the ratio between the weighting factor of the 𝑗th rule and the sum of all the weighting factors of the rules in the system. 𝜔𝑗(𝑥) = 𝜔𝑗(𝑥) ∑ 𝑗=1𝜔𝑗(𝑥) (A.3) The defuzzification layer provides a weighted output of the Takagi– Sugeno fuzzy rule if-then, obtained from the product of the normalization layer 𝑤𝑗 and the output of the 𝑗th rule (𝑓𝑗). 𝜔𝑗(𝑥)⋅𝑓𝑗(A.4a) 𝑓𝑗=𝑐0𝑗+𝑔1𝑗⋅𝑥1+𝑔2𝑗⋅𝑥2+ … 𝑔𝑖𝑗 ⋅𝑥𝑖(A.4b) The output layer calculates the sum of all weighted outputs of the rules. 𝑓= ∑ 𝑗=1 𝜔𝑗(𝑥)⋅𝑓𝑗(𝑥) = ∑ 𝑗=1𝜔𝑗(𝑥)⋅𝑓𝑗(𝑥) ∑ 𝑗=1𝜔𝑗(𝑥) (A.5) For a comprehensive explanation of the ANFIS architecture network, refer to [36,56]. Data availability Data will be made available on request. References [1] IEA. Net zero emissions by 2050 scenario (NZE) – global energy and climate model – analysis - IEA. Tech Rep 2023. URL https://www.iea.org/reports/globalenergy-and-climate-model/net-zero-emissions-by-2050-scenario-nze. [Accessed 09 February 2024]. [2] Shrivastava A, Sharma R, Kumar Saxena M, Shanmugasundaram V, Lal Rinawa M, Ankit. Solar energy capacity assessment and performance evaluation of a standalone PV system using PVSYST. Mater Today: Proc 2023;80:3385–92. http://dx.doi.org/10.1016/j.matpr.2021.07.258, SI:5 NANO 2021. [3] Yasmeen R, Yao X, Ul Haq Padda I, Shah WUH, Jie W. Exploring the role of solar energy and foreign direct investment for clean environment: Evidence from top 10 solar energy consuming countries. Renew Energy 2022;185:147–58. http://dx.doi.org/10.1016/j.renene.2021.12.048. [4] IEA. Renewables - energy system - IEA. Tech Rep 2023. URL https://www.iea. org/energy-system/renewables. [Accessed 09 February 2024]. [5] Sharif A, Meo MS, Chowdhury MAF, Sohag K. Role of solar energy in reducing ecological footprints: An empirical analysis. J Clean Prod 2021;292:126028. http://dx.doi.org/10.1016/j.jclepro.2021.126028. [6] Siddique M, Shahzad N, Umar S, Waqas A, Shakir S, Kashif Janjua A. Performance assessment of trombe wall and south fac¸ade as applications of building integrated photovoltaic systems. Sustain Energy Technol Assess 2023;57:103141. http://dx.doi.org/10.1016/j.seta.2023.103141. [7] Van Opstal W, Smeets A. Circular economy strategies as enablers for solar PV adoption in organizational market segments. Sustain Prod Consum 2023;35:40–54. http://dx.doi.org/10.1016/j.spc.2022.10.019. [8] Mehmood A, Ren J, Zhang L. Achieving energy sustainability by using solar PV: System modelling and comprehensive techno-economic-environmental analysis. Energy Strat Rev 2023;49:101126. http://dx.doi.org/10.1016/j.esr.2023.101126. [9] IEA. Renewables 2023, IEA, Paris. Tech Rep 2023. URL https://www.iea.org/ reports/renewables-2023. [Accessed 29 February 2024]. [10] Alhuyi Nazari M, Rungamornrat J, Prokop L, Blazek V, Misak S, Al-Bahrani M, Ahmadi MH. An updated review on integration of solar photovoltaic modules and heat pumps towards decarbonization of buildings. Energy Sustain Dev 2023;72:230–42. http://dx.doi.org/10.1016/j.esd.2022.12.018. [11] Batiyah S, Sharma R, Abdelwahed S, Zohrabi N. An MPC-based power management of standalone DC microgrid with energy storage. Int J Electr Power Energy Syst 2020;120:105949. http://dx.doi.org/10.1016/j.ijepes.2020.105949. [12] Chen T, Cao Y, Qing X, Zhang J, Sun Y, Amaratunga GA. Multi-energy microgrid robust energy management with a novel decision-making strategy. Energy 2022;239:121840. http://dx.doi.org/10.1016/j.energy.2021.121840. [13] Topa A, Cruz N, Álvarez J, Torres J. On the optimal demand-side management in microgrids through polygonal composition. Sustain Energy Grids Netw 2023;34:101066. http://dx.doi.org/10.1016/j.segan.2023.101066. [14] Topa A, Gil J, Álvarez J, Torres J. A hybrid-MPC based energy management system with time series constraints for a bioclimatic building. Energy 2024;287:129652. http://dx.doi.org/10.1016/j.energy.2023.129652. [15] Zhang H, Qi Q, Ji W, Tao F. An update method for digital twin multidimension models. Robot Comput-Integr Manuf 2023;80:102481. http://dx.doi. org/10.1016/j.rcim.2022.102481. [16] Guzman Razo DE, Müller B, Madsen H, Wittwer C. A genetic algorithm approach as a self-learning and optimization tool for PV power simulation and digital twinning. Energies 2020;13(24). http://dx.doi.org/10.3390/en13246712. [17] Chicaiza WD, Becerra-Mora YA, Escaño JM, Sánchez AJ, Acosta JÁ. Development of neurofuzzy and Gaussian regression models for a solar photovoltaic system. In: Lesot M-J, Vieira S, Reformat MZ, Carvalho JaP, Batista F, Bouchon-Meunier B, Yager RR, editors. Information processing and management of uncertainty in knowledge-based systems. Cham: Springer Nature Switzerland; 2025, p. 339–50. http://dx.doi.org/10.1007/978-3-031-73997-2_29. [18] Sun H, Liu N, Tan L, Han J. Data-driven energy sharing for multi-microgrids with building prosumers: A hybrid learning approach. IET Renew Power Gener n/a(n/a). http://dx.doi.org/10.1049/rpg2.12821. [19] Rosales-Asensio E, Icaza D, González-Cobos N, Borge-Diez D. Peak load reduction and resilience benefits through optimized dispatch, heating and cooling strategies in buildings with critical microgrids. J Build Eng 2023;68:106096. http://dx.doi. org/10.1016/j.jobe.2023.106096. [20] Gómez J, Chicaiza WD, Escaño JM, Bordons C. A renewable energy optimisation approach with production planning for a real industrial process: An application of genetic algorithms. Renew Energy 2023;215:118933. http://dx.doi.org/10.1016/ j.renene.2023.118933. Energy Conversion and Management 342 (2025) 120001 21
W.D. Chicaiza et al. [21] Wu Y, Guerrero JM, Wu Y, Bazmohammadi N, Vasquez JC, Cabrera AJ, Lu N. Digital twins for microgrids: Opening a new dimension in the power system. IEEE Power Energy Mag 2024;22(1):35–42. http://dx.doi.org/10.1109/ MPE.2023.3324296. [22] Chicaiza WD, Gómez J, Sánchez AJ, Escaño JM. El gemelo digital y su aplicación en la automática. Rev Iberoam Automat E Inform Ind 2024;21(2):91–115. http: //dx.doi.org/10.4995/riai.2024.20175. [23] Castilla M, Redondo J, Martínez A, Álvarez J. Artificial neural networkbased digital twin for a flat plate solar collector field. Eng Appl Artif Intell 2024;133:108387. http://dx.doi.org/10.1016/j.engappai.2024.108387. [24] Semeraro C, Lezoche M, Panetto H, Dassisti M. Digital twin paradigm: A systematic literature review. Comput Ind 2021;130:103469. http://dx.doi.org/ 10.1016/j.compind.2021.103469. [25] Schluse M, Priggemeyer M, Atorf L, Rossmann J. Experimentable digital twins— Streamlining simulation-based systems engineering for industry 4.0. IEEE Trans Ind Inform 2018;14(4):1722–31. http://dx.doi.org/10.1109/TII.2018.2804917. [26] Altair. 2022 digital twin global survey report. 2022, https://altair.com/resource/ digital-twin-summary-report. [Accessed 05 February 2023]. [27] Bazmohammadi N, Madary A, Vasquez JC, Mohammadi HB, Khan B, Wu Y, Guerrero JM. Microgrid digital twins: Concepts, applications, and future trends. IEEE Access 2022;10:2284–302. http://dx.doi.org/10.1109/ACCESS.2021.3138990. [28] Bordons C, Garcia-Torres F, Ridao MA. Model predictive control of microgrids. Cham: Springer International Publishing; 2020, http://dx.doi.org/10.1007/9783-030-24570-2. [29] Machado DO, Chicaiza WD, Escaño JM, Gallego AJ, de Andrade GA, NormeyRico JE, Bordons C, Camacho EF. Digital twin of a fresnel solar collector for solar cooling. Appl Energy 2023;339:120944. http://dx.doi.org/10.1016/j.apenergy. 2023.120944. [30] Gao Y, Chang D, Chen C-H, Sha M. A digital twin-based decision support approach for AGV scheduling. Eng Appl Artif Intell 2024;130:107687. http: //dx.doi.org/10.1016/j.engappai.2023.107687. [31] Kobayashi K, Daniell J, Alam SB. Improved generalization with deep neural operators for engineering systems: Path towards digital twin. Eng Appl Artif Intell 2024;131:107844. http://dx.doi.org/10.1016/j.engappai.2024.107844. [32] Keddouda A, Ihaddadene R, Boukhari A, Atia A, Arıcı M, Lebbihiat N, Ihaddadene N. Solar photovoltaic power prediction using artificial neural network and multiple regression considering ambient and operating conditions. Energy Convers Manage 2023;288:117186. http://dx.doi.org/10.1016/j.enconman.2023. 117186. [33] Amer HN, Dahlan NY, Azmi AM, Latip MFA, Onn MS, Tumian A. Solar power prediction based on artificial neural network guided by feature selection for large-scale solar photovoltaic plant. Energy Rep 2023;9:262–6. http://dx.doi.org/ 10.1016/j.egyr.2023.09.141, The 8th International Conference on Sustainable and Renewable Energy Engineering. [34] Rosiek S, Alonso-Montesinos J, Batlles F. Online 3-h forecasting of the power output from a BIPV system using satellite observations and ANN. Int J Electr Power Energy Syst 2018;99:261–72. http://dx.doi.org/10.1016/j.ijepes.2018.01. 025. [35] Trigo-González M, Cortés-Carmona M, Marzo A, Alonso-Montesinos J, MartínezDurbán M, López G, Portillo C, Batlles FJ. Photovoltaic power electricity generation nowcasting combining sky camera images and learning supervised algorithms in the Southern Spain. Renew Energy 2023;206:251–62. http://dx. doi.org/10.1016/j.renene.2023.01.111. [36] Karaboga D, Kaya E. Adaptive network based fuzzy inference system (ANFIS) training approaches: a comprehensive survey. Artif Intell Rev 2019;52(4):2263–93. http://dx.doi.org/10.1007/s10462-017-9610-2. [37] Machado DO, Chicaiza WD, Escaño JM, Gallego AJ, de Andrade GA, NormeyRico JE, Bordons C, Camacho EF. Digital twin of an absorption chiller for solar cooling. Renew Energy 2023;208:36–51. http://dx.doi.org/10.1016/j. renene.2023.03.048. [38] Costa EA, Rebello CM, Schnitman L, Loureiro JM, Ribeiro AM, Nogueira IB. Adaptive digital twin for pressure swing adsorption systems: Integrating a novel feedback tracking system, online learning and uncertainty assessment for enhanced performance. Eng Appl Artif Intell 2024;127:107364. http://dx.doi. org/10.1016/j.engappai.2023.107364. [39] AL-Qaysi AMM, Bozkurt A, Ates Y. Load forecasting based on genetic algorithm– artificial neural network-adaptive neuro-fuzzy inference systems: A case study in Iraq. Energies 2023;16(6). http://dx.doi.org/10.3390/en16062919. [40] Alkan S, Ates Y. Pilot scheme conceptual analysis of rooftop east–west-oriented solar energy system with optimizer. Energies 2023;16(5). http://dx.doi.org/10. 3390/en16052396. [41] Kolahi M, Esmailifar S, Moradi Sizkouhi A, Aghaei M. Digital-PV: A digital twin-based platform for autonomous aerial monitoring of large-scale photovoltaic power plants. Energy Convers Manage 2024;321:118963. http://dx.doi.org/10. 1016/j.enconman.2024.118963. [42] Olayiwola O, Cali U, Elsden M, Yadav P. Enhanced solar photovoltaic system management and integration: The digital twin concept. Solar 2025;5(1). http: //dx.doi.org/10.3390/solar5010007. [43] CIESOL. Centro de investigaciones en energía solar. 2024, https://www.ciesol.es. [Accessed 15 March 2024]. [44] Gong Y, Liu G, Xue Y, Li R, Meng L. A survey on dataset quality in machine learning. Inf Softw Technol 2023;162:107268. http://dx.doi.org/10.1016/j.infsof. 2023.107268. [45] Nisbet R, Miner G, Yale K. Handbook of statistical analysis and data mining applications (second edition). 2nd ed.. Boston: Academic Press; 2018, http: //dx.doi.org/10.1016/C2012-0-06451-4. [46] Soubdhan T, Ndong J, Ould-Baba H, Do M-T. A robust forecasting framework based on the Kalman filtering approach with a twofold parameter tuning procedure: Application to solar and photovoltaic prediction. Sol Energy 2016;131:246–59. http://dx.doi.org/10.1016/j.solener.2016.02.036. [47] Gallego A, Camacho E. Estimation of effective solar irradiation using an unscented Kalman filter in a parabolic-trough field. Sol Energy 2012;86(12):3512–8. http://dx.doi.org/10.1016/j.solener.2011.11.012, Solar Resources. [48] Beckman W, Blair N, Duffie J. Solar engineering of thermal processes, photovoltaics and wind, fifth edition. John Wiley & Sons, Ltd; 2020, p. 1–905. http://dx.doi.org/10.1002/9781119540328. [49] Qadeer A, Parvez M, Khan O, Jafri HZ, Lal S. Optimization of solar radiation on tilted surface in isotropic and anisotropic atmospheric conditions. Next Res 2024;1(2):100075. http://dx.doi.org/10.1016/j.nexres.2024.100075. [50] Driesse A, Jensen AR, Perez R. A continuous form of the Perez diffuse sky model for forward and reverse transposition. Sol Energy 2024;267:112093. http: //dx.doi.org/10.1016/j.solener.2023.112093. [51] Yang D. Solar radiation on inclined surfaces: Corrections and benchmarks. Sol Energy 2016;136:288–302. http://dx.doi.org/10.1016/j.solener.2016.06.062. [52] Li DH, Lou S. Review of solar irradiance and daylight illuminance modeling and sky classification. Renew Energy 2018;126:445–53. http://dx.doi.org/10.1016/j. renene.2018.03.063. [53] LONGi. Datasheet LR4-72hph-425-455m. 2022, https://solosolar.es/wp-content/ uploads/2022/04/ES_Datasheet_LR4-72HPH.pdf. [Accessed 12 December 2023]. [54] Chepp ED, Krenzinger A. A methodology for prediction and assessment of shading on PV systems. Sol Energy 2021;216:537–50. http://dx.doi.org/10.1016/ j.solener.2021.01.002. [55] Li S-Y, Han J-Y. The impact of shadow covering on the rooftop solar photovoltaic system for evaluating self-sufficiency rate in the concept of nearly zero energy building. Sustain Cities Soc 2022;80:103821. http://dx.doi.org/10.1016/j.scs. 2022.103821. [56] Jang J-S. ANFIS: adaptive-network-based fuzzy inference system. IEEE Trans Syst Man Cybern 1993;23(3):665–85. http://dx.doi.org/10.1109/21.256541. [57] Navarro-Almanza R, Sanchez MA, Castro JR, Mendoza O, Licea G. Interpretable mamdani neuro-fuzzy model through context awareness and linguistic adaptation. Expert Syst Appl 2022;189:116098. http://dx.doi.org/10.1016/j.eswa.2021. 116098. [58] Kerk YW, Jong CH, Chang WL, Tan CJ, Tay KM, Lim CP. Designing monotone Takagi-Sugeno-Kang fuzzy inference systems with new joint sufficient conditions. Fuzzy Sets and Systems 2025;502:109217. http://dx.doi.org/10.1016/j.fss.2024. 109217. [59] CENELEC. Programmable controllers - part 7: fuzzy control programming. ed 1.0. CENELEC; 2000. [60] Escaño JM, Bordons C, Witheephanich K, Gómez-Estern F. Fuzzy model predictive control: Complexity reduction for implementation in industrial systems. Int J Fuzzy Syst 2019;21(7):2008–20. http://dx.doi.org/10.1007/s40815-019-00693-z. [61] Schleich B, Anwer N, Mathieu L, Wartzack S. Shaping the digital twin for design and production engineering. CIRP Ann 2017;66(1):141–4. http://dx.doi.org/10. 1016/j.cirp.2017.04.040. [62] Jones D, Snider C, Nassehi A, Yon J, Hicks B. Characterising the digital twin: A systematic literature review. CIRP J Manuf Sci Technol 2020;29:36–52. http: //dx.doi.org/10.1016/j.cirpj.2020.02.002. [63] Cimino C, Negri E, Fumagalli L. Review of digital twin applications in manufacturing. Comput Ind 2019;113:103130. http://dx.doi.org/10.1016/j.compind. 2019.103130. [64] Rasheed A, San O, Kvamsdal T. Digital twin: Values, challenges and enablers from a modeling perspective. IEEE Access 2020;8:21980–2012. http://dx.doi.org/ 10.1109/ACCESS.2020.2970143. [65] Lu Y, Liu C, Wang KI-K, Huang H, Xu X. Digital twin-driven smart manufacturing: Connotation, reference model, applications and research issues. Robot Comput-Integr Manuf 2020;61:101837. http://dx.doi.org/10.1016/j.rcim.2019. 101837. [66] Allen BD. Digital twins and living models at nasa. 2021, https://ntrs.nasa.gov/ citations/20210023699. [67] Pereira M, Muñoz de la Peña D, Limon D. Robust economic model predictive control of a community micro-grid. Renew Energy 2017;100:3–17. http://dx. doi.org/10.1016/j.renene.2016.04.086, Special Issue: Control and Optimization of Renewable Energy Systems. [68] Alarcón MA, Alarcón RG, González AH, Ferramosca A. Economic model predictive control for energy management of a microgrid connected to the main electrical grid. J Process Control 2022;117:40–51. http://dx.doi.org/10.1016/j. jprocont.2022.07.004. Energy Conversion and Management 342 (2025) 120001 22