Full text
Contents lists available at ScienceDirect Engineering Applications of Artificial Intelligence journal homepage: www.elsevier.com/locate/engappai Research paper Reinforcement learning meets bioprocess control through behavior cloning: Real-world deployment in an industrial photobioreactor Juan D. Gila,∗, Ehecatl Antonio Del Rio Chanona b, José L. Guzmán a, Manuel Berenguel a aCentro Mixto CIESOL, ceia3, Department of Informatics, Universidad de Almería, Ctra. Sacramento s/n, Almería, 04120, Spain bSargent Centre for Process Systems Engineering, Imperial College London, SW7 2AZ, London, UK A R T I C L E I N F O Keywords: Data-driven control Model-free control Artificial intelligence-based control systems Offline reinforcement learning Control of photobioreactor systems A B S T R A C T The complexity of living cells as production units creates major challenges for maintaining optimal bioprocess conditions, especially in open Photobioreactors (PBRs) exposed to fluctuating environments. To address this, we propose a Reinforcement Learning (RL) controller for pH regulation in open PBR systems. This work represents, to the best of our knowledge, the first industrial validation of an RL-based control strategy in a bioprocess. Our method begins with an offline training stage in which the RL agent learns from trajectories generated by a nominal Proportional–Integral–Derivative (PID) controller, without direct interaction with the real system. This is followed by deployment, where the RL agent acts and collects data during the day and is fine-tuned every night, enabling adaptation to evolving process dynamics and stronger rejection of fast, transient disturbances. This continual policy adaptation allows the offline-trained RL to effectively handle the inherent nonlinearities and external disturbances in open PBRs. Simulation studies highlight the advantages of our method: the Integral of Absolute Error (IAE) was reduced by 8% compared to PID control, 6% relative to a classical Model Predictive Controller (MPC), and 5% relative to a standard off-policy RL. Moreover, control effort decreased substantially—by 54% compared to PID, 11% compared to MPC, and 7% compared to the standard off-policy RL. Finally, an 8-day experimental validation under varying environmental conditions confirmed the robustness and reliability of the proposed approach. Overall, this work demonstrates the potential of RL-based methods for bioprocess control and paves the way for their broader application to other nonlinear, disturbance-prone systems. 1. Introduction The importance of effective bioprocess control cannot be overstated. Unlike conventional chemical processes, where the reactor is the primary unit of control, bioprocesses rely on living cells as the actual manufacturing units. These cells are complex, autonomous systems with internal regulatory mechanisms and are heterogeneously distributed within the bioreactor. This introduces significant challenges, as the micro-scale dynamics of individual cells cannot be directly manipulated through macro-scale control variables (Luo et al., 2021). Maintaining stable environmental conditions—such as nutrient levels, pH, temperature, and Dissolved Oxygen (DO)—is essential to support cellular proliferation and productivity, yet these variables are in constant flux requiring the implementation of advanced control systems (Liu et al., 2024). Microalgae-based systems exemplify these challenges. Microalgae are photosynthetic microorganisms capable of thriving under adverse ∗Corresponding author. E-mail addresses: [email protected] (J.D. Gil), [email protected] (E.A. Del Rio Chanona), [email protected] (J.L. Guzmán), [email protected] (M. Berenguel). conditions by converting solar energy and carbon-based compounds such as CO2 into biomass, while releasing oxygen (Tarafdar et al., 2023). Their growth depends on the availability of essential nutrients like carbon, nitrogen, and phosphorus. CO2 is typically supplied via injection, serving both as a carbon source and a pH buffer—making pH control one of the most critical system variables. Nitrogen and phosphorus can be supplied directly or derived from the medium, particularly when wastewater is used. These biological and operational complexities underscore the need for advanced, automated control strategies to ensure process stability and regulatory compliance (Luo et al., 2021; Wang et al., 2022; Guzmán et al., 2025). Focusing on pH, this variable stands out as one of the most critical, as it directly influences the solubility and availability of both CO2 and nutrients, significantly affecting the metabolism of microalgae (Juneja et al., 2013; Nordio et al., 2023). Photosynthesis itself induces constant https://doi.org/10.1016/j.engappai.2025.113326 Received 6 September 2025; Received in revised form 17 November 2025; Accepted 23 November 2025 Engineering Applications of Articial Intelligence 164 (2026) 113326 0952-1976/© 2025 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY-NC-ND license ( http://creativecommons.org/licenses/bync-nd/4.0/ ).
J.D. Gil et al. pH fluctuations, further complicating its control. Typically, on/off control systems are employed (Rodríguez-Torres et al., 2021), which fail to capture the system’s dynamic behavior and external disturbances. Additionally, simple Proportional, Integral, and Derivative (PID) controllers with fixed parameters have been proposed as well (Fernández et al., 2010; Isiramen et al., 2022), yet their performance is often insufficient due to the nonlinearities, disturbances, and time-varying dynamics inherent to the process. The reliance on such control strategies largely stems from the challenges associated with developing accurate models that fully represent the complex process dynamics (Guzmán et al., 2021, 2025). Given the modeling challenges inherent in these types of processes, robust and adaptive control techniques have arisen in the literature. Several studies have addressed this issue from different perspectives. For instance, a robust adaptative feedforward tracking control was presented by Schaum et al. (2017). Feudjio Letchindjio et al. (2021) explored the use of Extremum Seeking Control (ESC) to drive the productivity of a continuous photobiorreactor (PBR) to optimal or suboptimal setpoints. Amaro et al. (2023) introduced an adaptative fuzzy strategy for the nominal model of a generalized Model Predictive Control (MPC) in charge of regulating the pH. In the work by Caparroz et al. (2024), an adaptive model based on a regression tree was presented, capable of predicting pH under varying operational conditions. This tree was subsequently used for the development and implementation of an adaptive control strategy based on a PID controller (Hägglund and Guzmán, 2024). In another study, Caparroz et al. (2023) explored the use of a Model Reference Adaptive Control (MRAC) strategy for pH regulation. This strategy was later hybridized with a PID controller (Caparroz et al., 2025a). Moreover, in Caparroz et al. (2025b), a new relay-based autotuning technique based on PID controllers that applies the relay signal to the setpoint was recently designed and experimentally validated for pH control in semi-industrial open PBR. Nevertheless, such adaptive strategies still face challenges related to their reliance on nominal models, limited generalization capability under highly nonlinear and time-varying dynamics, and the difficulty of accounting for the wide range of disturbances typically encountered in PBRs. Moreover, the only approach that can be regarded as truly model-free (Feudjio Letchindjio et al., 2021) targets steadystate optimization at a much slower timescale than the dynamic pH control problem addressed in this work. In light of the limitations of traditional adaptive approaches, datadriven learning techniques have emerged as a more flexible alternative. One of the first efforts in this direction was presented by Pataro et al. (2023), who developed an MPC strategy enhanced with an oracle function that learns online from data to adjust model uncertainties. This strategy demonstrated good adaptability to dynamic variations induced by different culture media. However, despite its learning component, this approach remains grounded in a nominal control structure and relies on a process model for prediction and optimization, thus requiring prior knowledge of the system dynamics. In contrast, Reinforcement Learning (RL) represents a further step toward full autonomy, as it directly learns the control policy through interaction with the environment or historical data, without requiring an explicit model representation or predefined control law (Petsagkourakis et al., 2020). Additional recent contributions further illustrate the expanding role of RL in process control and systems engineering. Notable examples include multi-setpoint and multi-trajectory tracking in bioprocesses (Espinel-Ríos et al., 2025), hybrid Machine Learning (ML) frameworks for industrial sensing and control (Lawrence et al., 2024), RL-based strategies for semi-batch reactor operation (Sass et al., 2022), physic informed neural network-based RL for gas-lifted well optimization (Faria et al., 2025), hybrid mechanistic-RL control of monoclonal antibody production (Monteiro et al., 2026), RL-based control of batch crystallization processes (Lima et al., 2025), and RL-augmented inexact Benders decomposition for large-scale optimization (LI2, 2026). These works highlight the increasing maturity and versatility of RL methodologies across diverse process systems and reinforce the relevance of advancing RL-based control. Despite its potential, RL still faces important challenges for control purposes (Buşoniu et al., 2018), which are specially relevant in the bioprocess engineering field. A key limitation is the data-intensive nature of training, which typically requires either large volumes of process data or accurate simulation models—reintroducing the modeling bottleneck. Moreover, RL agents often rely on ML models such as Artificial Neural Networks (ANNs) to parameterize both the value function and control policy, making them highly dependent on data quality and quantity. In this context, the distinction between on-policy and off-policy learning becomes critical. While on-policy methods require the agent to collect new data by acting according to its current policy (Sachio et al., 2021), off-policy approaches allow the agent to learn from data previously collected under different policies (Wang et al., 2025). This capability is especially valuable in PBR applications, where online experimentation is costly and risky. Off-policy RL seeks to learn policies that can be deployed in realworld settings without further exploration—an essential advantage when direct interaction with the environment is impractical or risky (Levine et al., 2020). Leveraging this property, several authors have designed model-free controllers for a variety of domains. For example, Deng et al. (2023) trained an offline RL agent with historical data from a steel plant and evaluated its performance in a simulator of the system. Similarly, Schepers et al. (2022) applied offline RL to regulate indoor temperatures in buildings, while Blad et al. (2022) explored a multi-agent offline RL approach for the same task. Although these works demonstrate the potential of offline RL, they rely exclusively on fixed datasets and are validated only in simulation. Consequently, the resulting policies cannot adapt when the system dynamics shift—a limitation that becomes acute in bioprocesses and other time-varying environments. Beyond reusing open-loop operational data, a key direction in offpolicy RL is to exploit experiences from expert systems—nominal controllers designed with established control theory or even human operator actions. A widely used technique for this purpose is Behavior Cloning (BC), a supervised-learning method that trains the policy to imitate optimal behaviors contained in an offline dataset (Levine et al., 2020). Because the data reflects stable and high-performing behaviors, BC directly transfers this domain knowledge to the agent, mitigating a long-standing limitation of RL algorithms (Buşoniu et al., 2018) while avoiding the risks and costs of online exploration. For instance, Park et al. (2025) applied an offline RL method trained on human operator actions to control an industrial dividing wall column. A further improvement of such BC strategies involves performing a fine-tuning stage after the initial offline training, leading to hybrid offline-to-online RL algorithms (Cheng et al., 2024). One of the pioneering studies applying this framework was presented by Hu et al. (2024), which implemented an effective offline-to-online RL paradigm for energy management in a hybrid electric vehicle. Wang et al. (2024) proposed an RL algorithm for active pantograph control in high-speed railways, combining offline meta-policy pretraining, online adaptation, and fine-tuning to ensure real-world deployment. Similarly, Seo et al. (2025) trained an RL agent for pressure control in a crude distillation unit using data from human operator decisions and subsequently improved it through online retraining. Chen and Luo (2025) introduced an offline-to-online RL framework for the industrial process control of penicillin fermentation and simulated moving bed processes. Wang et al. (2025) combined an offline RL policy, trained on closed-loop trajectories generated by an MPC controller, with an online fine-tuning phase that adapts to time-varying dynamics of the system. The objective was to optimize the production of an in-silico semi-batch bioprocess. Such hybrid approaches are particularly attractive in domains like bioprocess engineering, where system dynamics can drift, and purely offline policies quickly become sub-optimal. However, most of these Engineering Applications of Articial Intelligence 164 (2026) 113326 2
J.D. Gil et al. strategies have limited applicability to real-world process control so far, as the majority of studies have been restricted to simulations or relatively simple systems. This highlights a gap in the existing literature that must be addressed in order to advance the application of ML techniques to the control of complex systems. Building on the ideas of hybrid methodologies and motivated by the inherently dynamic and nonlinear behavior of bioprocesses such as microalgae cultivation, this work proposes an off-policy RL control system for pH regulation in microalgae PBRs. Specifically, the control algorithm is implemented as a Deep Deterministic Policy Gradient (DDPG) agent (Castilla et al., 2025; Rajasekhar et al., 2025). Unlike previous approaches, this RL framework operates fully model-free, without relying on a nominal process model or predefined control structure. The agent initially learns from historical data generated by conventional controllers, such as PID regulators, without requiring direct interaction with the real system. Once deployed, it can continue training periodically with new data, allowing adaptation to new operating conditions and system dynamics. The most significant achievement of this work is demonstrating, for the first time to our knowledge, the successful realworld deployment of a RL-based approach for controlling a complex, highly nonlinear, and multi-disturbed system. Moreover, compared to previous studies, the proposed approach stands out as: 1. A Partially Observable Markov Decision Process (POMDP) formulation is developed for a demonstration-scale bioprocess, explicitly accounting for fast disturbances such as solar irradiance fluctuations, air injections, and dilution rate variations. The formulation integrates both process and control-engineering variables, effectively bridging RL with conventional control concepts. This design enables the agent to incorporate an intrinsic feedforward control action while managing actuator constraints and the time-varying dynamics characteristic of real industrial bioprocesses. 2. From a RL perspective, a comprehensive methodology is proposed to develop a hybrid offline–online RL-based control strategy. First, the agent is trained offline using experience collected during the operation of a basic nominal PID-type controller. It is worth noting that PID controllers represent a form of nominal control commonly present in virtually all industrial processes (Hägglund and Guzmán, 2024). Second, an online fine-tuning stage is introduced to adapt to the evolving process dynamics and to improve the rejection of fast, transient disturbances. 3. The proposed approach is experimentally validated in an open PBR at demonstration-scale under industrially relevant conditions, during eight days of operation at the Solar Energy Research Center (CIESOL) facilities located at the Andalusian Institute for Agricultural, Fisheries, Food and Organic Production Research and Training (IFAPA) centre close to the University of Almería (UAL). To the best of our knowledge, this represents the first experimental validation of an RL-based control strategy in an industrial-scale bioprocess—providing a relevant contribution to the field of advanced bioprocess control. The paper is organized as follows: Section 2 presents the main materials used in this study, along with the theoretical background on reinforcement learning. Section 3 details the proposed methodology. Section 4 discusses the main results obtained from applying the proposed approach to the PBR systems. Finally, Section 5 provides the main conclusions. 2. Material and methods 2.1. System overview and control problem 2.1.1. System description The raceway PBR used in this study has a surface area of 80 m2 and is located at the IFAPA center of the Junta de Andalucía, near the UAL Fig. 1. Real open PBR facilities, in Almería, Spain. The two reactors have identical characteristics. The one used in this work is the reactor on the left. Fig. 2. PBR physical scheme, (a) top view, (b) side view. (see Fig. 1). This open PBR is primarily used for microalgae biomass production. It consists of two channels, each measuring 50 m in length, 1 m in width, and 0.3 m in depth (see Fig. 2). Mixing and circulation are achieved via a paddlewheel system with a diameter of 1.2 m and equipped with eight blades. CO2 is injected 1 m deep into a sump measuring 0.65 by 1 m, situated 1.8 m upstream from the paddlewheel. The paddlewheel’s role is to ensure continuous recirculation, flow, and homogenization of the culture medium. Flow velocity is regulated by a frequency inverter controlling the paddlewheel’s rotation speed, maintaining a constant velocity during operation. The system is equipped with comprehensive instrumentation that enables real-time data collection. Various process variables are recorded every second, including pH, DO, culture temperature and level, as well as environmental parameters such as solar irradiance, air temperature, wind speed, and relative humidity. Measurements of pH and DO are taken at two critical points: just after the sump and at the end of the channel, right before the paddlewheel. The latter point poses the greatest control challenge, as it is farther from the CO2 injection area and, therefore, is the main focus of the regulation strategies implemented in the system. 2.1.2. Control problem statement To analyze the control problem in PBR systems, it is essential to understand the physical and biological principles governing pH dynamics within the reactors. In microalgae cultivation using freshwater media supplemented with fertilizers (as in the PBR employed in this study—see Morillas-España et al. (2020) for details), the pH of the culture medium is primarily influenced by CO2 uptake during photosynthesis. During this process, microalgae consume CO2, producing O2 and causing an increase in pH. In contrast, the injected dissolved CO2 reacts to form carbonic acid, which in turn lowers the pH. Additionally, solar irradiance (I) serves as the principal energy source determining the rate of photosynthesis, directly affecting microalgae growth and thus influencing pH dynamics. Other important variables impacting photosynthesis include DO and culture medium temperature (T), Engineering Applications of Articial Intelligence 164 (2026) 113326 3
J.D. Gil et al. Fig. 3. External representation of the PBR system. both contributing to the overall biological activity and consequently affecting pH behavior in the system (Nordio et al., 2024). Nevertheless, accurately capturing pH dynamics requires complex measurements of several biological variables that are either unavailable online or unsuitable from a process control perspective, such as biomass concentration. The lack of these measurements greatly complicates precise modeling and the development of model-based control strategies, thereby justifying the use of techniques such as RL. From the perspective of an external representation of the system, and considering the available measurements in the existing facility, the pH control problem can be represented as shown in Fig. 3. In this representation, the previously described variables are supplemented with two additional problem disturbances, the air injection flow rate (Qair ), which is used to maintain DO at desired levels (disturbing pH), and the dilution flow rate (Qd). The dilution flow is added to the reactor after biomass harvesting or when the culture level drops due to evaporation, and thus does not follow any predefined pattern. These flow rates directly influence mass transfer and concentration balances within the system, affecting key parameters such as pH, thereby playing a crucial role in the overall dynamics and control of the bioprocess. Thus, the overall control problem is to maintain the optimal pH for the microalgae strain by regulating CO2 injection. The controller must reject disturbances affecting pH and compensate for the inherent nonlinear dynamics of the photosynthesis process. 2.2. Reinforcement learning background 2.2.1. Reinforcement learning Most RL algorithms assume that the environment can be modeled as a Markov Decision Process (MDP), formally defined by the quintuple (,, 𝑃 , 𝑅, 𝛾), where is the set of system states, the set of available actions, 𝑃∶××→[0,1] the state transition probability function, 𝑅∶×→R the reward function, and 𝛾∈ [0,1] the discount factor. In this setting, an agent interacts with the environment over discrete time steps: at each step 𝑡, it observes the current state 𝐱𝑡∈⊆R𝑛𝑥, selects an action 𝐮𝑡∈⊆R𝑛𝑢 following a policy, 𝜋, and the environment transitions to a new state 𝐱𝑡+1 accordingly. Fig. 4 illustrates this framework. However, in real-world scenarios—particularly in complex systems such as bioprocesses—full state observability is often unrealistic. Instead, the agent relies on observations 𝐨𝑡∈⊆R𝑛𝑜, which provide partial information about the underlying true state. Although it does not fundamentally alter the general interaction scheme of RL, partial observability necessitates a modified nomenclature, as depicted in Fig. 4. This scenario leads to the formulation of a POMDP. From now on, the POMDP formulation will be maintained in this work, as it is the one that defines the PBR system. To configure the RL agent within this POMDP framework, a commonly adopted approach is the actor-critic architecture, which integrates two components: an actor, responsible for selecting actions 𝐮𝑡 based on observations 𝐨𝑡 via a deterministic policy 𝜋(𝐨𝑡;𝜽); and a critic, which evaluates these actions using an action-value function 𝑄(𝐨𝑡,𝐮𝑡;𝜱). In this formulation, the parameters 𝜽 and 𝜱 denote the weights of the actor and critic networks, respectively, and are updated during training to optimize both the action-selection policy and its evaluation. 2.2.2. Deep deterministic policy gradient Among the most representative algorithms that implement the actorcritic paradigm in continuous action spaces is DDPG (Rajasekhar et al., 2025). DDPG specifically relies on neural networks to parameterize both the deterministic policy and the action-value function, making it well-suited for high-dimensional control tasks. For offline training of an agent following this algorithm, past experiences are stored in a memory structure known as the experience buffer, which consists of recorded observations, actions, and rewards at specific sampling times. During training, random mini-batches of size 𝑀 are sampled from the buffer to update the actor and critic models. Specifically, the critic is trained by minimizing the following loss function: 𝐿(𝜱) = 1 𝑀 𝑀 ∑ 𝑖=1 (𝑦𝑖−𝑄(𝐨𝑖,𝐮𝑖;𝜱))2,(1) 𝑦𝑖=𝑟𝑖+𝛾𝑄𝑇(𝐨𝑖+1, 𝜋𝑇(𝐨𝑖+1;𝜽𝑇); 𝜱𝑇).(2) The target value 𝑦𝑖 is computed by summing the immediate reward 𝑟𝑖 and the discounted future return, estimated using decoupled target networks 𝑄𝑇 and 𝜋𝑇. These networks are slowly updated copies of the main networks, used to stabilize learning and prevent oscillations in value estimates (Lillicrap et al., 2015). The actor, in turn, is updated by maximizing the expected accumulated reward using the following estimated gradient over a mini-batch of size 𝑀: ∇𝜃𝐽≈1 𝑀 𝑀 ∑ 𝑖=1 ∇𝐮𝑖𝑄(𝐨𝑖,𝐮𝑖;Φ)| | |𝐮𝑖=𝜋(𝐨𝑖;𝜽)∇𝜃𝜋(𝐨𝑖;𝜽),(3) where the first part of the expression in the sum denotes the gradient of the critic’s output with respect to the action generated by the actor network, whereas the second corresponds to the gradient of the actor’s output with respect to its own parameters. Note that the updates of the target network parameters are typically performed using a soft update strategy, as follows: 𝜽𝑇←𝜏𝜽+ (1 − 𝜏)𝜽𝑇, 𝜱𝑇←𝜏𝜱+ (1 − 𝜏)𝜱𝑇,(4) where 𝜏 is the smoothing factor. The algorithm table detailing the implementation of DDPG is provided in Algorithm 1. 3. RL-based control strategy for raceway PBR systems This section describes the proposed methodology to apply a DDPG based RL agent to the PBR system. First, the POMDP formulation is explained in detail, including the resulting control scheme representing the interaction of the agent with the real facility. Then, the methodology to train the agent and to transition to the fine-tuning is presented. Engineering Applications of Articial Intelligence 164 (2026) 113326 4
J.D. Gil et al. Fig. 4. Reinforcement learning framework. The left diagram shows the framework for an MDP, and the right one shows the framework for a POMDP. Algorithm 1: DDPG agent for offline training Input: Offline dataset = {(𝐨𝑖,𝐮𝑖, 𝑟𝑖,𝐨𝑖+1)}𝑁 𝑖=1, actor network 𝜋, critic network 𝑄, target networks 𝜋𝑇 and 𝑄𝑇, smoothing factor 𝜏, discount factor 𝛾, mini-batch size 𝑀 Initialize: Randomly initialize critic network 𝑄(𝐨,𝐮;𝜱) and actor 𝜋(𝐨;𝜽) with weights 𝜱 and 𝜽; Initialize target networks 𝑄𝑇 and 𝜋𝑇 with weights 𝜱𝑇←𝜱, 𝜽𝑇←𝜽; for each training iteration do Sample a mini-batch of transitions {(𝐨𝑖,𝐮𝑖, 𝑟𝑖,𝐨𝑖+1)}𝑀 𝑖=1 randomly from ; Compute target Q-values for critic update: 𝑦𝑖=𝑟𝑖+𝛾𝑄𝑇(𝐨𝑖+1, 𝜋𝑇(𝐨𝑖+1;𝜽𝑇); 𝜱𝑇) Update critic by minimizing loss: 𝐿(𝜱) = 1 𝑀∑𝑀 𝑖=1 (𝑦𝑖−𝑄(𝐨𝑖,𝐮𝑖;𝜱))2 Update actor policy using policy gradient: ∇𝜃𝐽≈1 𝑀∑𝑀 𝑖=1 ∇𝑢𝑖𝑄(𝐨𝑖,𝐮𝑖;Φ)| | |𝐮𝑖=𝜋(𝐨𝑖;𝜽)∇𝜃𝜋(𝐨𝑖;𝜽) Soft-update target networks with factor 𝜏 ≪ 1: 𝜽𝑇←𝜏𝜽+ (1 − 𝜏)𝜽𝑇,𝜱𝑇←𝜏𝜱+ (1 − 𝜏)𝜱𝑇, end 3.1. POMDP formulation 3.1.1. Observation space Observation space design in the context of microalgae control presents a significant challenge. As commented, the problem naturally takes the form of a POMDP: many critical variables are not directly observable in real time and must be inferred from other available variables. An effective observation space (⊆R𝑛𝑜) therefore has to fuse indirect measurements, domain knowledge, and control-engineering context so that the RL agent can infer the hidden states and learn policies that remain robust in real operating conditions. The proposed observation space for pH control in PBR systems is organized into three fundamental domains: •Variables measured directly from the process: The first component of the observation vector consists of variables that can be directly measured from the PBR system in real time. According to the schematic shown in Fig. 3, these include the reactor temperature (T), irradiance (I), dissolved oxygen (DO), dilution flow rate (𝑄𝑑), air flow rate (𝑄𝑎𝑖𝑟), and CO2 injection rate. These measurements reflect the physicochemical environment of the culture, the weather conditions, and the current control actions applied to the system. As such, they form the backbone of the observable state and provide the RL agent with essential, realtime information to infer the underlying process dynamics and make informed control decisions. •Temporal information: The second component of the observation vector incorporates temporal information, which is essential due to the periodic nature of PBR operation. Microalgae cultivation is tightly coupled to the day–night cycle, as photosynthesis depends on light availability and follows a predictable rhythm. Thus, including time-related features—such as time of day— enables the RL agent to align its decisions with the photosynthetic cycle. •Control variables: The third component of the observation vector is related to control variables. In addition to including current actuator values, this component incorporates the control error (i.e., the difference between the current pH and its setpoint) as well as the integral of the error over time. These terms are standard in classical control theory, particularly in PID-based strategies, as they capture both the instantaneous deviation and the accumulated offset from the desired value. Including them in the observation space provides the RL agent with a sense of performance history and helps guide the learning of corrective actions that not only stabilize the system but also minimize longterm deviations, making the policy more effective in a control regulation problem context. It is worth noting that some authors in the literature augment the observation space with historical measurements to address the issue of partial observability. However, in this case—given that it is a regulation control problem—this historical information is captured through control-related variables, particularly the integral of the error. This approach simplifies the implementation of the DDPG agent. 3.1.2. Action space The action space (⊆R𝑛𝑢) in this formulation is defined by the CO2 injection rate, as this is the primary actuator available for pH regulation in the PBR system. Injecting CO2 into the medium increases the concentration of dissolved carbon dioxide, which shifts the carbonate equilibrium and leads to a reduction in pH. It should be noted that the action space is constrained between a minimum and maximum CO2 injection rate, i.e., for each sampling time 𝑡, CO𝑚𝑖𝑛 2≤CO2,𝑡 ≤CO𝑚𝑎𝑥 2, reflecting the physical limitations of a real actuator used in the PBR system. This limitation can be directly integrated into the agent’s action space without introducing additional Engineering Applications of Articial Intelligence 164 (2026) 113326 5
J.D. Gil et al. Fig. 5. Comparison of rewards function with respect to the error. The logarithmic function is set with 𝜖= 10−6. Fig. 6. Proposed control structure. complexity. However, it is important to consider that actuator saturation may lead to integral windup in the error integration term included in the observation space. This phenomenon occurs when the integral component continues to accumulate error even though the actuator is out of its limits, potentially causing instability or delayed recovery once the system re-enters the controllable range. To mitigate this issue, an appropriate anti-windup mechanism based in clipping the integral term when is implemented in this study (Caparroz et al., 2025a). 3.1.3. Reward function Since the task is focused on control purposes, the most straightforward approach is to use a reward function (𝑅∶×→R) that directly depends on the error, such as the quadratic error, defined as: 𝑟𝑡= −(𝑒𝑡)2,(5) where 𝑒𝑡=𝑝𝐻SP −𝑝𝐻𝑡 represents the control error at sampling time 𝑡 between the desired set-point 𝑝𝐻SP and the current pH value, 𝑝𝐻𝑡. However, as the agent will be trained with data from a PID controller, where the error is expected to be closed to zero, this type of function can pose problems when calculating the gradients for training the agent in the DDPG algorithm (see Eq. (3)). For this reason, a logarithmic function of the error is proposed instead. This function is given by: 𝑟𝑡= −𝑙𝑜𝑔(𝑒2 𝑡+𝜖),(6) where 𝜖 is a parameter that adjust the maximum of the function (i.e., the maximum value that is reached when the error is zero). Fig. 5 illustrates the differences in reward values between the proposed logarithmic function and the quadratic function across varying error magnitudes. The logarithmic reward function reaches its maximum (i.e., the value achieved with an 𝜖 of 10−6) when the error is near zero. As the error subsequently increases, the function penalizes these deviations but smooths out large errors, a characteristic crucial in contexts like open PBRs. It is important to note that open PBR systems are exposed to environmental conditions, unmeasurable disturbances, and routine operational considerations, such as sensor’s calibration, which can cause large transient errors that, in turn, can destabilize the control loop. Conversely, it can be observed that the quadratic error function maintains values very close to zero with small errors, making it difficult to observe changes in the function’s gradient, and on the other hand, it quadratically penalizes large errors. 3.1.4. Resulting block diagram Taking into account the different components described in the previous subsections, the resulting block diagram representing the interaction between the agent and the real facility is presented in Fig. 6. As depicted in the figure, the agent operates within a classical feedback control scheme, with its behavior guided by the reward function. A key feature, however, is that the observation space—fed by direct measurements of process variables—effectively acts as a feedforward control structure, enabling the rejection of measurable disturbances (Guzmán and Hägglund, 2024). It is also worth noting that a clipping structure has been added to the error integral to prevent windup issues when the system enters actuator saturation. 3.2. RL algorithm training The methodology proposed in this work combines the offline training of the RL agent with subsequent online fine-tuning, and it is fully described in Algorithm 2. The offline component enables the agent to learn an initial control policy from historical data generated by Engineering Applications of Articial Intelligence 164 (2026) 113326 6
J.D. Gil et al. Fig. 7. Actor and critic structure. Algorithm 2: Development of the DDPG agent in an offline environment with real time fine-tuning Input: Historical dataset of the PBR under diverse operating conditions. Output: Robust and optimized control policy for pH regulation in the PBR. Step 1: Preparation of the historical dataset : 1. Collect historical operating data from the PBR using the expert system. 2. Preprocess the collected data and organize it into observation-action-reward tuples, that is, = {(𝐨𝑖,𝐮𝑖, 𝑟𝑖,𝐨𝑖+1)}𝑁 𝑖=1. Step 2: Offline training of the DDPG agent: 1. Randomly initialize the neural network parameters of the DDPG agent. 2. Apply the training process detailed in Algorithm 1. Step 3: Online fine-tuning of the DDPG agent: 1. Initialize the experience buffer using the historical dataset . 2. Continuously collect new online experiences during bioprocess operation at each sampling time 𝑡, that is, ∗ 𝑡= (𝐨𝑡,𝐮𝑡, 𝑟𝑡,𝐨𝑡+1). This will form the data set of new experiences: 𝑁𝑒𝑤 = {(𝐨𝑖,𝐮𝑖, 𝑟𝑖,𝐨𝑖+1)}𝑁𝑁𝑒𝑤 𝑖=1 3. Dynamically update the experience buffer, retaining the most recent experiences. 4. Perform fine-tuning of the DDPG agent using the training process described in Algorithm 11. an expert system, specifically a PID controller. The online fine-tuning phase, on the other hand, serves a dual objective: (i) to adapt the agent’s policy to the changing dynamics inherent in the PBR system, and (ii) to progressively learn from the effects of disturbances on the system. This continuous adaptation aims to significantly enhance the performance beyond that of simple PID-type controllers, from which the agent initially acquired its knowledge. It is important to note that online fine-tuning of a DDPG agent may lead to overfitting or destabilization, particularly due to excessive updates to the actor and critic networks. To mitigate this, the number of training epochs is limited during this phase, thereby constraining the frequency of gradient-based updates. This enables the agent to gradually adapt to changing conditions while maintaining policy stability and overall robustness. Furthermore, during experience replay in the fine-tuning phase (see 2 in Step 3 of Algorithm 2), the replay buffer is initially populated with historical data, which is then progressively replaced by new experiences collected during operation. This approach aims to dynamically update the agent’s policy using recent data that reflects operating conditions specific to the time of year and other dynamic characteristics of the bioprocess itself. 4. Results In this section, the computational implementation of the proposed RL-based control scheme is first described. Then, a simulation study is presented to highlight the benefits of the proposed approach compared to a PID controller, the expert system from which the RL agent was trained, and an RL agent without fine-tuning. Finally, the results obtained from the real facility during eight days of operation are presented. 4.1. Computational implementation With respect to the agent’s architecture, the actor (𝜋) and critic (𝑄) networks were established as represented in Fig. 7. The actor network consists of a recurrent neural network with three fully connected layers. The first two use a ReLU activation function, while the third is followed 1The training for fine-tuning follows Algorithm 1, but with a limited number of iterations or epochs to avoid overfitting issues or destabilization of the agent. Engineering Applications of Articial Intelligence 164 (2026) 113326 7
J.D. Gil et al. Fig. 8. Flow-chart representing the implementation of the proposed methodology. In the online phase, the red area represents the agent–environment interaction, which occurs with 𝑇𝑠= 10 s, while the yellow area represents the fine-tuning process, which is performed once every 24 h. The steps correlate with those reflected in Algorithm 2. by a Tanh activation function to ensure that the output actions remain within a bounded range. Finally, a scaling layer maps the Tanh output to the environment-specific action space, which, in this case, was between 0 and 10 L/min according to the actuator physical limits. On the other hand, the critic network consists of a recurrent neural network as well, and it employs a dual-input architecture designed to estimate the Q-value for a given observation-action pair. It employs two feature input layers: one for the observation and one for the action. These inputs are processed separately through fully connected layers and then merged using a concatenation layer. The merged representation is passed through additional fully connected layers with ReLU activations, culminating in a final fully connected layer that outputs the scalar Qvalue estimation. It should be noted that each layer, both in the actor and critic, contains 256 neurons. Regarding the parameters of the DDPG algorithm, the agent was configured using the Adam optimization method (Kingma and Ba, 2014), with learning rates of 10−4 for the critic and 10−5 for the actor. Moreover, the discount factor (𝛾) was set to 0.9, the smoothing target factor (𝜏) at 0.01, and the sampling time was set to 10 s, in accordance with that of the actual facility. The mini batch size (𝑀) was set at 64. The entire implementation was carried out in MATLAB 2023a (MATLAB, 2023). To deploy Algorithm 2, the scheme shown in Fig. 8 was used. Step 1 in the figure corresponds to obtaining historical data from the expert system, which was performed using a simple PID controller configured in the ideal form without the derivative term, with parameters 𝐾𝑝= −32 [L/min] and 𝑇𝑖= 1200 [s−1] (Caparroz et al., 2025a). These data were then used for offline training over 4000 epochs (Step 2 in Fig. 8). Finally, for online deployment and fine-tuning (Step 3 in Fig. 8), communication with the real facility was established through an industrial protocol, specifically via an OPC DA server. This protocol enabled the retrieval of observations from the real PBR system and the transmission of actions computed by the agent, based on the established sampling time of 𝑇𝑠= 10 s. During operation, data were also stored continuously. At the end of each day (23:59 h), fine-tuning was performed, with training limited to 50 epochs in order to prevent potential robustness issues in the agent. As a result, the fine-tuning procedure effectively operated with a 24-hour update cycle. Finally, it should be mentioned that several well-established operational aspects of the plant were also taken into account in the implementation. First, the pH setpoint was fixed at 8, which is optimal for the type of microalgae strain used. The controller was activated when solar radiation exceeded 100 W/m2 and the actual pH of the system was above the desired reference. It was then deactivated when radiation fell below the defined threshold. Air injections were maintained on demand using a on/off controller integrated into the system’s PLCs, to keep the DO level below a specified limit. This controller was not part of the designed control architecture, as it operated a discontinuous actuator. However, its effect was indirectly captured in the agent’s observation space through the measurements of DO and the recorded air injection events. 4.2. Comparative study in a representative simulation environment To highlight the advantages of the proposed methodology, a comparative simulation study was conducted using four control strategies: 1. A control strategy based on a nominal PID controller (denoted as PID), with the aim of comparing the agent’s performance with that of the expert system used during training. Engineering Applications of Articial Intelligence 164 (2026) 113326 8
J.D. Gil et al. Fig. 9. Operation of the simulated PBR system using the PID controller. (a) pH in the PBR system (pH), pH reference (pHSP), and PBR system temperature (T); (b) Dissolved oxygen (DO) and air flow injections (Qair ); (c) Global irradiance (I), and dilution flow injections (Qd); (d) CO2 flow (CO2). Fig. 10. Values of the reward function during the PID control of the simulated PBR system. 2. A nominal Generalized Predictive Controller (GPC), employing the same nominal model (as its internal model) used for the PID design, with Laplace-domain representation 𝐺(𝑠) = −0.102 1622𝑆+1 𝑒−300𝑠 (Caparroz et al., 2025a). This controller was included to provide an additional comparison with a more advanced control structure. 3. A control strategy using an offline-trained DDPG RL agent without online fine-tuning (denoted as RL). 4. A strategy that incorporated the proposed methodology with daily fine-tuning of the agent’s parameters at the end of each day (referred to as RL-FT). For a fair comparison, this study was carried out using a representative model of the system, validated for the PBR facilities available at the IFAPA research center. The model was fed with meteorological data collected from the actual facility. To increase the difficulty and robustness of the evaluation, the real data used in the study correspond to different time periods: training data were taken from April, while testing data correspond to November 2023. Please note that the used model was described in detail in Nordio et al. (2024) and SánchezZurano et al. (2021) and is available at Rodríguez Miranda et al. (2025). Fig. 9 presents the training data, i.e., the data obtained with the PID-expert system. In this case, two full days of operation were used for the training phase. As can be observed, during these two days there were variations in the main disturbances affecting the system, such as DO and air injection (see Fig. 9-(b)), or solar irradiance and dilution flow rate (see Fig. 9-(c)). To counteract the effects of disturbances and maintain the desired setpoint, the PID controller provided the CO2 injection signal shown in Fig. 9-(d). Moreover, Fig. 10 shows the values of the reward function during operation with the PID controller, which were later used for training the RL agent. Note that the moments when the value was zero correspond to times when the controller was not active, due to the activation conditions mentioned earlier. These instances were not taken into account during the training phase. The results obtained with the four controllers are presented in Fig. 11. Note that the variables acting as disturbances are the same for all four controllers; therefore, only the pH variable and the CO2 injection signal are shown for the four cases. In this test, data from a different time of year were used, which altered the system dynamics with respect to the ones observed in Fig. 10. Additionally, the dilution flow injection and the irradiance profiles (see Fig. 11-(c)) were different. This resulted in the PID controller exhibiting behavior that differed Engineering Applications of Articial Intelligence 164 (2026) 113326 9