scieee AI-readable full text Open interactive document viewer

Reinforcement Learning-Based Control for Robotic Flexible Element Disassembly

Tapia Sal Paz, Benjamín,Sorrosal Yarritu, Gorka,Mancisidor Barinagarrementeria, Aitziber,Calleja Elcoro, Carlos,Cabanes Axpe, Itziar

Abstract

This project was supported by the European Union’s Horizon 2020 research and innovation program under Marie Sklodowska-Curie grant agreement No. 955681 and by members of the Virtual Sensorization Research Group from the University of Basque Country (Basque Goverment Ref. IT1726-22).

Full text

Academic Editor: Yujiong Liu Received: 11 March 2025 Revised: 24 March 2025 Accepted: 27 March 2025 Published: 28 March 2025 Citation: Tapia Sal Paz, B.; Sorrosal, G.; Mancisidor, A.; Calleja, C.; Cabanes, I. Reinforcement Learning-Based Control for Robotic Flexible Element Disassembly. Mathematics 2025,13, 1120. https:// doi.org/10.3390/math13071120 Copyright: © 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/ licenses/by/4.0/). Article Reinforcement Learning-Based Control for Robotic Flexible Element Disassembly Benjamín Tapia Sal Paz 1,2,* , Gorka Sorrosal 1, Aitziber Mancisidor 2, Carlos Calleja 1and Itziar Cabanes 2 1Ikerlan Technology Research Centre, Basque Research and Technology Alliance (BRTA), 20500 Arrasate, Spain; [email protected] (G.S.); [email protected] (C.C.) 2Department of Automatic Control and System Engineering, Bilbao School of Engineering, University of the Basque Country (UPV/EHU), 48013 Bilbao, Spain; aitziber[email protected] (A.M.); itziar[email protected] (I.C.) *Correspondence: [email protected] Abstract: Disassembly plays a vital role in sustainable manufacturing and recycling processes, facilitating the recovery and reuse of valuable components. However, automating disassembly, especially for flexible elements such as cables and rubber seals, poses significant challenges due to their nonlinear behavior and dynamic properties. Traditional control systems struggle to handle these tasks efficiently, requiring adaptable solutions that can operate in unstructured environments that provide online adaptation. This paper presents a reinforcement learning (RL)-based control strategy for the robotic disassembly of flexible elements. The proposed method focuses on low-level control, in which the precise manipulation of the robot is essential to minimize force and avoid damage during extraction. An adaptive reward function is tailored to account for varying material properties, ensuring robust performance across different operational scenarios. The RL-based approach is evaluated in a simulation using soft actor–critic (SAC), deep deterministic policy gradient (DDPG), and proximal policy optimization (PPO) algorithms, benchmarking their effectiveness in dynamic environments. The experimental results indicate the satisfactory performance of the robot under operational conditions, achieving an adequate success rate and force minimization. Notably, there is at least a 20% reduction in force compared to traditional planning methods. The adaptive reward function further enhances the ability of the robotic system to generalize across a range of flexible element disassembly tasks, making it a promising solution for real-world applications. Keywords: intelligent control; robotic control; decision-making; reinforcement learning (RL); robotic disassembly MSC: 68T05 1. Introduction Disassembly is a critical stage in the lifecycle management of products spanning industries such as electronics, automotive, and household appliances. As global industries increasingly prioritize sustainability, disassembly has emerged as a key enabler of repair, recycling, and repurposing initiatives, aligned with the principles of the circular economy [ 1 ]. By recovering valuable materials and components, disassembly reduces waste and supports the reintegration of parts into manufacturing processes. However, automating disassembly remains a formidable challenge due to the inherent complexity, variability, and unpredictability of the tasks involved [2–5]. Key challenges include the following: Mathematics 2025,13, 1120 https://doi.org/10.3390/math13071120 Mathematics 2025,13, 1120 2 of 21 • Product complexity: Disassembly often involves products with numerous, intricately connected components. The complexity rises with the number of parts and the intricacy of their connections, requiring sophisticated handling to avoid damaging valuable elements. • Product variability: Variability across different products, or even between different versions of the same product, necessitates highly adaptable disassembly processes. Traditional automated systems struggle to accommodate this variability without extensive reconfiguration. • Condition of Components: The condition of the components of a product can vary widely. Parts may be damaged, worn out, or contaminated, complicating the disassembly process and requiring adaptable strategies to effectively handle them. Despite these challenges, manual disassembly remains widely used, as human operators excel at managing diverse and unpredictable scenarios. However, manual processes are inherently labor-intensive, time-consuming, and costly, highlighting the growing need for automated solutions. Robots, with their flexibility and advanced capabilities, offer a compelling alternative to traditional automation methods [ 4 , 6 ]. Yet, conventional robotic control techniques, which rely on predefined models and deterministic approaches, often struggle to accommodate the complex interactions between robotic manipulators and components, particularly in unstructured environments [ 4 , 7 , 8 ]. This limitation underscores the necessity of adaptive and intelligent control methods capable of handling the dynamic and unstructured nature of disassembly tasks [9–11]. Reinforcement learning has emerged as a powerful tool for robotic control in complex and dynamic environments, particularly in physical interaction tasks. Unlike traditional methods, RL enables robots to learn optimal policies through trial and error, eliminating the need for precise physical models. This capability is especially beneficial in disassembly tasks involving flexible elements such as cables, seals, and rubber components, which exhibit nonlinear and unpredictable behaviors that are difficult (if not impossible) to model accurately. Recent advancements in RL have demonstrated its effectiveness in solving high-dimensional control problems, making it well suited to physical interaction tasks that demand both precision and adaptability [ 9 , 12 – 18 ]. However, despite these advancements, the application of RL in robotic disassembly (particularly for handling flexible elements) remains underexplored [ 19 – 21 ]. Addressing this gap is crucial, as it presents unique challenges that must be overcome to enable more efficient and autonomous disassembly processes. This study addressed this gap by proposing an RL-based control strategy for the robotic disassembly of flexible elements. The primary objective was to develop a system capable of adapting to the dynamic and nonlinear behaviors of flexible materials while minimizing the interaction forces to prevent damage. This study focused on low-level control, in which the robot interacts directly with unknown flexible elements, making on-the-fly adjustments to ensure safe and efficient extraction. The key contributions of this study are as follows. 1. RL-based control strategy: the design and implementation of an RL-based control strategy tailored to the disassembly of flexible elements, emphasizing force minimization and adaptability. 2. Adaptive reward function: the introduction of an adaptive reward function that normalizes task complexity based on material properties, ensuring consistent performance across varying elasticities. 3. Algorithm comparison: A comparative analysis of state-of-the-art RL algorithms (SAC DDPG, and PPO) to evaluate their effectiveness in dynamic disassembly environments. By benchmarking these algorithms, this work provides practical insights into their Mathematics 2025,13, 1120 3 of 21 applicability for real-world disassembly tasks while also identifying key limitations, such as challenges in generalizing to the unseen direction of extraction scenarios. 4. experimental validation: A comprehensive experimental evaluation in a simulated environment, demonstrating the ability to generalize across different disassembly scenarios and material characteristics. This paper is organized as follows. Section 2provides a comprehensive review of related works, focusing on advancements in robotic disassembly and RL-based control strategies. Section 3presents the problem formulation (Section 3.1) and elaborates on the design of the reward function (Section 3.2). In Section 4, the experimental setup is described in detail, including the implementation specifics (Section 4.1) and experiments conducted (Section 4.2). Section 5analyses the experimental results and provides insights and observations. Finally, Section 6concludes the paper by summarizing the findings and discussing the existing challenges and potential avenues for future research in this domain. 2. Related Work Robotic disassembly has attracted significant attention as industries seek to automate the recovery and recycling of valuable components from end-of-life products. Traditional approaches primarily rely on predefined sequences and deterministic control methods. For example, ref. [ 3 ] explored the use of structured assembly data to guide robotic disassembly, highlighting challenges such as product variability and the need for adaptable systems. However, these methods often fall short in unstructured environments, where product conditions and configurations exhibit high variability [9]. To address these limitations, recent advancements have focused on improving flexibility and adaptability in robotic disassembly. For instance, ref. [ 4 ] proposed a hybrid approach that combines rule-based methods with machine learning to enhance system adaptability. Despite this progress, the disassembly of flexible elements (such as cables and soft materials) remains a major challenge due to their complex and unpredictable interactions [19–21]. Traditional robotic systems, primarily designed for manipulating rigid objects, struggle with the complexities introduced via flexible elements. Various approaches have been explored to overcome this issue, ranging from model-based control to data-driven techniques [ 5 ]. Model-based methods require precise physical models of flexible elements, which are often difficult to obtain and may lack generalizability across different materials and configurations. In contrast, data-driven approaches, particularly those leveraging artificial intelligence (AI), have shown promise in adapting to the variability of flexible elements by training on large datasets. Several studies have focused on specific tasks within flexible object manipulation, such as cable routing and soft object grasping. For instance, ref. [ 22 ] developed a method for cable routing that combines visual feedback with machine learning, enabling robots to adapt to diverse cable types and routing paths. Despite these advancements, the application of such techniques to disassembly remains limited, particularly in scenarios where flexible elements are entangled with rigid components or require precise manipulation to prevent damage. Learning-based techniques provide a promising solution to these challenges by offering the flexibility and adaptability required for complex, dynamic applications. In robotics, two primary methodologies (learning from demonstration (LfD) and reinforcement learning) are commonly employed. While LfD is effective when human demonstrations can guide robotic behavior, it is less suitable for disassembly tasks, which are highly variable and unpredictable. The uniqueness of each disassembly scenario makes it impractical to account for all possible cases through demonstrations alone. Consequently, a more Mathematics 2025,13, 1120 4 of 21 autonomous approach is needed—one that enables robots to navigate and respond to task complexities without extensive human intervention. Several approaches attempt to mitigate these limitations, including generative adversarial imitation learning (GAIL) [ 23 ]. However, GAIL is heavily dependent on the quality and diversity of expert demonstrations. When the test environment deviates from training examples, performance deteriorates, requiring additional training or an expanded dataset. Similarly, inverse reinforcement learning (IRL) derives an expert’s cost function before optimizing policies through reinforcement learning. While effective in some contexts, IRL is computationally expensive, struggles with generalization, and often requires substantial expert data and dedicated hardware [24]. For complex and dynamic tasks such as disassembly, RL presents a compelling alternative to methods reliant on predefined examples. RL allows robotic agents to learn through accumulated experience, rather than explicit demonstrations, enabling them to adapt to the unpredictable nature of disassembly tasks [ 21 , 25 ]. In RL, the agent continuously refines its control policies based on environmental feedback, optimizing performance over time through trial and error [ 26 ]. This paradigm is particularly advantageous in disassembly operations, where precise adjustments in force modulation, compliance, and positioning are essential. By leveraging RL, robotic systems can autonomously manage environmental variability, effectively addressing challenges that traditional control techniques and other learning-based approaches struggle to overcome. Reinforcement learning has proven to be a powerful tool for robotic control in physical interaction tasks, particularly where conventional methods fall short due to scenario complexity and unpredictability [ 21 , 25 ]. By enabling robots to learn control policies through direct interaction with the environment, RL is especially well suited to tasks in unstructured or dynamic settings. The theoretical foundations of RL, established in [ 26 ], have since been extended to a wide range of robotic applications. In robotic disassembly, RL has been applied in two primary areas: high-level sequence planning and low-level control. High-level planning focuses on optimizing the order of disassembly actions, as demonstrated in [ 27 ], where RL was used to determine optimal sequences for electronic device disassembly. This approach prioritizes macro-level decisions, such as minimizing the disassembly time or maximizing material recovery. Conversely, lowlevel control involves the precise manipulation of individual components. RL has shown effectiveness in fine motor control tasks such as grasping and manipulation [ 21 , 28 ]. For instance, ref. [ 29 ] applied RL to the manipulation of soft objects, outperforming traditional control methods in scenarios where object behavior is difficult to model. While significant progress has been made in both robotic physical interaction tasks and reinforcement learning, their intersection remains underexplored. In particular, the application of RL for low-level control in flexible element disassembly has not been comprehensively investigated. Most existing studies focus on related physical interaction tasks, such as assembly, or emphasize high-level planning, with relatively few addressing the unique challenges posed by low-level control in flexible element disassembly [19–21]. Notably, previous research has primarily focused on the disassembly of rigid objects, where problem formulation is simplified by setting the objective force to zero. However, these studies frequently highlight the inability to handle flexible elements as a key limitation. The need for rapid adaptation to varying elastic properties and the complex interactions involved in flexible element disassembly make RL a particularly promising yet underexplored approach. This paper aims to bridge these gaps by proposing an RL-based control strategy specifically designed for flexible element disassembly. Unlike prior studies focused on high-level planning or rigid object manipulation, this research emphasizes low-level control, Mathematics 2025,13, 1120 5 of 21 developing a robust and adaptable system capable of handling the complexities inherent in flexible element disassembly. 3. Problem Formulation The methodology proposed in this work leverages reinforcement learning to develop a robotic control strategy for disassembling flexible elements. This section describes the problem formulation (Section 3.1) and the design of the reward function, including an adaptive reward mechanism for handling varying elasticities (Section 3.2). 3.1. Problem Formulation The disassembly task is formulated as a Markov decision process (MDP), where the robot interacts with its environment to learn an optimal control policy. The MDP is defined by the tuple (O,A,P,R,γ), where the following applies: •Ois the state space, representing the robot’s observations of its environment. •Ais the action space, consisting of the robot’s possible movements. •Pis the transition probability function, describing the dynamics of the environment. •Ris the reward function, providing feedback to the robot based on its actions. •γis the discount factor, balancing immediate and future rewards. The environment in this study is defined by the flexible element to be disassembled and the state of the robot, which includes the position of the end effector and applied forces. The RL agent’s primary objective is to extract the flexible element while minimizing applied forces to prevent damage to both the environment and the robotic system. This problem formulation is intentionally designed for simplicity, ensuring efficient learning and practical implementation. However, it is the outcome of extensive evaluations of more complex representations, which ultimately did not yield significant improvements. The rationale behind these design choices and their implications will be further explored in the discussion section. 3.1.1. State Space (O) The state space, O , represents the robot’s observations of its environment, which include the following: • The position of the end effector relative to the grasping point (eex,eey,eez). • The Cartesian force exerted via the end effector (Fee) , computed as the Euclidean norm of the force components: Fee =qF2 xee +F2 yee +F2 zee . • The distance, d , between the end-effector position (eeposition) and the grasping point (Gposition): d=∥Gposition −eeposition∥. These observations are critical for the robot to monitor its progress, adjust its actions, and ensure safe and efficient disassembly. The state space is formally defined as follows: O −→ eex,eey,eez Fee = Fxee +Fyee +Fzee   d= Gposition −eeposition  (1) Mathematics 2025,13, 1120 6 of 21 3.1.2. Action Space (A) The action space A consists of continuous Cartesian movements of the end effector of the robot (ax , ay , az) while maintaining a fixed orientation. These actions allow for finegrained control over the movements of the robot, enabling precise adjustments during the disassembly process. The action space is defined as follows: A → ax,ay,az∈ ℜ[0, 0.05](2) where the range [ 0,0.05 ] ensures that the movements of the robot are incremental and controlled, minimizing the risk of an excessive force application. 3.2. Reward Function Design The reinforcement learning agent receives feedback from the environment through a reward function, which is fundamental in shaping the learning process of the RL-based controller. The reward function is designed to implicitly encode the task objective by assigning a numerical value to each state–action pair. This value quantifies the immediate benefit or penalty associated with the chosen action by the agent. The goal of the RL agent is to learn a policy that maximizes the cumulative reward over time, thereby optimizing task performance. In the context of flexible element disassembly, the reward function R is formulated to balance two key factors: task progress and force minimization. The objective is to guide the robot toward efficient disassembly while minimizing physical interaction forces to prevent damage to both the flexible element and the robotic system. To achieve this, the proposed reward function is defined as follows: R=α×d−β×Fee2(3) where the following applies: • ( d ) represents the progress made in the disassembly task, measured using the distance between the grasping point and the current position of the end-effector. • ( Fee ) denotes the physical interaction forces exerted via the robot, which should be minimized to prevent damage to the flexible elements and ensure safe handling. • ( α ) and ( β ) are fixed weighting coefficients that govern the trade-off between task progress and force minimization. These coefficients determine the relative importance of each objective in the reward function, ensuring a balanced optimization strategy. The coefficients α = 2.5 and β = 1 were determined through a systematic analysis to produce a reward surface (illustrated in Figure 1) that aligns with the expected physical behavior of the system. This analysis assumed a flexible element with an elastic property of k = 50 [N/m]. The reward function is structured to reflect real-world disassembly scenarios, where successful extraction typically occurs when the end-effector reaches approximately 0.3 m from the grasping point. This formulation ensures that the reward function effectively incentivizes both efficiency and safety in the disassembly process. Figure 1shows the distribution of the reward function within a plane defined by the grasping point and the preferred extraction direction. In the left subfigure, higher reward values are observed in regions aligned with the extraction trajectory, reinforcing the importance of following an optimal path. The right subfigure, which presents a parallel view along the extraction direction, highlights the existence of a peak reward point along the extraction path. This peak corresponds to the location where successful disassembly occurs, demonstrating how the reward function directs the RL agent toward optimal performance. Mathematics 2025,13, 1120 7 of 21 −0.2 −0.4 Grasping Point −0.5 −1.0 −1.5 −2.0 −2.5 −3.0 −0.4 −0.2 −0.4 −0.2 Y axis Extraction Direction Extraction Direction 0.0 0.2 0.4 0.6 0.8 1.0 Figure 1. Reward function value distribution considering a plane that contains the grasping point and the preferred direction of extraction. (Left) shows the reward distribution in a plane that includes the grasping point and the preferred extraction direction. (Right) provides a view parallel to the extraction direction. Adaptive Reward Function The performance of a PPO-based RL agent trained using the fixed reward function defined in Equation (3) is illustrated in Figure 2. The agent was evaluated using four flexible elements with distinct elastic properties ( k = 10, 50, 100, and 200 [N/m]). While the agent successfully extracts the element with the elastic property used during training (k= 50 [N/m]) , its performance deteriorates when tested on elements with different elasticities. Specifically, for the element with k = 10 [N/m], the robot’s end-effector remains near the grasping point, failing to complete the extraction. Conversely, for elements with higher elasticities ( k = 100 and k = 200 [N/m]), the robot applies excessive force even after extraction, increasing the risk of damaging the element. These results highlight the limitations of a fixed reward function in handling materials with varying elastic properties. Similar conclusions were drawn when evaluating the SAC and DDPG algorithms. Distance [m] Distance [m] Figure 2. Performance of the extraction task for flexible elements with different elastic properties using a PPO RL agent trained with the fixed reward function in Equation (3). The agent was trained on an element with k= 50 [N/m]. The plot demonstrates that the agent performs efficiently only for the k= 50 [N/m] element while struggling to adapt to elements with other elastic properties. To address these limitations, an adaptive reward function is introduced. This approach dynamically normalizes the force component of the reward function based on the elastic properties of the element being disassembled in each episode. By incorporating the elastic property of the material, the adaptive reward function ensures that the RL agent can adjust its behavior to the specific characteristics of each flexible element, enabling consistent learning and performance across a wide range of materials. Mathematics 2025,13, 1120 8 of 21 The adaptive reward function is defined as follows: Rnorm =2R − Rmin Rmax − Rmin −1 (4) where the following applies: •Ris the reward computed using Equation (3). •Rmin and Rmax are the minimum and maximum expected reward values for the episode, estimated based on the elastic constant of the flexible element. This normalization scales the reward function within a fixed range ( − 1, 1), ensuring that the reward signal remains independent of the material’s elasticity. Consequently, the RL agent receives appropriately scaled feedback, regardless of the elastic properties of the element. Additionally, the weighting coefficients α and β , which control the trade-off between task progress and force minimization, are dynamically adjusted based on the elastic properties of the flexible element. Specifically, α is defined as α=k2 1.5 , where k represents the elastic constant of the element, while β remains fixed at β= 1. This dynamic adaptation ensures that the reward function scales appropriately with the elastic property of the material, preserving a consistent reward distribution, as illustrated in Figure 1. By incorporating an adaptive reward function, the system maintains robust performance even when faced with significant variations in material properties, which is a common challenge in real-world disassembly tasks. This enhancement enables the RL agent to generalize more effectively across materials with different elasticities, addressing the shortcomings of the fixed reward function. As a result, task performance improves while reducing the risk of damage to both the flexible elements and the robotic system, making the proposed approach more suitable for practical applications. 4. Methodology This section outlines the experimental setup, procedure, and evaluation metrics used to assess the performance of the proposed RL-based control strategy for the disassembly of flexible elements. The experiments were conducted in a simulated environment designed to replicate real-world conditions and challenges. A key aspect of the methodology, consistent with the problem formulation, is the emphasis on maintaining simplicity in both environment design and the information required for task execution. This decision was motivated by the primary objective of this study: to evaluate whether the proposed approach can effectively handle task uncertainties and adapt to real-world conditions. By reducing environmental and state-space complexity, the study aims to demonstrate that robust and adaptive control strategies can be developed even with limited prior knowledge or highly simplified models. This approach enhances the practical applicability of the solution while also providing insights into the ability of the system to generalize across diverse and unpredictable scenarios. 4.1. Experimental Setup For the implementation of the proposed control, this work selects as a use case the disassembly of sealing elements in refrigerators (Figure 3). This is a representative task where the disassembly of these elements requires dynamic actions and the adaptation of the system according to the current state of the flexible element. Mathematics 2025,13, 1120 9 of 21 Gripper Prefered Extraction Direction Uncertain Extraction Direction Flexible Element Grasping Point Interaction Forces (b) a) (a) Flexible Element (c) º Figure 3. Disassembly task use case: Extraction of the sealing element of a fridge door. (a) realworld application; (b) Experimental setup used for simulation, where the gripper replicates the attached points in the real application. (c) Simulation environment used to replicate the real world interaction forces. 4.1.1. Simulated Environment This work uses a simulated environment to validate the proposed methodology in the disassembly task. For that, the whole system is implemented using the ROS 2 (Robot Operating System) framework, simulated in the Gazebo using the collaborative robot KUKA LBR iiwa14 (KUKA, Augsburg, Germany), using a computer with a processor Intel i7 (Intel, Atlanta, GA, USA) and an Nvidia RTX 4080 graphic card (GIGABYTE, Singapore). The KUKA LBR iiwa14 robot was chosen for its advanced kinematic and dynamic capabilities, which are essential for executing the precise and adaptive movements required in flexible element disassembly tasks. Additionally, the KUKA LBR iiwa14 is a collaborative robot designed to operate safely alongside humans. This feature not only enhances its adaptability but also facilitates future integration into human–robot workspaces, making it a versatile choice for applications requiring close collaboration. The simulation replicated the real-world disassembly scenario, as shown in Figure 3, where the setup focused on replicating the physical and dynamic conditions of the flexible element extraction. The main aspects considered in the simulations are as follows: • Kinematics and dynamics: the simulation includes the kinematic and dynamic models of the KUKA LBR iiwa14 robot, ensuring realistic interaction with the flexible elements. • Use case workspace: the workspace mimics the real-world setup, including the constraints and preferred extraction direction for the flexible element. • Interaction forces: The forces exerted during extraction are simulated using two main components; the reaction force of the gripper ( FGripper ) and the flexible element’s elastic force ( Felastic ). These forces are modeled to replicate the physical interactions between the robot and the flexible element during disassembly. However, it is important to note that the main sim-to-real gaps are expected in this aspect, as real-world conditions may introduce additional complexities, such as unmodeled friction, material imperfections, or dynamic perturbations, which are not fully captured in the simulation. Here, the elastic force is modeled as follows: Felastic =Kelastic ×d(5) where Kelastic is the elastic constant of the flexible element, and d is the distance of the end-effector from the grasping point (d=0→Felastic =0). And the gripper reaction force (FGripper) is modeled as follows: Mathematics 2025,13, 1120 16 of 21 Table 2. Evaluation of learned strategies under the combination of different environment configurations: operation range (O), structured configuration (S), and unexplored configuration (U). Evaluation of Learned Strategies Under Different Environment Configurations. Algorithm Training Force (k) Training Direction TestForce (k) Test Direction Mean Reward Success Rate SAC S S S S 0.85 1.00 SAC O O S S 0.60 1.00 SAC O O O O 0.61 1.00 SAC O O U O 0.48 0.00 SAC O O O U −0.08 0.00 SAC O O U U −0.05 0.00 DDPG S S S S 0.75 1.00 DDPG O O S S 0.44 1.00 DDPG O O O O 0.44 1.00 DDPG O O U O 0.48 0.57 DDPG O O O U −0.25 0.00 DDPG O O U U −0.02 0.00 PPO S S S S 0.80 1.00 PPO O O S S 0.62 1.00 PPO O O O O 0.62 1.00 PPO O O U O 0.62 1.00 PPO O O O U −0.46 0.00 PPO O O U U −0.47 0.00 5.2.2. Adaptability and Generalization An in-depth analysis was conducted using various combinations of environmental configurations (Table 2). The study aimed to evaluate how effectively these agents could manage different conditions, both within their training range and beyond their training scenarios. The results in Table 2and Figure 9show that all agents were able to successfully learn the task when trained and tested within their operational range ( O ). As illustrated in Figure 9, the agents displayed the correct reward patterns during episodes, consistently achieving a success rate of 1.00 across all structured environment configurations. This indicates that the agents, particularly when dealing with familiar conditions, were able to execute the task with complete accuracy and reliability. Figure 9. Reward curves for PPO, SAC, and DDPG algorithms trained in the operation range ( O ) and tested in structured ( S ) and operational ( O ) configurations. The curves show the reward progression during testing across a variety of elastic properties and extraction directions within the training range. The three algorithms show similar behaviors in the O range. The shaded regions represent the variance across multiple (100) test runs. Mathematics 2025,13, 1120 17 of 21 However, when tested in previously unexplored environments ( U ), there was a noticeable drop in both success rate and reward values, as shown in Figure 10. This performance drop was particularly evident when the agents encountered new extraction directions that were absent from the training data. In these scenarios, all three agents consistently failed to complete the task, as reflected in the negative rewards (approaching − 1, the lowest possible value due to normalization) and a zero success rate. On the other hand, when faced with unseen elastic properties of the flexible elements, the agents were able to perform the task, but with a reduced success rate and lower rewards compared to familiar scenarios. This suggests that the agents were better equipped to handle variations in material characteristics than drastic changes in task dynamics, such as extraction direction. Figure 10. Reward curves for PPO, SAC, and DDPG algorithms trained in the operational range ( O ) and tested in unexplored environment configurations ( U ). The curves illustrate the agents performance when tested on elastic properties and extraction directions outside the training range. All algorithms show similar performances facing an unseen k but significant performance degradation in scenarios with unfamiliar extraction directions. The shaded regions represent the variance across multiple (100) test runs. 5.3. Discussion The experimental results confirm the effectiveness of the proposed RL-based control strategy for flexible element disassembly. All three algorithms (PPO, SAC, and DDPG) maintained high success rates while minimizing the exerted force, an essential factor in preserving the integrity of flexible components. This aspect is shown in the force signature of flexible element extraction of Figure 8, where a 20% force reduction is achieved against forcible classical methodologies. These findings highlight RL-based control as a promising solution for real-world robotic disassembly. The comprehensive evaluation presented in Table 2provides several key insights: • As expected, the agent performs optimally when trained and tested in structured conditions, achieving the highest success rates. • In operational conditions, the agent also demonstrates strong performance, achieving a perfect success rate. This is a crucial finding, as it validates the proposed approach and supports its potential transfer to real-world experiments. • A significant observation is that the only cases where the agent fails to complete the task involve unknown extraction directions. However, when faced with unknown elastic properties, the agent successfully adapts, demonstrating its ability to generalize across different material conditions. • The adaptive reward function played a crucial role in this success, particularly in handling varying elastic properties of flexible elements. The results indicate that dynamically normalizing the force component of the reward function based on material elasticity allows the RL agent to generalize effectively across diverse scenarios. By Mathematics 2025,13, 1120 18 of 21 scaling the reward function according to the elasticity ( k ) of the element, the system ensures consistent and meaningful feedback, regardless of material properties. • These results were achieved despite using simplified force models and state representations, aligning with the objective of validating the use of simplifications without compromising task success. Furthermore, since the agent successfully overcame these simplifications through the adaptive reward function, this finding suggests that the system may also be capable of bridging the sim-to-real gap and handling real-world uncertainties such as noise and unmodeled environmental factors. Future research will focus on real-world testing, specifically on the disassembly of refrigerator door seals, to validate the robustness of the approach beyond simulation. Among the evaluated reinforcement learning algorithms, PPO emerged as the most practical for real-world implementation, offering a balance between training efficiency, stability, and computational demands. PPO exhibited faster and more stable convergence compared to SAC and DDPG, making it well suited for real-time disassembly tasks. While SAC and DDPG also performed well, their higher computational requirements may limit scalability in practical applications. Despite promising results, certain limitations were identified. The performance of the RL agent deteriorated in extreme cases where the expected extraction direction varied significantly. However, it is important to emphasize that these limitations were observed only in highly extreme scenarios. Within the operational range, which encompasses a broad spectrum of realistic configurations, the agent performed effectively. Addressing these challenges will be essential for developing more robust and generalizable RL agents capable of handling a wider range of disassembly tasks. Future investigations will focus on enhancing the adaptability to highly dynamic environments, ultimately paving the way for RL-driven automation across diverse industrial applications. 6. Conclusions This paper has presented a reinforcement learning-based control strategy for the robotic disassembly of flexible elements, addressing key challenges inherent to dynamic and unstructured environments. The approach centered on low-level control, enabling the robotic system to learn adaptive strategies for extracting flexible components, such as cables and rubber seals, through low-force trajectories. By utilizing RL algorithms, including SAC, PPO, and DDPG, the study demonstrated significant advancements in adaptability, efficiency, and overall task performance compared to traditional control methods. The experimental results showcased the effectiveness of the RL-based control approach in achieving high success rates, consistently minimizing force exertion, particularly in scenarios where predefined thresholds are not possible, and maintaining efficient optimized trajectories across a range of task conditions. A key contributor to this success was the adaptive reward function, which allowed the RL agents to maintain reliable performance when dealing with elements of varying elastic properties. However, the study also identified some limitations, particularly in the ability of the RL agent to generalize beyond their trained task configurations, especially in highly complex or unforeseen situations. These limitations point toward future research opportunities, including exploring hybrid control strategies that combine RL with model-based techniques and expanding the RL framework to multi-agent systems, which could further enhance adaptability and efficiency in complex disassembly scenarios. Mathematics 2025,13, 1120 19 of 21 Overall, this research advances intelligent robotic disassembly by demonstrating the potential of RL-based control to handle the complexities of flexible element disassembly in unpredictable and dynamic environments, which is an existing gap in current robotic disassembly studies [ 19 – 21 ]. The findings underscore the feasibility of automating flexible element disassembly using RL-based controllers, contributing to more sustainable and efficient manufacturing and recycling practices. Future work will aim to further refine these control strategies and expand their applicability across a broader range of disassembly tasks and real-world scenarios. Author Contributions: Conceptualization, B.T.S.P., G.S. and A.M.; methodology, B.T.S.P., G.S. and A.M.; software, B.T.S.P.; validation, B.T.S.P., G.S. and A.M.; formal analysis, B.T.S.P., G.S. and A.M.; investigation, B.T.S.P.; resources, B.T.S.P., G.S., A.M., C.C. and I.C.; data curation, B.T.S.P.; writing—original draft preparation, B.T.S.P.; writing—review and editing, B.T.S.P., G.S., A.M., C.C. and I.C.; visualization, B.T.S.P.; supervision, G.S., C.C. and A.M.; project administration, G.S., C.C. and I.C.; funding acquisition, G.S., C.C. and I.C. All authors have read and agreed to the published version of the manuscript. Funding: This project was supported by the European Union’s Horizon 2020 research and innovation program under Marie Sklodowska-Curie grant agreement No. 955681 and by members of the Virtual Sensorization Research Group from the University of Basque Country (Basque Goverment Ref. IT1726-22). Data Availability Statement: The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author. Conflicts of Interest: The authors declare no conflicts of interest. Abbreviations The following abbreviations are used in this manuscript: RL Reinforcement Learning SAC Soft Actor–Critic DDPG Deep Deterministic Policy Gradient PPO Proximal Policy Optimization AI Artificial Intelligence LfD Learning from Demonstration IRL Inverse Reinforcement Learning MDP Markov Decision Process ROS2 Robot Operating System kModulus of Elasticity References 1. Li, J.; Barwood, M.; Rahimifard, S. Robotic disassembly for increased recovery of strategically important materials from electrical vehicles. Robot. Comput.-Integr. Manuf. 2018,50, 203–212. [CrossRef] 2. Foo, G.; Kara, S.; Pagnucco, M. Challenges of robotic disassembly in practice. Procedia CIRP 2022,105, 513–518. [CrossRef] 3. Vongbunyong, S.; Kara, S.; Pagnucco, M. Application of cognitive robotics in disassembly of products. CIRP Ann.-Manuf. Technol. 2013,62, 31–34. [CrossRef] 4. Poschmann, H.; Brüggemann, H.; Goldmann, D. Disassembly 4.0: A Review on Using Robotics in Disassembly Tasks as a Way of Automation. Chem. Ing.-Tech. 2020,92, 341–359. [CrossRef] 5. Li, F.; Jiang, Q.; Zhang, S.; Wei, M.; Song, R. Robot skill acquisition in assembly process using deep reinforcement learning. Neurocomputing 2019,345, 92–102. [CrossRef] 6. Hjorth, S.; Chrysostomou, D. Human–robot collaboration in industrial environments: A literature review on non-destructive disassembly. Robot. Comput.-Integr. Manuf. 2022,73, 102208. [CrossRef] 7. Wan, A.; Xu, J.; Chen, H.; Zhang, S.; Chen, K. Optimal Path Planning and Control of Assembly Robots for Hard-Measuring Easy-Deformation Assemblies. IEEE/ASME Trans. Mechatron. 2017,22, 1600–1609. [CrossRef] Mathematics 2025,13, 1120 20 of 21 8. Schneider, D.; Schomer, E.; Wolpert, N. A motion planning algorithm for the invalid initial state disassembly problem. In Proceedings of the MMAR: 2015 20th International Conference on Methods and Models in Automation and Robotics, Miedzyzdroje, Poland, 24–27 August 2015; Institute of Electrical and Electronics Engineers: Miedzyzdroje, Poland, 2015; p. 839. 9. Elguea-Aguinaco, Í.; Serrano-Muñoz, A.; Chrysostomou, D.; Inziarte-Hidalgo, I.; Bøgh, S.; Arana-Arexolaleiba, N. A review on reinforcement learning for contact-rich robotic manipulation tasks. Robot. Comput.-Integr. Manuf. 2023,81, 102517. [CrossRef] 10. Duan, J.; Gan, Y.; Chen, M.; Dai, X. Adaptive variable impedance control for dynamic contact force tracking in uncertain environment. Robot. Auton. Syst. 2018,102, 54–65. [CrossRef] 11. Wang, W.; Guo, Q.; Yang, Z.; Jiang, Y.; Xu, J. A state-of-the-art review on robotic milling of complex parts with high efficiency and precision. Robot. Comput.-Integr. Manuf. 2023,79, 102436. [CrossRef] 12. Martín-Martín, R.; Lee, M.A.; Gardner, R.; Savarese, S.; Bohg, J.; Garg, A. Variable Impedance Control in End-Effector Space: An Action Space for Reinforcement Learning in Contact-Rich Tasks. In Proceedings of the 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China, 3–8 November 2019. 13. Schoettler, G.; Nair, A.; Luo, J.; Bahl, S.; Ojea, J.A.; Solowjow, E.; Levine, S. Deep Reinforcement Learning for Industrial Insertion Tasks with Visual Inputs and Natural Rewards. arXiv 2019, arXiv:1906.05841. [CrossRef] 14. Zhang, H.; Wang, W.; Zhang, S.; Zhang, Y.; Zhou, J.; Wang, Z.; Huang, B.; Huang, R. A novel method based on deep reinforcement learning for machining process route planning. Robot. Comput.-Integr. Manuf. 2024,86, 102688. [CrossRef] 15. Englert, P.; Toussaint, M. Learning manipulation skills from a single demonstration. Int. J. Robot. Res. 2018,37, 137–154. [CrossRef] 16. Levine, S.; Wagener, N.; Abbeel, P. Learning Contact-Rich Manipulation Skills with Guided Policy Search. arXiv 2015, arXiv:1501.05611. 17. Chebotar, Y.; Kalakrishnan, M.; Yahya, A.; Li, A.; Schaal, S.; Levine, S. Path Integral Guided Policy Search. In Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA), Singapore, 29 May–3 June 2018. 18. Huang, Y.; Liu, D.; Liu, Z.; Wang, K.; Wang, Q.; Tan, J. A novel robotic grasping method for moving objects based on multi-agent deep reinforcement learning. Robot. Comput.-Integr. Manuf. 2024,86, 102644. [CrossRef] 19. Qu, M.; Wang, Y.; Pham, D.T. Robotic Disassembly Task Training and Skill Transfer Using Reinforcement Learning. IEEE Trans. Ind. Inform. 2023,19, 10934–10943. . [CrossRef] 20. Qu, M.; Pham, D.T.; Altumi, F.; Gbadebo, A.; Hartono, N.; Jiang, K.; Kerin, M.; Lan, F.; Micheli, M.; Xu, S.; et al. Robotic Disassembly Platform for Disassembly of a Plug-In Hybrid Electric Vehicle Battery: A Case Study. Automation 2024,5, 50–67. [CrossRef] 21. Serrano-Muñoz, A.; Arana-Arexolaleiba, N.; Chrysostomou, D.; Bøgh, S. Learning and generalising object extraction skill for contact-rich disassembly tasks: An introductory study. Int. J. Adv. Manuf. Technol. 2023,124, 3171–3183. [CrossRef] 22. Zhang, X.; Sun, L.; Kuang, Z.; Tomizuka, M. Learning Variable Impedance Control via Inverse Reinforcement Learning for Force-Related Tasks. IEEE Robot. Autom. Lett. 2021,6, 2225–2232. [CrossRef] 23. Ho, J.; Ermon, S. Generative Adversarial Imitation Learning. In Advances in Neural Information Processing Systems; Springer: Berlin/Heidelberg, Germany, 2016. 24. Zhao, T.Z.; Kumar, V.; Levine, S.; Finn, C. Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv 2023, arXiv:2304.13705. 25. Beltran-Hernandez, C.C.; Petit, D.; Ramirez-Alpizar, I.G.; Nishi, T.; Kikuchi, S.; Matsubara, T.; Harada, K. Learning Force Control for Contact-Rich Manipulation Tasks with Rigid Position-Controlled Robots. IEEE Robot. Autom. Lett. 2020,5, 5709–5716. [CrossRef] 26. Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction, 2nd ed.; The MIT Press: Cambridge, MA, USA, 2018; p. 526. 27. Chen, H.; Liu, Y. Robotic assembly automation using robust compliant control. Robot. Comput.-Integr. Manuf. 2013,29, 293–300. [CrossRef] 28. Kristensen, C.B.; Sørensen, F.A.; Nielsen, H.B.; Andersen, M.S.; Bendtsen, S.P.; Bøgh, S. Towards a Robot Simulation Framework for E-Waste Disassembly Using Reinforcement Learning; Elsevier: Amsterdam, The Netherlands, 2019; Volume 38, pp. 225–232. [CrossRef] 29. Kroemer, O.; Niekum, S.; Konidaris, G. A Review of Robot Learning for Manipulation: Challenges, Representations, and Algorithms. J. Mach. Learn. Res. 2021,22, 1–82. 30. Tapia Sal Paz, B.; Sorrosal, G.; Mancisidor, A. Hybrid Robotic Control for Flexible Element Disassembly. In Proceedings of the European Robotics Forum 2024, Rimini, Italy, 13–15 March 2014; Secchi, C., Marconi, L., Eds.; Springer: Cham, Switzerland, 2024; pp. 180–185. 31. Lillicrap, T.P.; Hunt, J.J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; Wierstra, D. Continuous control with deep reinforcement learning. arXiv 2015, arXiv:1509.02971. Mathematics 2025,13, 1120 21 of 21 32. Duan, Y.; Chen, X.; Houthooft, R.; Schulman, J.; Abbeel, P. Benchmarking Deep Reinforcement Learning for Continuous Control. In Proceedings of the International Conference on Machine Learning, New York, NY, USA, 20–22 June 2016. 33. Akiba, T.; Sano, S.; Yanase, T.; Ohta, T.; Koyama, M. Optuna: A Next-generation Hyperparameter Optimization Framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Anchorage, AK, USA, 4–8 August 2019. Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.