scieee AI-readable full text Open interactive document viewer

Emerging locomotion behavior by maximizing entropy-based information metrics

ZEREIK, ENRICA; Zanetti, Luca; Bonsignorio, Fabio

Abstract

The Snakebot, proposed by I.Tanev about 20 years ago, is an inspiring illustration of emergent behavior in a robotic application. It demonstrates how the behaviors of a snake-like robotic platform self-organize when certain information metrics are maximized. The integration of self-organization principles, specifically emergence, with morphological computation can facilitate the development of relatively straightforward yet robust controllers for complex structures, enhancing their flexibility and adaptability to environmental changes and dynamic conditions. Our study focuses on achieving self-organization of the behaviors of the robots by optimizing relevant information metrics and guiding evolution using an appropriate learning methodology. In our work we have studied by extensive simulation a WormBot, modeled after the Snakebot, and we employ deep reinforcement learning techniques to enhance the sensory-motor coordination and locomotion behaviors of the WormBot. We developed an appropriate learning system capable of acquiring several motion models for the robot, by optimizing information metrics, and producing an embodied agent that can move consistently (in this case, in a worm-like manner) in accordance with the robot’s morphology. Our results show that the WormBot may acquire coherent locomotionpatterns without explicit motion function constraints, relying solely on a metric associated with predictive information and the interactions among the loosely coupled components of the agent (in our case the sections constituting the WormBot).

Full text

Emerging locomotion behavior by maximizing entropy-based information metrics Enrica Zereik1, Luca Zanetti1and Fabio Bonsignorio2 Abstract— The Snakebot, proposed by I.Tanev about 20 years ago, is an inspiring illustration of emergent behavior in a robotic application. It demonstrates how the behaviors of a snake-like robotic platform self-organize when certain information metrics are maximized. The integration of self-organization principles, specifically emergence, with morphological computation can facilitate the development of relatively straightforward yet robust controllers for complex structures, enhancing their flexibility and adaptability to environmental changes and dynamic conditions. Our study focuses on achieving self-organization of the behaviors of the robots by optimizing relevant information metrics and guiding evolution using an appropriate learning methodology. In our work we have studied by extensive simulation a WormBot, modeled after the Snakebot, and we employ deep reinforcement learning techniques to enhance the sensory-motor coordination and locomotion behaviors of the WormBot. We developed an appropriate learning system capable of acquiring several motion models for the robot, by optimizing information metrics, and producing an embodied agent that can move consistently (in this case, in a worm-like manner) in accordance with the robot’s morphology. Our results show that the WormBot may acquire coherent locomotion patterns without explicit motion function constraints, relying solely on a metric associated with predictive information and the interactions among the loosely coupled components of the agent (in our case the sections constituting the WormBot). I. INTRODUCTION The resilience and flexibility of intelligent robotic solutions that are currently being developed are still insufficient, despite the dramatic rate of innovation in robotics and artificial intelligence. Their comparison with the naturally evolved intelligent physical agents that we can observe in nature (such as animals and plants) still shows enormous gaps in terms of adaptivity, dexterity and flexibility. We believe that to get robotic platforms that can truly function well in complex open-ended real-world applications, we need a paradigm change from the top-down ‘Cartesian’ (assuming a neat division in between mind and body) mechatronic paradigm to a deeply bioinspired paradigm leveraging on morphological computation and emergence. Although learning techniques are helpful, real-time learning from data-rich time series, interaction with an unstructured environment, and collaboration with humans—who then learn alongside the intelligent service robots—needs novel algorithms and systems. The main assumption is that while “computing” is a feature of all material systems, intelligence and meaning 1Enrica Zereik and Luca Zanetti are with the Institute of Marine Engineering of the Italian National Research Council, 16149, Genoa, Italy [email protected], [email protected] 2Fabio Bonsignorio is with University of Zagreb Faculty of Electrical Engineering and Computing, AIFORS LAB, 10000, Zagreb, Croatia and Heron Robots, 16121, Genoa, Italy. [email protected] are created by synchronization processes among networks of independent agents that are “situated” in their surroundings. Investigating the primary unresolved issues in the field of evolutionary emergent information processing in multibody (rigid, compliant, or hybrid) systems is the goal of our research project. Robotics has recently gained many capabilities for being effective and supportive to people within various contexts and applications, even if confined in structured environments and controlled conditions. As soon as these robots have to face a more unpredicted and dynamically-changing environment, their performance is more likely hindered in terms of adaptability, and their (centralized) control should become much more complex to cope with these unstructured situations. In nature, complex collective behaviors (and consequently adaptability) often emerge from much simpler decentralized procedures and subtask execution. Emergent behavior is much desirable in swarms of robots (or robots made by “swarms of modules/agents”) in which many peer subsystems interact, each executing its own tasks with its own algorithms, concurring to the final requested cooperation and producing a sort of collective intelligence. One of the possible ways to steer the emergence of this collective intelligence consists in the maximization of some metrics of information about the robot, the environment and the mutual interaction between them (how robot actions affect the environment which is in turn perceived back by the robot itself). The Snakebot [1] is an interesting example of emergent behavior in a robotic application. It shows a self-organization of the behaviors of a snake-like robotic platform emerging from the maximization of some information metrics. To this aim, Lie Groups were exploited to reduce the computational burden of information-driven self-organization [2], expanding upon previous work [3], [4], [5]. The combination of concepts from self-organization (i.e. emergence) of intelligent behaviors and morphological computation can lead to the design of relatively ‘simple’ robust controllers for complex structures, more flexible and able to adapt to the environment and cope with dynamic changes. In our research, we are working to implement this robotic selforganization by driving the evolution through an appropriate learning approach, grounded on the maximization of suitable information metrics [6], [7]. With respect to Tanev and Prokopenko’s work [8], we consider a WormBot, inspired by their Snakebot, and with an experimental set up similar to [9]) we implement deep reinforcement learning methods to improve the sensorimotor coordination of the WormBot and its locomotion behaviors. The main goal of our research is the understanding of how complex behaviours emerge from natural embodied agents, and the replication of these mechanisms on their artificial robotic counterparts. We believe that our approach could help to disclose the relation between the dynamics of an embodied agent and its information processing capabilities, addressing the challenges highlighted in [10]. In section II we describe our experimental setup and methodology, in III we show our simulation results, finally discussing them in section IV, and drawing some conclusions and outlining our future work. II. SYSTEM SETUP AND METHODOLOGY Designing locomotion controllers for complex robots is challenging, especially when accurate models are not available [9]. A particularly interesting metrics is provided by the predictive information, Ip[11]. The Predictive information, expressed in equation 1, quantifies how much knowledge of the past of a time series helps us predict its future. By definition, it is the mutual information between the random variables’ “past” and “future” in the series. Intuitively, this means we measure how much observing past values reduces our uncertainty about future values. If a series has clear patterns or memory, then past observations carry useful clues about what comes next, so the predictive information is large. If the series is completely random (no correlations), then knowing the past gives no advantage, and the predictive information is zero. Locomotion is inherently sequential and dynamic actions taken now (like joint angles or muscle activations) influence what happens in the next few moments (e.g., body orientation, foot placement, balance). Predictive information measures how much of that future behavior is already contained in the past and current observations. This is useful, as efficient locomotion often appears to rely on rhythmic, repeatable movement patterns (such as walking gaits). Predictive information may help identify components of the system’s state that exhibit predictable structure, potentially including cyclic motor patterns. Ip=Ixt+1 t−τ+1;xt t−τ=1 TPT τ=1 H(xt+1−τ) + −H(xt+1−τ|xt−τ) (1) In our work we addressed a simple embodiment, dealing with a worm-like robot, made of a number of different physical modules. Our primary goal was to understand if and to what extent an information metrics could aid the development of suitable locomotion gaits, given a series of situated cooperating agents (the modules composing the WormBot). We assume that the WormBot moves on a flat terrain with steady dynamical friction. Our aim is to evaluate how information metrics can shape robot learning, not to replicate real-world physics with high fidelity. Accordingly, we use a standard dynamic friction model: contacts oppose tangential motion with a default sliding coefficient (µ≈ 1.0), constrained by a pyramidal friction cone and solved via the Newton method. Slip may occur when tangential forces exceed this limit. This simplified model ensures stable, consistent simulations while keeping computational demands low, aligning with our focus on metric analysis rather than physical accuracy. We first implemented an architecture similar to that proposed in Tanev’s Snakebot [1] and first of all tried to reproduce and validate his work, to update it to the new technologies and computing capabilities, and to make the approach more effective and efficient. We have developed a different approach that gave improved and interesting results. Both are described in the following subsection IIA. A. WormBot - System Overview Fig. 1: Block scheme of the WormBot architecture. The original Tanev’s architecture relied on predefined primitive functions, such as sine and cosine, which inherently constrained the system to oscillatory or periodic behaviors. In this work, we remove this constraint by adopting a deep reinforcement learning framework that learns locomotion policies for the entire structure directly from interaction with the environment. This approach operates without prior knowledge of effective functions and avoids low-level motion generators such as central pattern generators (CPGs) [12]. Instead, behavior emerges solely from the maximization of previously defined information-theoretic metrics. Within this framework, the updated WormBot simulation shifts the focus from isolated locomotion patterns to the execution of a meaningful, goal-directed task. Specifically, the robot must perform target reaching, navigating toward and making contact with a box placed at a fixed distance from its starting position. The key motivation is to investigate whether the inclusion of the predictive information into the reward function can naturally encourage smoother and more efficient motion. Since maximizing predictive information—defined as the mutual information between past and future sensorimotor states—directly increases the predictability of future joint configurations from the past ones, we hypothesize that it will lead WormBot to exhibit more coordinated and fluid movements, ultimately improving its ability to accomplish the target-reaching task. The WormBot is a modular, wormlike robot consisting of 8spherical segments, connected by simulated universal joints, implemented using pairs of perpendicular hinge joints and small auxiliary bodies. We modeled the WormBot by a comparatively small number of spherical segments. This configuration results in a total of 16 controllable joints. The joint positions directly define the agent’s action space, with the policy output specifying target positions for each joint at every control step. Each joint is then controlled in position by a low-level PD controller.The system is implemented and simulated using the MuJoCo physics engine [13], integrated within the Gymnasium framework [14], with policy training performed using the Proximal Policy Optimization (PPO) algorithm [15], as provided by the Stable-Baselines3 (SB3) library [16]. We started with a number of spherical segments equal to that of the original Tanev’s Snakebot, we then refined the optimal number of segments by extensive simulations. Similar optimizations could be performed on the diameter of the spherical segments and the other kinematical and dynamical characteristics of the agents, and of course different body and segment morphologies and in particular the compliance of different body sections will in general affect the optimal number of segments for a given hybrid rigid soft agent. The aim of this article is to identify a broadly applicable approach and to validate it on an initial toy problem, that allows to understand its weaknesses and strenghts at a very fundamental level. To determine an appropriate morphological configuration for the WormBot, we empirically evaluated the number of spheres in the body by testing the target-reaching task on a flat surface. The configurations tested ranged from 5 to 16 spheres. The simulations conducted to determine the optimal number of spheres employed a dense reward function defined as rt=1 1+d2, where d2denotes the Euclidean distance from the head of the robot to the target. Although this formulation was later replaced—since it does not encourage learning relative to progress made between consecutive timesteps—it remains a useful metric for benchmarking performance in this preliminary evaluation. We observed, as shown in Figure 2, that when the number of spheres exceeds 10, the simulator exhibits numerical stability issues, which degrade performance. Considering both stability constraints and task success rates, we selected an 8-sphere configuration, as it not only avoided stability problems but also achieved the highest performance in the target-reaching trials. In addition, we were also seeking a balance between locomotion performance and the complexity of the system’s informational demands. The 8-sphere setup appeared to offer an optimal trade-off—maintaining sufficient morphological complexity to enable effective control and adaptability, while still maintaining a sufficient locomotion performance. B. Learning System Description and Reward Shaping The observation space of the proposed system is a 47dimensional vector containing the Cartesian position and quaternion orientation of the reference sphere, the joint angles of the 14 actuated joints, the linear and angular velocities of the head, the angular velocities of all joints, and the relative Euclidean distance from the head to the target. Fig. 2: Target-reaching performance of WormBot with 5–16 spheres using a dense reward 1 1+d2Configurations over 10 spheres showed stability issues. The 8-sphere setup achieved the best trade-off, balancing morphological complexity and locomotion capability. Including the Euclidean distance in the observation space enables the policy to generalize its behavior across varying target positions, rather than being overfitted to a fixed goal location. The policy is trained using the Proximal Policy Optimization (PPO) algorithm with a multilayer perceptron architecture. The network maps observations through two fully connected hidden layers of 64 units each to produce a 14-dimensional action vector, where each element specifies the desired target position of a corresponding joint. A complete list of simulation and training hyperparameters is provided in Table III. Each training episode has a fixed length of 8192 time steps. During each episode, as shown in equation 2, the reward function combines two key objectives: •Locomotion reward rloc, defined as the change in Euclidean distance between the head of the snake and the target, measured before and after a single step of the reinforcement learning environment. •Information reward ri, i.e. the predictive information between the past and current joint angle values. In the reward function, the combination of the two distinct components is performed via a product rather than a sum, in order to prevent one term from dominating or being neglected during optimization. This multiplicative approach to reward shaping has been shown to be effective in previous work by Montufar et al. [9]. Additionally, the reward is transformed using a square function to suppress low values and emphasize higher ones. To properly handle negative values in the locomotion reward, the squared term is scaled by the sign of the locomotion component, ensuring consistent gradient direction. No such adjustment is required for the predictive information term, as it is intrinsically non-negative. rt=sign(rloc,t)·(|rξ loc,t|r1−ξ i,t )2(2) Table II briefly summarizes the main variables and definitions used in the article. To evaluate the robustness of our approach, we tested the proposed environment on four distinct ground surface types: flat ground, low roughness, medium roughness, and high roughness, as illustrated in Figure 3. The rough surfaces were implemented in MuJoCo using a heightfield generated from a random grayscale PNG image. In this representation, each pixel’s intensity value encodes a height level, with darker pixels corresponding to lower elevations and lighter pixels to higher elevations. By using a randomly generated pattern, the resulting terrain introduces irregularities that are not biased toward any specific shape or frequency, enabling a more general assessment of the system’s adaptability to unpredictable surface variations. The predictive information was calculated using JAX, and simulations were conducted within SB3-JAX, a JAX-based implementation of Stable Baselines3. This setup ensured both tasks ran in a consistent execution environment and benefited from JAX’s compilation and differentiation features for improved computational efficiency. III. SIMULATION RESULTS (a) Flat surface (b) Low-roughness terrain (c) Medium-roughness terrain (d) High-roughness terrain Fig. 3: Types of terrain used in the simulation environment. During experimentation, we encountered significant challenges in steering the evolutionary process toward synchronized motion between the vertical and horizontal axes of the joints. Specifically, the evolved solutions often lacked coherent coordination across these axes, resulting in erratic or inefficient locomotion patterns. Due to these limitations and the inability to achieve satisfactory results with the evolutionary approach under the tested conditions, we have opted not to include detailed results from these experiments in this paper. For the experimental evaluation, each training run was conducted for a total of 8.192 million simulation steps. The simulation uses a timestep of 0.001 seconds, while control actions are applied every 0.01 seconds. To investigate the influence of the predictive information term on the emergence of behavior, we performed a parameter sweep over the weighting factor ξ, which scales the contribution of predictive information in the reward function w.r.t. the locomotion speed, see subsection II-B above. Specifically, we evaluated four values: ξ= [0.25,0.5,0.75,1.0] with purely locomotion reward-driven learning when ξ= 1.0.At each control step, predictive information is computed based on the previous 50 control steps, according to equation 1 (τ= 50). This evaluation allows us to systematically assess how different levels of information drive affect the locomotion capabilities of the worm-like robot, and more broadly, to investigate the role of predictive information in shaping coordinated, self-organized movement in the absence of predefined motion primitives. . Screenshots of the simulations are depicted in Figure 4, which shows the emergence of a suitable locomotion behaviour for the WormBot, whose current positions are shown at time instants t= [0,12,25,37,50,62,75,87,100] s. The relevant role of the predictive information term in the reward function is confirmed also by the computation of the permutation entropy, which indicates the predictability within the data time series: this entropy decreases when the time series is more predictable given the past values. Figure 6 shows the mean of the all permutation entropies computed on all the WormBot joints. It is clear from the plot that there is a minimum in this entropy when a greater weight is assigned to the information-driven terms in the learning system reward function (ξ= 0.25), while the maximum value is reached in correspondence of purely locomotion reward-driven learning (ξ= 1). Task completion performance was evaluated by measuring the final distance to the target at the end of each episode. As shown in Figure 7d, the system progressively improves its ability to approach the target as training advances. In rough environments, the target is often not reached—not due to limitations in control capability, but because episode durations are too short relative to the robot’s maximum velocity. This was confirmed in test runs with extended time duration, where the trained policy was evaluated over 20 runs for each terrain type. In these tests, the target was successfully reached, and the corresponding timesteps are reported in Table I. Screenshots of these tests can be found in Fig 4. Across all environments, performance was consistently worse when relying solely on the locomotion reward, highlighting that the inclusion of predictive information can significantly support learning in locomotion tasks. In order TABLE I: Timestep at which the target is reached in extended test runs Environment Timestep Reached Flat terrain 8,354 ±523 Low roughness 9,781 ±697 Medium roughness 11,434 ±855 High roughness 13,954 ±1,324 to assess the hypothesized smoothing effect of predictive information on the system’s trajectories, we computed the autocorrelation [17] score for each condition. Specifically, we measured the autocorrelation of the time series data (e.g., (a) t= 0 s(b) t= 12 s(c) t= 25 s(d) Navigating in high roughness terrain (e) t= 37 s(f) t= 50 s(g) t= 62 s(h) WormBot close to the target (i) t= 75 s(j) t= 87 s(k) t= 100 s(l) WormBot at target Fig. 4: (a)-(c), (e)-(g), (i)-(k) Screenshots of the simulations for ξ= 0.75 at time instants t= [0,12,25,37,50,62,75,87,100] s. (d), (h), (l) Screenshots of the test runs with extended time. (a) Mean episode reward during training across different values of ξ, averaged over all terrain types. (b) Mean episode reward across terrains of varying difficulty, averaged over all ξvalues. Fig. 5: Incorporating predictive information with a weighting factor ξ= 0.25 consistently increases the reward compared to locomotion-only rewards. Results are smoothed using a moving time average to enhance trend visibility. Performance is reported as the mean over 5 runs per configuration. joint angles’ time series data) across 50 time lags, i.e. those used to compute the predictive information. Autocorrelation captures how well the values at one time point can be related to earlier values, and can therefore be a preliminary proxy Fig. 6: Mean of the permutation entropy in all WormBot joints computed versus the parameter ξalong the simulation. for smoothness or regularity in the signal. Figure 8 shows the autocorrelation scores for 3 adjacent joints, varied as a function of the parameter ξ. These examples were selected because they capture the overall temporal trend observed across all joints—every other joint follows a qualitatively similar pattern—so showing additional plots would be redundant. These results show a general trend: as ξincreases and predictive information decreases, the average autocorrelation scores across lags tends to drop. This indicates that the resulting behaviors become less temporally structured—i.e., less smooth or predictable over time. While this does not directly establish a causal relationship, it suggests that promoting predictive information during learning may be linked to the emergence of smoother, more temporally coherent motion patterns. IV. DISCUSSION AND CONCLUSIONS Our aim in this work was to investigate the emergence of a sensory motor coordination behaviors in an agent made of a loosely coupled set of simpler embodied agents, a kind of little ’swarm’, and to explore the role of entropy-based information metrics in the emergence of suitable locomotion behaviors, specifically applied to a simple worm-like robot. The architecture that we have studied confirmed that the WormBot is able to learn coherent locomotion patterns without any specific forcing of motion functions, by only introducing a metrics related to predictive information, and thanks to the interaction of the agents (e.g. the worm nodes) among them. We developed a suitable learning system able to learn different motion models for the robot, maximizing the information metrics, and obtaining an individual able to move in a consistent way (in this case worm-wise) with the robot morphology. This should provide examples of intelligent behaviors through morphological computation and emergence from loosely coupled networks of embodied agents. In the future, we are planning to perform more simulation experiments with the WormBot architecture, in order to have a deeper insight on the system behaviour, and to obtain (a) Flat terrain (b) Low-roughness terrain (c) Medium-roughness terrain (d) High-roughness terrain Fig. 7: The figure shows that the robot learns to approach the target over time. In rough environments, it fails to reach it due to short episode lengths relative to its maximum velocity. In all environments, performance declines when only the locomotion reward is used. Fig. 8: Autocorrelation of the joint angles’ time series data across 50 time lags. The joints shown in the picture qh2, qv2, qv3 correspond to the second horizontal joint and to the second and third vertical joints. The autocorrelation is calculated for 50 time lags corresponding to the time instant used to calculate the predictive information (τ= 50). The first score is omitted as it autocorrelates the timeseries with itself, so it provides no useful information. The results are very robust to training and exhibit limited variances. more statistically significant results to support the findings of our research activity, aiming to investigate the emergence of intelligent behaviors in networks of cooperating agents with distributed control. Further, we will focus on improving and optimizing the execution performance and to reduce the burden in terms of computational resources needed; to this aim, we will exploit a direct formulation of the predictive information that can be found in [2]. Moreover, we will also explore other information-related metrics to understand how the exchange of information in a network of embodied agents is related to the emergence of suitable behavior. ACKNOWLEDGMENT This work has been partially supported by the PRIN project BEASTIE (2022FZJ7MR), funded by the Italian Ministry of University and Research and by the European Union – NextGenerationEU, and supported by EU project Horizon 2020 AIFORS 952275. REFERENCES [1] I. Tanev, T. Ray, and A. Buller, “Automated evolutionary design, robustness, and adaptation of sidewinding locomotion of a simulated snake-like robot,” IEEE Transactions on robotics, vol. 21, no. 4, pp. 632–645, 2005. [2] F. Bonsignorio, “Quantifying the evolutionary self-structuring of embodied cognitive networks,” Artificial life, vol. 19, no. 2, pp. 267–289, 2013. [3] G. S. Chirikjian and J. W. Burdick, “A modal approach to hyperredundant manipulator kinematics,” IEEE Transactions on Robotics and Automation, vol. 10, no. 3, pp. 343–354, 1994. [4] G. S. Chirikjian, Stochastic Models, Information Theory, and Lie Groups, Volume 1: Classical Results and Geometric Methods. Springer Science & Business Media, 2009. [5] G. S. Chirikjian and G. S. Chirikjian, “Information, communication, and group theory,” Stochastic Models, Information Theory, and Lie Groups, Volume 2: Analytic Methods and Modern Applications, pp. 271–312, 2012. [6] M. Lungarella and O. Sporns, “Information self-structuring: Key principle for learning and development,” in Proceedings. The 4th International Conference on Development and Learning, 2005. IEEE, 2005, pp. 25–30. [7] M. Lungarella, T. Pegors, D. Bulwinkle, and O. Sporns, “Methods for quantifying the informational structure of sensory and motor data,” Neuroinformatics, vol. 3, pp. 243–262, 2005. [8] M. Prokopenko, V. Gerasimov, and I. Tanev, “Evolving spatiotemporal coordination in a modular robotic system,” in From Animals to Animats 9: 9th International Conference on the Simulation of Adaptive Behavior (SAB 2006) Lecture Notes in Computer Science, S. Nolfi, G. Baldassarre, R. Calabretta, J. Hallam, D. Marocco, J. Meyer, O. Miglino, and D. Parisi, Eds., vol. 4095. Berlin–Heidelberg, Germany: Springer, 2007, pp. 558–569. [9] G. Mont´ ufar, K. Ghazi-Zahedi, and N. Ay, “Information theoretically aided reinforcement learning for embodied agents,” CoRR, vol. abs/1605.09735, 2016. [Online]. Available: http://arxiv.org/abs/1605.09735 [10] F. Bonsignorio, “Preliminary considerations for a quantitative theory of networked embodied intelligence,” in Fifty years of AI, M. Lungarella, F. Iida, J. Bongard, and R. Pfeifer, Eds. Berlin–Heidelberg, Germany: Springer, 2007. [11] W. Bialek, I. Nemenman, and N. Tishby, “Predictability, complexity and learning,” 2001. [Online]. Available: https://arxiv.org/abs/physics/0007070 [12] A. J. Ijspeert, “Central pattern generators for locomotion control in animals and robots: A review,” Neural Networks, vol. 21, no. 4, pp. 642–653, May 2008. [13] E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE, 2012, pp. 5026–5033. [14] M. Towers, A. Kwiatkowski, J. Terry, J. U. Balis, G. D. Cola, T. Deleu, M. Goul˜ ao, A. Kallinteris, M. Krimmel, A. KG, R. Perez-Vicente, A. Pierr´ e, S. Schulhoff, J. J. Tai, H. Tan, and O. G. Younis, “Gymnasium: A standard interface for reinforcement learning environments,” 2024. [Online]. Available: https://arxiv.org/abs/2407.17032 [15] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, TABLE II: Variables used in the system Variable Definition IpPredictive information (how much observations of the past help predict the future). I(xt+1 t−τ+1;xt t−τ)Mutual information between the random variables “past” and “future”. TTime duration of the observations. H(x)Shannon entropy of a stochastic variable x, defined as H(x) = Px∈Xp(x) ln p(x). p(x)Probability distribution associated with a stochastic variable x. p(x|y)Conditional probability of event xgiven that event yhas occurred. h1Entropy rate: additional information needed to represent the next observation of one of the two stochastic processes. h2Same as above, but assuming xis independent of y. TE Transfer entropy (information flow): directed, time-asymmetric information transfer between two random processes. TABLE III: Hyperparameters used in the simulations Hyperparameter Value Mujoco timestep 0.001 s Control frequency 100 Hz Mini batch-size 64 Total training steps 8,192,000 Rollout buffer 2048 Number of epochs 10 Learning rate Linear scheduler starting from 0.0003, decreasing proportionally to (1 −%pr), where pr is the training progress Discount factor 0.99 GAE factor 0.95 Episode steps 8192 Clip range 0.2 Parallel environments 2 and O. Klimov, “Proximal policy optimization algorithms,” CoRR, vol. abs/1707.06347, 2017. [Online]. Available: http://arxiv.org/abs/1707.06347 [16] A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann, “Stable-baselines3: Reliable reinforcement learning implementations,” Journal of Machine Learning Research, vol. 22, no. 268, pp. 1–8, 2021. [Online]. Available: http://jmlr.org/papers/v22/201364.html [17] Y. Dodge, Autocorrelation. New York, NY: Springer New York, 2008, pp. 24–26. [Online]. Available: https://doi.org/10.1007/978-0387-32833-1 17