A Digital Twin Based Reconfigurable Intelligent Surface Phase Adaptation Using Spiking Reinforcement Learning Policy Optimization
Abstract
This demo presents a digital twin of a reconfigurable intelligent surface-empowered wireless system that employs spiking reinforcement learning (SRL) optimization policy for phase adaptation in order to maximize the network coverage, while minimizing the energy consumption at both the microcontroller and transmission related processes. The demo assesses the efficiency of SRL against conventional deep reinforcement learning approaches in terms of (i) energy consumption, (ii) reduction of training latency, (iii) probability of outage, and (iv) bit error rate.
Full text
A Digital Twin Based Reconfigurable Intelligent Surface Phase Adaptation Using Spiking Reinforcement Learning Policy Optimization Ilias Crysovergis Department of Informatics &Telecommunications, University of Thessaly, Lamia, Greece ichrysover[email protected] Stylianos E. Trevlakis InnoCube IKE Thessaloniki, Greece tre[email protected]g Dimitris Kleitsas METATOPIA LP Thessaloniki, Greece [email protected] Alexandros-Apostolos A. Boulogeorgos Department of Electrical and Computer Engineering University of Western Macedonia Kozani, Greece [email protected] Theodoros A. Tsiftsis Department of Informatics & Telecommunications University of Thessaly Lamia, Greece [email protected] Dusit Niyato College of Computing and Data Science Nanyang Technological University Nanyang Avenue, Singapore [email protected] Abstract—This demo presents a digital twin of a reconfigurable intelligent surface-empowered wireless system that employs spiking reinforcement learning (SRL) optimization policy for phase adaptation in order to maximize the network coverage, while minimizing the energy consumption at both the microcontroller and transmission related processes. The demo assesses the efficiency of SRL against conventional deep reinforcement learning approaches in terms of (i) energy consumption, (ii) reduction of training latency, (iii) probability of outage, and (iv) bit error rate. Index Terms—Digital twin (DT), spiking neural networks (SNN), spiking reinforcement learning (SNL), deep reinforcement learning (DRL). I. INTRODUCTION,MOTIVATION AND RATIONALIZATION Digital twins (DTs) are expected to become a key element of next-generation wireless networks, since they are capable of turning the stochastic propagation environment to a deterministic system. Meanwhile, reconfigurable intelligent surfaces (RISs) are capable of exploting the wireless channel randomess, creating favorable propagation conditions that boost the performance of the wireless system and network. The marriage of DTs and RISs is expected to turn wireless systems into software platforms with flexible and energy efficient manipulations. However, as the wireless environment becomes increasingly complex, the efficiency and performance of conventional optimization approaches become questionable. To cover this gap, a solution that is widely investigated is the use of deep reinforcement learning (DRL) methods, which retrieve information from the DT, and provide either optimal or sub-optimal configurations. These DRL approaches are expected to be executed in the RIS control/configuration unit. From the control unit point of view, several different computing architectures, including application-specific integrated circuit (ASIC) low-power (LP) complementary metaloxide semiconductor (CMOS), ASIC CMOS, ASIC fin fieldeffect transistor (FinFET) CMOS, field programmable gate array (FPGA), computing processing unit (CPU), and graphical processing unit (GPU), were employed. However, the aforementioned architectures either experience high response time, i.e., in the orders of µs to ms, or have dramatically high power consumption that may even reach some hundreds Watts. Fortunately, a new brain-inspired architecture has very recently been developed, namely neuromorphic computing architecture. Neuromoprhic computing processing achieves significantly low response time, which is in the orders of some ns, while their power consumption is in the order mW. The use of DRL in neuromorphic processing units has been shown to be neither energy-efficient nor time-efficient. Thus, new approaches need to be tested. This motivates a paradigm shift towards spiking reinforcement learning (SRL). SRL integrates spiking neural networks with reinforcement learning (RL), offer important advantages over algorithms like proximal policy optimization (PPO) and soft actor-critic (SAC) in terms of enhanced energy efficiency due to the sparse firing nature of spiking neurons, and improved robustness from the inherent noise tolerance in spiking networks. As a consequence, SRLs are expected to converge faster in certain RIS-tailored scenarios than both PPO and SAC. Additionally, SRL leverages hardware-efficient computations in neuromorphic processors. In this direction, this demo presents a DT of an RIS-aided wireless system that employs SRL policy optimization for phase adaptation. The demo will demonstrate the efficiency of the SRL against conventional DRL approaches in terms of (i) energy consumption, (ii) training latency reduction, (iii) outage probability, and (iv) bit error rate. The SRL model will retrieve information from a DT that operate in real time. In other words, the users will have the opportunity to: This is the accepted manuscript version of the paper prior to IEEE formatting and copyediting. © 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses.
•Watch a time-evolving DT; •See quantified benefits of SRL against conventional DRL; and •Observe the impact of the use of RIS coupled with DT and SRL approaches on the performance of the communication system. The rest of the paper is organized as follows: The overall architecture including the DT ans SRL frameworks are articulated in Section II. Initial indicative results are presented in Section III. II. OVERALL ARCHITECTURE The blueprint of the proposed DT-empowered SRL-training framework is presented in Fig. 1. It consists of two classes, i.e., i) environment, which is a ray-tracing based time-evolving simulation framework, and ii) an agent that is where the SRL model is trained and provide it outcomes to the environment. The remainder of this section focuses on articulating the functionalities and operations of each one of the aforementioned classes. A. A DT based simulation framework Fig. 2 presents a three dimensional screenshot of the digital twin (DT). The geographical data was derived by open-street maps. The area of interest is in Barcelona and is defined by the following coordinates: •(41.38885246115754,2.1969617049647394), •(41.39566172681428,2.206242147917081), •(41.39980649791837,2.200641695785468),and •(41.393214980646256,2.191318337812094). The ground is presented in light gray and has been modeled as a concrete material. The electromagnetic properties for the ground are based on the corresponding ITU recommendation [1]. The walls are colored beige and are modeled as bricks [1]. The top of the buildings is painted darg gray and is modeled as concrete. A secondary cylinder coordination system (CCS) was defined for each node of the network in order to define their orientation. The origin of the secondary CCS is consistent with the node position. The xand yaxes of the corresponding Cartesian coordination system are parallel to the horizontal and vertical directions of Fig. 2. The zaxis is perpendicular to the x−yplane. In Fig. 2, we represent the gNodeB by light blue circle. The gNodeB is located at (41.39452591852527,2.2005211400272082) on a roof of height 45 m. The gNodeB orientation is given by the vector (0,0,0). The gNodeB is equipped by a tr38901 antenna [2] of 4×4elements and a beam-pattern that is depicted in Fig. 3. The antenna gain is equal to 15 dBi. The distance between neighboring elements is equal to λ/2, where lambda stands for the transmission wavelength. Finally, we assumed a vertical polarization. A downlink scenario is considered in which the gNodeB communicates with a mobile user equipment (MUE) through a neuromoprhic-controlled-RIS (neuroRIS) [3]. The RIS is located at the front of a building. The RIS coordinates are (41.39647643475048,2.1982684063213482) and at a height of 40 m. The neuroRIS is equipped with 20 ×20 metaatoms that are connected through a switching circuit to the neuromorphic processor, which is responsible for predicting the MUE position in the next timeslot. To achieve this, a control interface between a distributed unit of the open-radio access network (O-RAN) and the gNodeB is established. The interface may use either a low-data rate radio link or a power line communication system and forward signal quality indicators (SQIs) to the neuroRIS microcontroller. Note, this is a realistic assumption as both the gNodeB and the neuroRIS are connected to the same power grid. Moreover, notice that since the neuroRIS microcontroller receives SQIs, the second type of quantization errors exists. The motion of the MUE is a constricted random walk in that is defined by the coordinates: •(41.397285537979805,2.1992953701040503), •(41.3949792088333,2.1962280281227984), •(41.39485627358036,2.1963099664103347) and •(41.39714657260773,2.19939155852855). The MUE height is considered constant and equal to h. We assume that between two consequence points (Aiand Ai+1) of the random walk process is linear of constant velocity that is equal to vi,i+1. The velocity is modeled as a random process that follows uniform distribution and is in the range of [0.8,1.2] m/s. For the MUE orientation, we consider that is given by a deterministic and two random processes, which define the vector (r, ϕ, θ).ris set to 15 cm,ϕand θare independent identical random processes that follow uniform distribution in the range of [−π/8, π/8]. The MUE is equipped with an isotropic antenna. The MUE antenna pattern is presented in Fig. 4. In Fig. 2, we use light green color to represent the position of the MUE. In this screenshot, the MUE position is (41.396962340649566,2.1990733259398847) and it is in a height of 1.7m. In the screenshot presented in Fig 2, the MUE orientation is given by the vector (0,0,0). Next, we articulate the communication protocol. We set the transmission frequency at 3 GHz. Without loss of generality, we use a binary process to generate the information bits at the gNodeB. Low-density parity-check (LDPC) coding is employed with coding rate equals to 0.5in order to enhance the communication system reliability. Moreover, pilot bits are added to the transmission bit frame. The bits are first mapped through a 16 quadrature amplitude modulator (QAM) mapper. The extracted symbols are organized in an orthogonal frequency division modulation (OFDM) stream that consists of 14 subcarriers and is depicted in Fig. 5, while the corresponding OFDM resource grid is illustrated in Fig. 6. Of note zeropadding is employed for mitigating the impact of inter-symbol interference. A direct current (DC) carrier is loaded in the central transmission of each resource grid to enable hardware imperfections mitigation. The OFDM symbols are forwarded towards the digital-toanalog converter, which output is inputted in the transmission filter. The transmission filter response is a raised cosine. The output of the transmission filter is forwarded towards the
Fig. 1. The blueprint of the proposed DT-empowered SRL-training framework. Fig. 2. Simulation scenario power amplifier (PA) that ensures that the transmission power is equal to 30 dBm. We assume a linear PA operation. The PA output is inputted to the transmission antenna. The voltage signal is converted by the transmission antenna to an electromagnetic wave that travels towards the neuroRIS. Due to the long gNodeB-neuroRIS transmission distance, the propagation wave has far-field characteristics. To model the wireless channel, we employ an RT approach. The maximum number of reflection that we account for is equal to 5. Figure 7 Fig. 3. gNodeB beam-pattern. presents a screenshot of the power delay profile (PDP) of the gNodeB-neuroRIS channel. The neuroRIS steers the incident wave towards the predicted position of the MUE. Figure 8 presents a screenshot of the radiation pattern of the RIS. By accounting the relative position of the MUE receiver and the neuroRIS, as well as the beam pattern of the neuroRIS, af-
Fig. 4. MUE beam-pattern. Fig. 5. OFDM transmission stream. Fig. 6. OFDM resource grid. Fig. 7. PDP of he gNodeB-neuroRIS channel. Fig. 8. RIS radiation pattern. ter the configuration, we evaluate the reflected power. Note that the neuroRIS configuration is performed by an SRL model that is running in the neuroRIS micro-controller. The SRL model has been previously trained offline. A ray-tracing procedure is employed to compute the PDP between the RIS and the MUE. The MUE antenna transforms the captured wave into a voltage signal and forwards it towards the low-noise amplifier (LNA). We assume that LNA work in the linear operation area. However, both the antenna and the LNA add white Gaussian noise to the received signal. Note that the additive white Gaussian noise (AWGN) is model as a zero-mean Gaussian process. The variance of AWGN is calculated based on the atmospheric and electronics temperature as well as the signal bandwidth. After the LNA, a reception filter is employed to mitigate the impact of inter-symbol interference, which is followed by an analog-to-digital converter (ADC). The output
Fig. 9. The agent architecture. samples of the ADC are forwarded toward the receivers digital part that performs guard and DC carrier removal, in order to fully mitigate the impact of inter-symbol interference and hardware imperfections, and inverse fast Fourier transform, in order to extract the (distorted due to fading and adaptive white Gaussian noise) received symbols. In this point, the received sample consists of pilot and information signals that are saved in different matrices. The i−th row of each matrix carries the symbol that was loaded at the i−th subcarrier, while the j−th column, the one that send at the j−th time-slot. Two copies of the received symbols, extracted from the OFDM pilots, are created. The first copy is forwarded to a conventional received power estimator. The received power estimator, which leverages prior-knowledge of the noise power by performing power estimation during periods of receiver inactivity. As a consequence, it is able to extract the received SNR. The second copy is fed into the channel estimator. The channel estimator additionaly employing a minimum least squares approach. The channel estimator additionally incorporates the estimated SNR as a confidence metric, where higher SNR values indicate greater reliability and lead to improved channel estimation accuracy. The estimated baseband equivalent channel coefficients, along with the information symbols are forwarded to the equalizer. The equalizer processes each information symbol at the k-th subcarrier by dividing it by the corresponding estimated channel coefficient. The equalizer output is then passed to the maximum likelihood detector (MLD), which produces an estimation of the transmitted symbol. B. SNN for RIS phase adaptation As depicted in Fig. 9, the agent consists of two SNNs, which are designed to optimize an RIS in a wireless communication system, working together to maximize the total data rate for multiple MUEs. The first SNN decides which MUEs should use the RIS by predicting the association matrix, a binary vector of length equal to the number of the MUEs. It takes as input the channel gains with and without the RIS for all MUEs, processes them through hidden layers and NRoutput layers Fig. 10. NeuroRIS coverage gain map. (producing 2 options per receiver: 0or 1), and outputs spike counts over Ttime steps to sample the association decisions. The second SNN determines the phase shifts for each of the PRIS elements to enhance signal strength for the receivers using the RIS. It takes as input the 3D positions of only those receivers connected to the RIS, as specified by the association matrix, resulting in a dynamic input size that depends on the number of connected MUEs. This SNN processes the input through hidden layers and Poutput layers (one per RIS element, each outputting A options for phase shifts), generating spike counts over Ttime steps to sample phase shift actions (e.g., 0, π/2,π,3pi/2). Both SNNs use leaky integrate-and-fire (LIF) neurons with a decay factor of 0.9and a surrogate gradient for training, and they are trained using policy gradients to maximize the total data rate, with the first SNN guiding which MUEs the second SNN optimizes. III. INDICATIVE RESULTS & DEMONSTRATOR’S FRONT-END Figure 10 presents a screenshot of the coverage gain that is experienced by the use of the RIS in comparison with the coverage in the absence of the neuroRIS. It is worth reporting that the neuroRIS assists in limiting the secondary lobes and focusing the transmission towards the MU. As a consequence, a coverage gain of more than 120 dB is observed at the MU reception area. Figure 11 presents the expected energy consumption benefits in terms of average energy consumption as a function of the number of meta-atoms for different type of neural networks (NNs). From this figure, it becomes evident that as the number of meta-atoms increases, the energy consumption gain of spiking neural networks instead of artificial neural networks increases.
8 16 32 64 N 10-3 10-2 10-1 100 101 102 Average Energy ( J) SNN ANN Fig. 11. Average energy consumption vs number of RIS for different type of neural networks [3]. Fig. 12. Demonstrator software front-end screenshot. Finally, Fig. 12 presents a screenshot of the demonstrator front-end. The geographical area, transmission and reception parameters, type of antennas, and neuroRIS characteristics, and machine learning models are adjustable. IV. REQUESTED EQUIPMENT The demo setup consists of a laptop with a powerful graphic processing unit (GPU) that will be provided by the presenters. The laptop need to be connected to the internet through either an Ethernet or a wireless fidelity (wi-fi) connection. A 100 Mbps data-rate would be desirable. Additionally, we are going to need: •One power socket 230V/50Hz AC (please bring your power stripe if you need multiple sockets). •One table (180 cm ×50 cm) and 2chairs. ACKNOWLEDGMENT The work of S. E. Trevlakis and and T. Tsiftsis are supported by the European Unions Horizon-CL4-2021 research and innovation programme under grant agreement No. 101070181 (TALON). The contribution of A.-A. A. Boulogeorgos was supported by the research project MINOAS. The project MINOAS is implemented in the framework of H.F.R.I called “Basic Research Financing (Horizontal support of all Sciences); under the National Recovery and Resilience Plan “Greece 2.0” funded by the European Union – NextGenerationEU (H.F.R.I. Project Number: 15857). ORGANIZER/AUTHOR/PRESENTER OF DEMO Table I summarizes the organizers and presenters information. Next, the short CVs of the organizers are presented. Ilias Crysovergis holds a 5-year diploma from the ECE Department of Aristotle University of Thessaloniki (AUTh) and an MSc in Communications & Signal Processing from Imperial College London. He is currently a PhD candidate in the Department of Informatics and Telecommunications at the University of Thessaly, Greece. Ilias is a software architect with expertise in Extended Reality, Digital Twins, and Artificial Intelligence. He has received notable awards and distinctions from organizations such as Microsoft, Hong Kong Polytechnic University, the Chief of the Greek Army, the Rector of AUTh, and the Minister of Economics of Cyprus. Stylianos E. Trevlakis (Member, IEEE) was born in Thessaloniki, Greece in 1991. He received the ECE diploma and PhD both from the AUTh in 2016 and 2022. From April 2022 until now, Dr. Trevlakis works at InnoCube as research director. His research interests lie in the area of Wireless Communications, with emphasis on conventional & AI-enabled Wireless Communication Systems, as well as Communications & Signal Processing for Biomedical Engineering. Dimitris Kleitsas is a Machine Learning researcher at Metatopia and a senior undergraduate student in the Department of Computer Science at AuTH. He specializes in AI, with experience in applying advanced AI algorithms to telecommunications systems and fine-tuning language models for automated content review. His expertise encompasses deep supervised and unsupervised learning, reinforcement learning, large language models, and knowledge systems. Alexandros-Apostolos A. Boulogeorgos (Senior Member, IEEE) received the Diploma degree in Electrical & Computer Engineering (ECE) and the Ph.D. degree in wireless communications from Aristotle University of Thessaloniki in 2012 and 2016, respectively. From 2022, he is an Assistant Professor at the Department ECE of the University of Western Macedonia, Greece. He is listed in “World’s Top 2% Scientists” for the Year 2022 and 2023,” which is published by Stanford University and Elsevier. He is an IEEE Senior Member. He is an associate Editor for IEEE Transactions on Wireless Communications and IEEE Communications letters. His research interests fall into the broad areas of communication theory, with an emphasis on wireless communications theory, optical wireless communications, non-conventional communications and neuromorphic-empowered wireless systems. Theodoros A. Tsiftsis (Senior Member, IEEE) received the PhD degree in electrical engineering from the University of Patras, Greece, in 2006. He is a professor with the
TABLE I ORGANIZER//PRESENTER OF DEMO Name Position in the organization Affiliation Country E-mail address Ilias Crysovergis Ph. D. Candidate Department of Informatics & Telecommunications, University of Thessally Greece ichrysov[email protected] Stylianos E. Trevlakis Research director InnoCube IKE Greece tre[email protected] Dimitris Kleitsas Researchers Metatopia LTD Greece [email protected] Alexandros-Apostolos A. Boulogeorgos Assistant Professor Department of Electrical and Computer Engineering, University of Western Macedonia Greece [email protected] Theodoros Tsiftsis Professor Department of Informatics & Telecommunications, University of Thessally Greece [email protected] Dusit (Tao) Niyato Professor College of Computing and Data Science, Nanyang Technological University Singapure [email protected] Department of Informatics & Telecommunications, University of Thessaly, Greece, and also an honorary professor with Shandong Jiaotong University, Jinan City, China. His research interests fall into the broad areas of communication theory and wireless communications, with an emphasis on wireless communications theory, reconfigurable intelligent surfaces, optical wireless communications, and physical layer security. He served on the Editorial Boards of the IEEE TRANSACTIONS ON COMMUNICATIONS, IEEE TRANSACTIONS ON VEHICULAR TECHNOLOGY, IEEE COMMUNICATIONS LETTERS, and IEEE TRANSACTIONS ON MOBILE COMPUTING. He is an Associate Editor of the IEEE TRANSACTIONS ON WIRELESS COMMUNICATIONS, and Specialty Chief Editor for Networks and Communications of Frontiers in Computer Science. He was appointed as an IEEE Vehicular Technology Society Distinguished Lecturer for two terms (2018–2022) and was recently appointed as an IEEE Communications Society Distinguished Lecturer (2024–2025). Dusit Niyato (Fellow, IEEE) is a professor in the College of Computing and Data Science, at Nanyang Technological University, Singapore. He received B.Eng. from King Mongkuts Institute of Technology Ladkrabang (KMITL), Thailand and Ph.D. in Electrical and Computer Engineering from the University of Manitoba, Canada. His research interests are in the areas of mobile generative AI, edge intelligence, decentralized machine learning, and incentive mechanism design. REFERENCES [1] International Telecommunication Union Radiocommunication Sector (ITU-R), “Effects of building materials and structures on radiowave propagation above about 100mhz,” ITU-R, Recommendation P.2040, August 2023. [Online]. Available: https://www.itu.int/rec/R-REC-P.2040 [2] ETSI, “Study on channel model for frequencies from 0.5 to 100 ghz (3gpp tr 38.901 version 16.1.0 release 16),” European Telecommunications Standards Institute (ETSI), Technical Report TR 138 901, July 2020, v16.1.0 (2020-07). [Online]. Available: https://www.etsi.org/deliver/etsi tr/138900 138999/138901/16.01.00 60/tr 138901v160100p.pdf [3] C. G. Tsinos, A.-A. A. Boulogeorgos, and T. A. Tsiftsis, “Neuroris: Neuromorphic-inspired metasurfaces,” IEEE Wireless Communications Letters, vol. 13, no. 7, pp. 1878–1882, 2024.