TA12-378: Fault Resilience Analysis of the Dynamic-TMR RISC-V Fault Tolerant Microprocessor Design
Abstract
RADNEXT Transnational Access Summary Report
Full text
https://radnext.web.cern.ch/ https://www.linkedin.com/company/radnext EDMS NO. 3355743 VALIDITY Released REV. 1.0 DOI: 10.5281/zenodo.17288071 RADNEXT Transnational Access Summary Report Project title Fault Resilience Analysis of the Dynamic-TMR RISC-V Fault Tolerant Microprocessor Design Project TA identifier TA12-378 General application Space, high-reliability ground level Type of test SEE Group leader, Institute Marcello Barbirotta, Sapienza University of Rome Co-authors, Institutes Marco Angioli, Mauro Olivieri, Antornio Mastrandrea Date(s) of the experiment 24/06/2025 – 25/06/2025 Facility HollandPTC Amount of access granted 8 h Objectives of the experiments Fault Tolerance (FT) ensures continuous operation despite faults by enabling isolation and recovery without system failure. Growing interest in FT-COTS embedded microprocessors has led to various studies over the years testing FPGAs under different radiation injection facilities. While current FPGA technologies can be characterized with near-accuracy, most designs implemented on them still rely on traditional redundancy methods, like temporal or spatial redundancy, without introducing new innovations. In this study, we propose investigating the fault reliability of an innovative approach involving a microarchitecture Interleave Multi-Threading (IMT) RISC-V design capable of dynamically switching from Dual Modular Redundancy (DMR) to Triple Modular Redundancy (TMR) in case of faults. This proposal fits the vast scenario of microprocessor devices for FT COTS embedded applications, having as its basis the first detailed exploitation of the Interleave Multi-Threading (IMT) execution scheme for implementing fault tolerant processors on FPGA. This work fills a gap in the literature by analyzing the FT potential of IMT architectures: • It validates and demonstrates the FT advantage of the proposed Dynamic-TMR method over existing lockstep techniques by functioning as a DMR-lockstep dual core. This approach avoids the overhead of duplicating all hardware and aims to significantly reduce checkpointing and rollback routines. • It quantifies the efficiency of the technique through extensive FI campaigns under radiation conditions, showing that the resilience achievable is comparable to a lockstep dual core. This makes the method suitable for systems with high fault rates or critical applications. • It compares the implementation inside the Klessydra-dfT03 architecture with the unprotected version and the protected Buffered TMR Klessydra-fT03 core, validating the IMT FT approaches across the Klessydra family and paving the way for future comparisons with other RISC-V FT processors. • The significance of this work lies in its potential to produce Failure Probability data comparable with results from Single-Event-Upset fault injection tests, useful for metrics such as Mean-Time-to-Failures EDMS 3355743 v.1 status In Work access Restricted PDF from TA12-378-Zenodo.doc modified 2025-10-07 16:12
https://radnext.web.cern.ch/ https://www.linkedin.com/company/radnext EDMS NO. 3355743 VALIDITY Released REV. 1.0 (MTTF), Mean-Work-to-Failures (MWBF), Architectural-Correct-Execution (ACE) bits, and Architectural-Vulnerability-Factor (AVF). The overall approach reflects the tests previously conducted at PSI’s Proton Irradiation Facility. Although FT processors are ideally tested with heavy ions that simulate space conditions, we chose to repeat these tests using the same proton sources. This permits a clear comparison with earlier data and related research on RISC-V processors with proton radiation. Proton testing was selected due to the limited existing literature, which predominantly focuses on heavy ions or neutrons. Experiment test report The tests were carried out on different FPGA boards, numbered 1, 3, 4, and 5, from the same FPGA family, Xilinx Artix7, on a CMODA7 Digilent Board. First, the beam was calibrated with High energy = 70 MeV with the following fluence and flux: Current= 3nA - flux= 1,22x10^7 p/cm^2/s - fluence = 1,43x10^9 p/cm^2 Current= 21nA - flux= 8,55x10^7 p/cm^2/s - fluence = 1x10^10 p/cm^2 Current= 7nA - flux= 2,85x10^7 p/cm^2/s - fluence = 3,3x10^9 p/cm^2 Current= 14nA - flux= 5,70x10^7 p/cm^2/s - fluence = 6,67x10^9 p/cm^2 Current= 17nA - flux= 6,92x10^7 p/cm^2/s - fluence = 3,32x10^9 p/cm^2 Current= 10nA - flux= 4,07x10^7 p/cm^2/s - fluence = 1,95x10^9 p/cm^2 Current= 13nA - flux= 5,29x10^7 p/cm^2/s - fluence = 2,54x10^9 p/cm^2 Second, we set up the environment by putting the DUT under the Beam and connecting it to its control board, as shown in the image below.
https://radnext.web.cern.ch/ https://www.linkedin.com/company/radnext EDMS NO. 3355743 VALIDITY Released REV. 1.0 After a few simple runs, the beam control software and the control board manually synchronized each test for the activation or deactivation of both the board and the beam, leading to data readback from the FPGA. The 8 hours were divided into 2 different test campaigns for the two different days, each with different runs for each design version (T03 and dfT03). In particular, were carried out: Day1 70MeV 7 runs for dfT03 -> 3nA -> Board n°1 12 runs for dfT03 -> 21nA -> Board n°1 7 runs for dfT03 -> 3nA -> Board n°3 8 runs for dfT03 -> 21nA -> Board n°3 Day2 70MeV 4 runs for dfT03 -> 14nA -> Board n°3 10 runs for dfT03 -> 17nA -> Board n°3 10 runs for T03 -> 10nA -> Board n°5 9 runs for T03 -> 13nA -> Board n°5 10 runs for T03 -> 17nA -> Board n°4 10 runs for T03 -> 10nA -> Board n°4 12 runs for T03 -> 13nA -> Board n°4 At the beginning of each test, golden runs were carried out for direct comparison with subsequent faulted runs. Moreover, before and after each run, bitstream readback was performed for bit comparison and error checkout. The final 99 collected tests will be subjected to post-processing analysis and cataloguing of the results.
https://radnext.web.cern.ch/ https://www.linkedin.com/company/radnext EDMS NO. 3355743 VALIDITY Released REV. 1.0 Outcome of the experiments Please indicate what the experiment is likely to lead to by putting an ‘X’ next to one or more of the possible outcomes below. Journal publication X Data for Thesis Follow-up experiment at same facility Follow-up experiment at another facility Other As a RADNEXT user, we encourage you to submit the scientific results of your experiments to journals as well as to the NSREC and RADECS data workshops. Please remember to include the RADNEXT acknowledgment into your publications! RADNEXT acknowledgment: This project has received funding from the European Union's Horizon 2020 research and innovation programme under grant agreement No 101008126.