Scalable Sensor Fusion for Motion Localization in Large RF Sensing Networks
Full text
Charting the Intelligence Frontiers Edge AI Systems Nexus
RIVER PUBLISHERS SERIES IN COMMUNICATIONS AND NETWORKING Series Editors: ABBAS JAMALIPOUR MARINA RUGGIERI The University of Sydney University of Rome Tor Vergata Australia Italy MARKO JURCEVIC University of Zagreb Croatia The “River Publishers Series in Communications and Networking” is a series of comprehensive academic and professional books which focus on communication and network systems. Topics range from the theory and use of systems involving all terminals, computers, and information processors to wired and wireless networks and network layouts, protocols, architectures, and implementations. Also covered are developments stemming from new market demands in systems, products, and technologies such as personal communications services, multimedia systems, enterprise networks, and optical communications. The series includes research monographs, edited volumes, handbooks and textbooks, providing professionals, researchers, educators, and advanced students in the field with an invaluable insight into the latest research and developments. Topics included in this series include: • Communication theory • Multimedia systems • Network architecture • Optical communications • Personal communication services • Telecoms networks • Wifi network protocols For a list of other books in this series, visit www.riverpublishers.com
Charting the Intelligence Frontiers Edge AI Systems Nexus Editors Ovidiu Vermesan SINTEF, Norway Alain Pagani German Research Center for Artificial Intelligence, Germany Paolo Meloni University of Cagliari, Italy River Publishers
Published, sold and distributed by: River Publishers Broagervej 10 9260 Gistrup Denmark www.riverpublishers.com ISBN: 978-87-4380-884-8 (Hardback) 978-87-4380-883-1 (Ebook) ©The Editor(s) and The Author(s) 2025. This book is published open access. Open Access This book is distributed under the terms of the Creative Commons Attribution-NonCommercial 4.0 International License, CC-BY-NC 4.0 (http://creativecommons.org/ licenses/by/4.0/), which permits use, duplication, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, a link is provided to the Creative Commons license and any changes made are indicated. The images or other third party material in this book are included in the work’s Creative Commons license, unless indicated otherwise in the credit line; if such material is not included in the work’s Creative Commons license and the respective action is not permitted by statutory regulation, users will need to obtain permission from the license holder to duplicate, adapt, or reproduce the material. The use of general descriptive names, registered names, trademarks, service marks, etc. in this publication does not imply, even in the absence of a specific statement, that such names are exempt from the relevant protective laws and regulations and therefore free for general use. The publisher, the authors and the editors are safe to assume that the advice and information in this book are believed to be true and accurate at the date of publication. Neither the publisher nor the authors or the editors give a warranty, express or implied, with respect to the material contained herein or for any errors or omissions that may have been made. Printed on acid-free paper.
Dedication “Time stays long enough for those who use it.” – Leonardo da Vinci “Pleasure in the job puts perfection in the work.” – Aristotle “Success in the AI era will belong to those who adapt, learn, and innovate continuously.” – Anonymous “It is not the strongest species that survive, nor the most intelligent, but the ones most responsive to change.” – Charles Darwin Acknowledgement The editors would like to thank all the contributors for their support in the planning and preparation of this book. The recommendations and opinions expressed in the book are those of the editors, authors, and contributors and do not necessarily represent those of any organizations, employers, or companies. Ovidiu Vermesan Alain Pagani Paolo Meloni
Contents Preface xix List of Figures xxiii List of Tables xxix List of Contributors xxxi List of Abbreviations xxxv 1 Edge AI Systems Verification and Validation 1 Ovidiu Vermesan, Alain Pagani, Roy Bahr, Marcello Antonio Coppola, and Giulio Urlini 1.1 Introduction and Background . . . . . . . . . . . . . . . . . 2 1.2 Foundational Concepts and Edge AI Verification and ValidationTaxonomy ........................ 6 1.2.1 Agentic AI and AI Agents . . . . . . . . . . . . . . 12 1.3 Defining Verification and Validation per Standard . . . . . . 14 1.4 Key Elements for Edge AI Verification and Validation . . . 17 1.4.1 Core Elements for AI Verification . . . . . . . . . . 19 1.4.1.1 Data Verification . . . . . . . . . . . . . . 19 1.4.1.2 Model Verification . . . . . . . . . . . . . 20 1.4.1.3 System-Level Verification . . . . . . . . . 21 1.4.1.4 Process and Governance Verification . . . 22 1.4.2 Core Elements Subject to AI Validation . . . . . . . 23 1.4.2.1 Ensuring Fitness for Intended Purpose and Operational Context . . . . . . . . . . . . 23 1.4.2.2 Meeting User Needs and Stakeholder Expectations . . . . . . . . . . . . . . . . 23 1.4.2.3 Assessing Real-World Effectiveness and Outcomes ................. 24 vii
viii Contents 1.4.2.4 Evaluating Usability and Human-AI Interaction ................... 24 1.4.2.5 Validating Ethical Alignment and Societal Impact................... 24 1.4.2.6 Data Quality and Suitability . . . . . . . . 25 1.5 The Edge AI Verification and Validation Lifecycle . . . . . . 26 1.6 Failure Case Behaviour in Edge-based Machine Vision Systems............................ 29 1.7 Research Challenges in Edge AI Verification andValidation......................... 31 1.8 Trends and Methodologies in Edge AI Verification and Validation........................... 39 1.9 Conclusion .......................... 41 2 Pioneering the Hybridization of Federated Learning in Human Activity Recognition 53 Alfonso Esposito, Yasamin Moghbelan, Ivan Zyrianoff, Leonardo Ciabattini, Federico Montori, and Marco Di Felice 2.1 Introduction and Background . . . . . . . . . . . . . . . . . 53 2.2 Hybrid FL Architecture . . . . . . . . . . . . . . . . . . . . 55 2.3 Evaluation Methodology and Metrics . . . . . . . . . . . . . 57 2.4 Evaluation Results . . . . . . . . . . . . . . . . . . . . . . 59 2.5 Conclusion and Future Works . . . . . . . . . . . . . . . . . 62 3 Edge Intelligence Architecture for Distributed and Federated Learning Systems 65 Pierluigi Dell’Acqua, Lorenzo Carnevale, and Massimo Villari 3.1 Introduction.......................... 66 3.2 RelatedWorks......................... 67 3.2.1 Edge Intelligence . . . . . . . . . . . . . . . . . . . 67 3.2.2 Federated Learning . . . . . . . . . . . . . . . . . . 69 3.2.3 Model Compression . . . . . . . . . . . . . . . . . 70 3.2.4 Beyond the State of the Art . . . . . . . . . . . . . . 73 3.3 UseCase ........................... 73 3.4 Architecture Proposal . . . . . . . . . . . . . . . . . . . . . 75 3.4.1 Assumption...................... 75 3.4.2 Cluster Aggregator . . . . . . . . . . . . . . . . . . 75 3.4.3 Cloud Components . . . . . . . . . . . . . . . . . . 78 3.4.4 Distributed Agent . . . . . . . . . . . . . . . . . . . 81
Contents xv 16.3.1 Layer 1: Automation . . . . . . . . . . . . . . . . . 314 16.3.2 Layer 2: Self-awareness . . . . . . . . . . . . . . . 316 16.3.3 Layer 3: High-level communication and coordination ..................... 317 16.3.4 Layer 4: System goals . . . . . . . . . . . . . . . . 318 16.4CaseStudy .......................... 319 16.4.1 Vertical farming module . . . . . . . . . . . . . . . 319 16.4.2 HVAC system of a cruise ship . . . . . . . . . . . . 320 16.5 Discussion and Conclusion . . . . . . . . . . . . . . . . . . 321 17 Neuromorphic IoT Architecture for Efficient Water Management 325 Mugdim Bublin, Heimo Hirner, Antoine-Martin Lanners, and Radu Grosu 17.1 Introduction and Background . . . . . . . . . . . . . . . . . 326 17.2 Neuromorphic IoT Architecture . . . . . . . . . . . . . . . 327 17.2.1 Design principles . . . . . . . . . . . . . . . . . . . 327 17.2.2 Hierarchical distributed control and learning . . . . 327 17.3 Free Energy Principle . . . . . . . . . . . . . . . . . . . . . 329 17.4 Asynchronous Processing and Event-driven Communication........................ 330 17.5 The Role of Thresholds in Hierarchical IoT Model . . . . . 331 17.5.1 Setting adaptive thresholds using prediction errors ......................... 331 17.5.2 Incorporating actions into threshold setting . . . . . 332 17.5.3 Optimizing thresholds using free energy minimization ..................... 333 17.5.4 Threshold setting summary . . . . . . . . . . . . . . 333 17.6Implementation........................ 333 17.7 Case Study: Smart Village Water Management . . . . . . . . 335 17.7.1 Context and objectives . . . . . . . . . . . . . . . . 335 17.7.2 Data collection and preprocessing . . . . . . . . . . 336 17.7.3 Prediction models and performance . . . . . . . . . 337 17.7.4 Anomaly detection . . . . . . . . . . . . . . . . . . 338 17.8Discussion........................... 339 17.8.1 Energy efficiency and communication overhead . . . 339 17.8.2 System responsiveness and latency . . . . . . . . . . 339 17.8.3 Safety & security . . . . . . . . . . . . . . . . . . . 339 17.8.4 Practical implications . . . . . . . . . . . . . . . . . 339
xvi Contents 17.9Conclusion .......................... 340 17.10FutureWork ......................... 340 18 Online AI Benchmarking on Remote Board Farms 343 Maïck Huguenin, Baptiste Dupertuis, Robin Frund, Margaux Divernois, and Nuria Pazos 18.1 Introduction and Novelty Aspect . . . . . . . . . . . . . . . 344 18.2 State-of-the-art . . . . . . . . . . . . . . . . . . . . . . . . 345 18.3 dAIEdge-VLab Architecture . . . . . . . . . . . . . . . . . 347 18.4 dAIEdge-VLab Implementation . . . . . . . . . . . . . . . 352 18.5Conclusion .......................... 357 19 Optimising Neural Networks for Water Stress Prediction in Europe: A Sustainable Approach 361 Laura Sanz-Martín, Manal Jammal, and Javier Parra-Domínguez 19.1 Introduction and Background . . . . . . . . . . . . . . . . . 362 19.2StateoftheArt ........................ 364 19.3 Material and Methods . . . . . . . . . . . . . . . . . . . . . 365 19.3.1 Data.......................... 365 19.3.2 Methodology . . . . . . . . . . . . . . . . . . . . . 368 19.3.2.1 Data .................... 368 19.3.2.2 Neural network architecture . . . . . . . . 368 19.3.2.3 Model optimization . . . . . . . . . . . . 369 19.3.2.4 Evaluation and metrics . . . . . . . . . . 372 19.4Results............................. 373 19.5Conclusions.......................... 376 20 The Accountability Strikes Back: Decentralizing the Key Generation in CL-PKC with Traceable Ring Signatures 379 Varesh Mishra, Aysajan Abidin, and Bart Preneel 20.1Introduction.......................... 380 20.2 Related Work and Contributions . . . . . . . . . . . . . . . 380 20.3Preliminaries ......................... 381 20.3.1 Pairing Based Cryptography . . . . . . . . . . . . . 381 20.3.2 Certificateless Public Key Cryptography (CL-PKC) . 382 20.3.3 Traceable Ring Signatures . . . . . . . . . . . . . . 382 20.3.4 Merkle Patricia Trie . . . . . . . . . . . . . . . . . 383 20.4ProposedModel........................ 383 20.4.1 Notation ....................... 383
Contents xvii 20.4.2 Network Architecture . . . . . . . . . . . . . . . . . 384 20.4.3 Protocol........................ 386 20.4.3.1 Network Bootstrapping . . . . . . . . . . 386 20.4.3.2 Key Generation . . . . . . . . . . . . . . 387 20.4.3.3 Full Private Key Generation . . . . . . . . 389 20.4.3.4 Update ST ................. 390 20.4.3.5 Audit Algorithm . . . . . . . . . . . . . . 391 20.5 Empirical Results and Analysis . . . . . . . . . . . . . . . . 396 20.6 Conclusions and Future Works . . . . . . . . . . . . . . . . 398 Index 403 About the Editors 409
Preface A New Edge AI Reality This book is the result of the rich exchanges of ideas and presentations at the European Conference on EDGE AI Technologies and Applications (EEAI) held on 21-23 October 2024 in Cagliari, Sardinia, Italy, offering a panoramic snapshot and a technical deep dive into the contemporary landscape of edge AI. With twenty selected chapters, it encapsulates the convergence of fundamental concepts, technical advancements, and real-world deployments that define the edge AI continuum. Collectively, the book serves as a reference for the field, capturing the current state-of-the-art and anticipating future trends in hyperautomation, generative AI, connectivity, autonomy, and security mesh architectures. Whether you are seeking in-depth technical knowledge, inspiration for novel applications, or a strategic overview of the edge AI landscape, you will find invaluable insights from thought researchers and practitioners at the forefront of the field of edge AI. A brief overview of each of the twenty chapters is provided below, highlighting the research and applications of edge AI that underscore the book’s commitment to both technological and societal impact. Edge AI Systems Verification and Validation: This chapter explores the challenges of verifying and validating complex edge AI systems, which integrate hardware, software, and data. It proposes a structured framework that combines modeland data-driven engineering to ensure these systems are reliable, robust, and meet regulatory standards. Pioneering the Hybridization of Federated Learning: This work introduces a hybrid federated learning framework for human activity recognition, where some clients agree to share a portion of their data. The research assesses whether this partial data sharing can improve the overall classification accuracy of the collective model while maintaining user privacy. xix
xx Preface Edge Intelligence Architecture for Distributed and Federated Learning: This chapter proposes a novel architecture for monitoring Electric Vehicles (EVs) by combining Federated Learning, Knowledge Distillation, and model compression. This approach enables the creation of efficient, privacy-preserving AI models that can be deployed on resource-constrained edge devices for applications like predictive maintenance. Challenges and Performance of SLAM Algorithms on Resource-Constrained Devices: This study evaluates the performance of various visual-based SLAM (Simultaneous Localisation and Mapping) algorithms on resourceconstrained hardware, such as the NVIDIA Jetson. It benchmarks several deep learning-based systems on metrics such as accuracy, energy consumption, and resource usage to assess their real-world viability. Designing Accelerated Edge AI Systems with Model-Based Methodology: This chapter presents a Model-Based Cybertronic System Engineering (MBCSE) methodology for designing optimal edge AI systems with bespoke hardware accelerators. This approach enables a holistic analysis that balances performance, power, and cost, ensuring AI algorithms can be deployed effectively within tight system constraints. Edge AI Acceleration for Critical Systems: Focusing on the demanding environment of satellites, this work discusses hardware solutions, such as FPGAs and CGRAs, for real-time, autonomous AI processing. The research addresses critical system challenges, including power constraints and radiation tolerance, and details the design of an FPGA-based GPU and an AI accelerator framework. Model Selection and Prompting Strategies for LLM-Based Robotic Systems: This chapter examines the challenges of selecting and implementing Large Language Models (LLMs) in resource-constrained robotic systems. It highlights that changing model weights or precision often requires significant modifications to prompting strategies, complicating the development of modular, weight-agnostic systems. Optimising ViT for Edge Deployment: This research presents a hybrid token reduction method, combining token merging and pruning, to make Vision Transformers (ViT) more efficient for semantic segmentation on edge devices. This approach significantly reduces computational complexity with only a minimal drop in accuracy, though it highlights challenges in exporting pruned models.
Preface xxi Recent Trends in Edge AI: This chapter provides a comprehensive overview of recent techniques for efficiently designing, training, and deploying machine learning models on edge devices. It covers scalable architectures, neural architecture search, and compression methods, such as quantisation and pruning, to enable energy-efficient AI in resource-limited environments. Scalable Sensor Fusion for Motion Localization in Large RF Sensing Networks: This work addresses the challenge of accurate motion localisation in large-scale wireless sensing networks by using a probabilistic model. It demonstrates that variational Bayesian techniques offer a scalable solution for sensor fusion, enabling localised updates that model non-local effects efficiently. Multi-Step Object Re-Identification on Edge Devices: This chapter proposes a pipeline for vehicle re-identification on edge devices using a multi-step feature extraction and matching process. The system detects an object, converts it to a vector embedding, and queries a database to find matches, achieving high precision in real-world camera network scenarios. A TinyMLOps Framework for Real-World Applications: This work introduces a TinyMLOps framework to streamline the optimisation and deployment of AI models on microcontrollers. The framework uses cloud resources for intensive tasks while gathering real-time performance metrics from target devices, ensuring an accurate and scalable solution for deploying AI in constrained environments. Transfer and Self-Learning in Probabilistic Models: This chapter explores the integration of transfer-learning and self-learning techniques within a single probabilistic model. The research finds that this synergy can be achieved through prior optimisation, enabling models to adapt across different environments where they are deployed. A Novel Hierarchical Approach for On-Device Energy Efficient Fault Classification: This work proposes a hierarchical architecture utilising multiple smaller neural networks to perform energy-efficient fault classification directly on edge devices. By dividing the problem into smaller sub-tasks, the approach achieves a nine-fold reduction in energy consumption with comparable accuracy to a non-hierarchical model. Discovering and Classifying Defects at the Edge: This chapter presents an AIbased optical inspection solution for detecting defects in digital and wooden industry products. Using YOLO and ResNet models deployed on edge
xxii Preface devices, the system achieves high accuracy in identifying defect positions and classifying defect types, with explainability tools clarifying the model’s decisions. Conscious Agents Interaction Framework for Industrial Automation: This paper examines the integration of human cognitive models into industrial automation, aiming to create flexible, multi-agent systems where humans and machines collaborate as equal partners. Case studies in vertical farming and HVAC control demonstrate how agents can reason and negotiate to achieve both collective and individual goals. Neuromorphic IoT Architecture for Efficient Water Management: This work proposes a neuromorphic IoT architecture inspired by biological systems to address the energy and communication challenges of traditional IoT networks. A case study on water management demonstrates how this eventdriven, asynchronous approach can be realised with neuromorphic hardware to create a more efficient and responsive system. Online AI Benchmarking on Remote Board Farms: This project aims to create a collaborative platform, dAIEdge - VLab, that enables researchers to benchmark AI models on a range of remote edge devices. This virtual laboratory will provide access to shared resources and tools, enabling users without deep-embedded expertise to conduct live AI experiments. Optimising Neural Networks for Water Stress Prediction in Europe: This study compares various neural network architectures and optimisers to predict water stress, a key sustainability indicator accurately. The findings show that a three-layer architecture with an Adam optimiser provides the highest accuracy, offering a valuable tool for informed water resource management. Decentralising Key Generation in CL-PKC with Traceable Ring Signatures: This chapter addresses a key vulnerability in Federated Learning by proposing a mechanism to decentralise key generation in Certificateless Public Key Cryptography. Using traceable ring signatures and blockchain infrastructure, the model provides accountability and disincentivises malicious behaviour among trusted authorities.
List of Figures Figure 1.1 Edge AI advantages. . . . . . . . . . . . . . . . . 3 Figure 1.2 Edge AI verification and validation process. . . . . 7 Figure 1.3 Edge AI dependability – Trustworthiness. . . . . . 11 Figure 1.4 Edge AI dependability - Trustworthiness extended properties....................... 12 Figure 1.5 Verification and validation. . . . . . . . . . . . . . 17 Figure 1.6 Verification and validation framework. . . . . . . . 19 Figure 1.7 Edge AI system W-Model (adapted from [92]) . . . 28 Figure 2.1 Typical Federated Learning Architecture . . . . . . 56 Figure 2.2 Vertical Hybrid Federated Learning Architecture . . 57 Figure 2.3 Horizontal Hybrid Federated Learning Architecture..................... 57 Figure 2.4 Implementation of the Vertical Hybridization in Flower........................ 59 Figure 2.5 Horizontal Hybridization Results for UCI HAR . . 60 Figure 2.6 Vertical Hybridization Results for UCI HAR . . . . 61 Figure 2.7 Horizontal Hybridization Results for FEMNIST . . 61 Figure 2.8 Vertical Hybridization Results for FEMNIST . . . 62 Figure 3.1 Six-level rating for EI described in [28]. . . . . . . 68 Figure 3.2 In a Federated Learning scenario, each client trains its model leveraging its own private data and sends its model parameters to a central server. The central server aggregates the parameters received from each client to enhance the performance of the central global model, which is then sent back to the clients......................... 70 Figure 3.3 The schema illustrates the fundamental concept of KD: during the training of a simplified neural network, knowledge from a larger network is transferred to the smaller one. . . . . . . . . . . . . 72 Figure 3.4 Use case scenario. . . . . . . . . . . . . . . . . . . 74 xxiii
xxiv List of Figures Figure 3.5 Cluster Aggregator schema designed to handle FL central aggregator tasks and to implement a distillation framework adaptable during the training process. ....................... 76 Figure 3.6 Software components deployed in the Cloud. . . . 79 Figure 3.7 The final architecture includes a Cluster Aggregator, deployed in the Cloud, and Distributed Agents, deployed on resource-constrained edge devices. . . 82 Figure 4.1 Overview of RDS-SLAM [1]. . . . . . . . . . . . . 94 Figure 4.2 Overview of VDO-SLAM [2]. . . . . . . . . . . . 95 Figure 4.3 Trajectory predictions: each color denotes a different tested system. . . . . . . . . . . . . . . . . . . 98 Figure 4.4 SLAM Block Execution Time Breakdown. Left: VDO-SLAM; Right: RDS-SLAM. . . . . . . . . . 100 Figure 5.1 Model-Based Cybertronics Systems Engineering Methodology .................... 115 Figure 5.2 MBCSE Process for AI system design . . . . . . . 116 Figure 5.3 Performance exploration . . . . . . . . . . . . . . 117 Figure 5.4 Alternative micro-architectures from a single source, based on tools settings and constraints . . . . . . . 120 Figure 6.1 FPG-AI block diagram. . . . . . . . . . . . . . . . 132 Figure 6.2 Overview of the System-on-Chip based on GPU@SAT...................... 136 Figure 6.3 Overview of the GPU@SAT architecture. . . . . . 137 Figure 6.4 CGR-AI Engine block diagram. . . . . . . . . . . 140 Figure 7.1 Scematic of the MMS first demonstrator’s HLP system. ....................... 153 Figure 7.2 Correctly passed tests for quantization precision comparison...................... 157 Figure 7.3 Correctly passed tests by various LLMs. . . . . . . 158 Figure 7.4 Planning success rates for the various models. . . . 159 Figure 7.5 VRAM usage of the tested models. . . . . . . . . . 160 Figure 7.6 Testing results relative to model performance on the BFCL......................... 160 Figure 8.1 Outline of the Proposed Hybrid Token Optimization Technique....................... 171 Figure 8.2 Results of Patch merging: grouped patches in blue, individual patches in red. . . . . . . . . . . . . . . 173
List of Contributors Abidin, Aysajan, COSIC, KU Leuven, Belgium Antonelli, Fabio, Fondazione Bruno Kessler, Italy Antonini, Mattia, Fondazione Bruno Kessler, Italy Antonio Coppola, Marcello, STMicroelectronics, France Arents, Janis, Institute of Electronics and Computer Science (EDI), Latvia Bahr, Roy, SINTEF AS, Norway Bocchi, Tommaso, University of Pisa, Italy Bublin, Mugdim, University of Applied Science FH Campus Wien, Austria Bureka, Anzelika, Institute of Electronics and Computer Science (EDI), Latvia Cancelliere, Francesco, University of Catania, Italy Carnevale, Lorenzo, University of Messina, Italy Ciabattini, Leonardo, University of Bologna, Italy Dell’Acqua, Pierluigi, University of Messina, Italy Deutel, Mark, Friedrich-Alexander-Universität Erlangen-Nürnberg, Germany Di Felice, Marco, University of Bologna, Advanced Research Center for Electronic Systems, Italy Divernois, Margaux, Haute Ecole Arc – HES-SO, Switzerland Dupertuis, Baptiste, Haute Ecole Arc – HES-SO, Switzerland Eduards Zinars, Toms, Institute of Electronics and Computer Science (EDI), Latvia Esposito, Alfonso, University of Bologna, Italy xxxi
xxxii List of Contributors Fanucci, Luca, University of Pisa, Italy Faro, Robin, Deepsensing SRL, Italy Frund, Robin, Haute Ecole Arc – HES-SO, Switzerland Galagain, Calvin, Université Paris-Saclay, CEA-List, ENSTA Paris, France Goulette, François, ENSTA Paris, France Greitans, Modris, Institute of Electronics and Computer Science (EDI), Latvia Grosu, Radu, Technische Universität Wien, Austria Haroun, Karim, Université Paris-Saclay, CEA-List, Université Côte d’Azur, France Hirner, Heimo, University of Applied Science FH Campus Wien, Austria Huguenin, Maïck, Haute Ecole Arc – HES-SO, Switzerland Jammal, Manal, IoT Digital Innovation Hub, Spain Judvaitis, Janis, Institute of Electronics and Computer Science (EDI), Latvia Klein, Russell, Siemens EDA, USA Lanners, Antoine-Martin, University of Applied Science FH Campus Wien, Austria Mallah, Maen, Fraunhofer IIS, Fraunhofer Institute for Integrated Circuits, Germany Mishra, Varesh, COSIC, KU Leuven, Belgium Moghbelan, Yasamin, University of Bologna, Italy Monopoli, Matteo, University of Pisa, Italy Montori, Federico, University of Bologna, Advanced Research Center for Electronic Systems, Italy Nannipieri, Pietro, University of Pisa, Italy Ovsiannikova, Polina, Aalto University, Finland Pacini, Tommaso, University of Pisa, Italy Pagani, Alain, German Research Center for Artificial Intelligence (DFKI), Germany
List of Contributors xxxiii Parra-Domínguez, Javier, IoT Digital Innovation Hub, Spain Pazos, Nuria, Haute Ecole Arc – HES-SO, Switzerland Pijlman, Fetze, Signify, Eindhoven University of Technology, The Netherlands Poreba, Martyna, Université Paris-Saclay, CEA-List, France Preneel, Bart, COSIC, KU Leuven, Belgium Proust, Mathilde, Université Paris-Saclay, CEA-List, France Racinskis, Peteris, Institute of Electronics and Computer Science (EDI), Latvia Sanz-Martín, Laura, University of Salamanca, Spain Scheele, Stephan, Ostbayerische Technische Hochschule Regensburg, Germany Solanti, Petri, Siemens EDA, Germany Strano, Alessandro, Deepsensing SRL, Italy Szczepanski, Michal, Université Paris-Saclay, CEA-List, France Urlini, Giulio, STMicroelectronics, Italy Vashishth, Devesh, University of Applied Sciences, Heilbronn, Germany Vecchio, Massimo, Fondazione Bruno Kessler, Italy Vermesan, Ovidiu, SINTEF AS, Norway Villari, Massimo, University of Messina, Italy Vismanis, Oskars, Institute of Electronics and Computer Science (EDI), Latvia Vyatkin, Valeriy, Aalto University, Finland; Luleå Tekniska Universitet, Sweden Wagner, Marco, University of Applied Sciences, Heilbronn, Germany Wissing, Julio, Fraunhofer IIS, Fraunhofer Institute for Integrated Circuits, Germany Zulberti, Luca, University of Pisa, Italy Zutis, Tomass, Institute of Electronics and Computer Science (EDI), Latvia Zyrianoff, Ivan, University of Bologna, Italy
List of Abbreviations AHU Air handling unit AI Artificial Intelligence ANN Artificial Neural Network ASIC Application-Specific Integrated Circuit AXI Advanced eXtensible Interface BDI Belief-desire-intention CGRA Coarse-Grained Reconfigurable Array CNN Convolutional Neural Network COTS Commercial Off-The-Shelf CPU Central Processing Unit CPU Central Processing Unit CTS Content-aware Token Sharing DL Deep Learning DMA Direct Memory Access DNN Deep Neural Network DSE Design Space Exploration DSP Digital Signal Processing DToP Dynamic Token Pruning EI Edge Intelligence ESA European Space Agency EV Electric Vehicle FL Federated Learning FP Floating-point FPGA Field Programmable Gate Array FPS Frames Per Second FU Functional Unit FXP Fixed-point GB Gigabyte GEO Geosynchronous Earth Orbit GPGPU General-Purpose Computing on Graphic Processing Units xxxv
xxxvi List of Abbreviations GPU Graphics Processing Unit HCI Human-computer interaction HDL Hardware Description Language HMI Human-machine interface HVAC Heating, ventilation, and air conditioning IoT Internet of Things KD Knowledge Distillation LEO Low Earth Orbit LIB LI-ion Battery LUT Look Up Table MAC Multiply And Accumulation MAS Multiagent systems MCU Microcontroller Unit MDE Modular Deep Learning Engine MES Manufacturing execution system mIoU Mean Intersection Over Union ML Machine Learning MM Memory-Mapped mmseg MMSegmentation, an open-source semantic segmentation toolbox MPU Microprocessor Unit NN Neural Network NPU Neural Processing Unit OTA Over-the-air PA Power/AreaRH Radiation-Hardened RBf Radial Basis Functions RHBD Radiation-Hardened by Design RL Reinforcement learning RNN Recurrent Neural Network RT Radiation-Tolerant SAN Stochastic Activity Network SEL Single Event Latchup SEU Single Event Upset SLAM Simultaneous Localization and Mapping SoC System-on-Chip SoC State of Charge SoH State of Health SVD Singular Value Decomposition TCM Tightly-Coupled Memory
List of Abbreviations xxxvii TMR Triple Modular Redundancy VIO Visual-Inertial Odometry ViT Vision Transformer VO Visual Odometry VPU Vision Processing Unit VRAM video random-access memory
1 Edge AI Systems Verification and Validation Ovidiu Vermesan1, Alain Pagani2, Roy Bahr1, Marcello Antonio Coppola3, and Giulio Urlini4 1SINTEF AS, Norway 2German Research Center for Artificial Intelligence (DFKI), Germany 3STMicroelectronics, France 4STMicroelectronics, Italy Abstract The integration of edge artificial intelligence (AI) into different complex systems presents unique challenges, particularly concerning their reliability, robustness, safety, and transparency. Edge AI systems must function as intended and meet regulatory and technical standards. Traditional verification and validation (V&V) methodologies, which are well-suited for conventional software (SW) and hardware (HW) systems, do not fully address the unique characteristics of edge AI-based systems that include hardware, software, elements of edge AI technology stack and data. The chapter delves into the challenges and methodologies for edge AI verification and validation to identify the unique elements required to develop verifiable edge AI systems based on a structured verification and validation framework integrated with modeland data-driven engineering principles, assurance cases, and domain-specific requirements. It highlights the terminology and concepts for edge AI as a technology that integrates HW, SW, and edge AI technology and data while presenting the challenges of the convergence of these technologies in developing verification and validation solutions. Keywords: edge AI, edge AI system, verification, validation, machine learning, deep learning, AI agents, agentic AI, system engineering, small language models. 1
2Edge AI Systems Verification and Validation 1.1 Introduction and Background Edge AI has become a cornerstone of innovation in various industries, driving advancements in automation, decision-making, and predictive analysis. Edge AI systems applying machine learning (ML), deep learning (DL), and data processing at the edge involving deep neural networks (DNN) present significant challenges for ensuring the reliability, safety, and effectiveness of intelligent embedded devices across the edge AI computing continuum, ranging from microto deepand meta-edge. Edge AI can be either deterministic or non-deterministic, based on the typical application and design choices involved. Many edge AI applications prioritise real-time, deterministic behaviour for critical tasks, such as control algorithms. Other applications can leverage the non-deterministic nature of AI to deliver more adaptable and creative solutions as the non-deterministic nature of edge AI means it can offer different interpretations based on context. In real-time applications, edge AI systems require precise timing and consistent response times. This is demanded for tasks where milliseconds of delay can be critical. Deterministic edge AI is appropriate for applications that demand predictability and consistency, while non-deterministic approaches are advantageous for applications that require adaptability, creativity, and continuous learning. The choice between using a deterministic or non-deterministic approach finally depends on the detailed requirements of the application and the expected trade-offs among predictability, adaptability, and computational cost. The advancement of edge AI technologies and the ubiquity of automated AI-based tools have created complex operational environments. Edge AI systems are evolving towards engineering advanced adaptive systems and require new concepts for verification and validation to address the challenging multidimensional integration of HW, SW, AI models, algorithms, datasets, and the multimodality of data. The advantages of leveraging edge AI in many industrial applications include real-time processing, enhanced privacy and data security, reduced latency, optimised bandwidth, reliability, and scalability, as illustrated in Figure 1.1. Edge AI technology stack combines AI and IoT with edge computing, allowing data processing and edge AI algorithm execution to occur directly on devices located at the edge of the network. By bringing AI closer to the source of data generation, edge AI enables more efficient and responsive decision-making across a wide range of applications. AI systems, particularly those based on machine learning (ML), pose unique challenges that differ from traditional software. Unlike conventional
1.2 Foundational Concepts and Edge AI Verification and Validation Taxonomy 9 In this context, the principal verification process involves several methodical steps: • Requirement Analysis: Clearly define and document edge AI systems’ functional, performance, and safety requirements. • Verification Planning: Establishing a structured plan that details verification strategies, methods, criteria, and resources. • Model and Code Inspection: Applying manual or automated inspections and formal verification techniques to analyse AI model structures and implementation code for correctness. • Test Development: Generating extensive and varied test cases covering all possible usage scenarios, operational environments, and stress conditions. • Verification Execution: Systematically conducting tests and verification activities, rigorously analysing outcomes against specified acceptance criteria. • Reporting and Review: Documenting detailed verification outcomes, identifying discrepancies, and facilitating stakeholder review to ensure comprehensive verification coverage. • Iterative Refinement: Addressing identified issues through iterative model adjustments, re-verification cycles, and continual improvement to achieve specified verification goals. Validation refers to the set of the activities that ensure that the edge AI system that has been built is traceable to the requirements and the right edge AI system is built to meet user needs. Validation is the process of checking whether the edge AI system is up to the mark or, in other words, if the product has high-level requirements. It is the process of checking the validation of the edge AI system, e.g., it checks if what we are developing is the right edge AI system. It is validation of the actual and expected edge AI systems. Validation is a form of dynamic testing. Validation means answering the question: are we building the right edge AI system? Validation of edge AI systems is a critical and systematic process intended to ensure that the developed AI system meets stakeholders’ and end-users’ specific needs and expectations, as explicitly outlined in ISO/IEC 22989 (Information Technology — Artificial Intelligence — Concepts and Terminology). According to ISO standards, validation involves confirming through objective evidence that the requirements for a specific intended use or application have been fulfilled. In AI, validation goes beyond verifying
10 Edge AI Systems Verification and Validation compliance with technical specifications—it assesses whether the system performs suitably in real-world conditions and scenarios. The principal elements involved in the validation of edge AI systems encompass several dimensions: • The identification of intended use and user requirements is foundational. Clear articulation and comprehensive understanding of user needs, operational contexts, and usage environments are paramount. This involves gathering input from stakeholders and end-users to form a robust basis for subsequent validation activities. • Operational scenario definition is critical. Edge AI systems must be validated within scenarios that accurately represent real-world operational contexts. Scenarios are typically derived from realistic usage conditions, including normal operational states, boundary conditions, and potential abnormal or edge cases. • Performance evaluation under realistic conditions is essential. Validation combines simulated environments and real-world testing to ensure edge AI systems perform reliably and effectively. Performance metrics, such as accuracy, precision, recall, robustness, resilience, and usability, form the basis for evaluating system performance and alignment with stakeholder expectations. • Human-machine interaction and usability assessment are integral to validation. Edge AI systems are validated to ensure effective and intuitive interactions with human operators or users. Usability testing, user experience assessments, and feedback loops with real users facilitate comprehensive evaluations of the AI system’s ease of use and accessibility. • Safety, security, and ethical considerations are central elements of the validation process. These assessments verify that edge AI systems function correctly and comply with safety standards, security protocols, data privacy laws, and ethical guidelines, aligning with international frameworks and societal expectations. The edge AI validation process typically involves structured, methodical steps: • Requirement and Expectation Definition: Establishing clear validation criteria and user expectations, documenting them rigorously. • Validation Planning: Creating detailed validation plans that specify methodologies, scenarios, test environments, and acceptance criteria.
1.2 Foundational Concepts and Edge AI Verification and Validation Taxonomy 11 • Scenario Development: Defining realistic operational scenarios and selecting representative use-cases and edge-cases for comprehensive validation. • Simulation and Real-world Testing: Controlled simulations are conducted, followed by real-world trials to evaluate AI system performance against established criteria. • Performance and Usability Assessment: Analysing performance outcomes, usability data, and user feedback to ascertain compliance with expectations and user requirements. • Safety, Security, and Ethical Evaluation: Systematically reviewing compliance with safety and security standards, data protection requirements, and ethical norms. • Reporting and Continuous Improvement: Compiling comprehensive validation reports, documenting findings and recommendations, and establishing iterative cycles for continuous system refinement. Verification and validation of AI and edge AI models and data are required in safety-critical applications to ensure the trustworthiness of edge AI-enabled systems (e.g., reliability, availability, maintainability, safety, security, resilience, connectability, explainability, interpretability, transparency, etc.) as illustrated in Figure 1.3 and Figure 1.4. Dependable edge AI systems involve using systems and software engineering principles to systematically guarantee dependability during the edge AI system’s construction, V&V, and operation and consider legal and normative requirements directly from the start. Figure 1.3 Edge AI dependability – Trustworthiness.
12 Edge AI Systems Verification and Validation Figure 1.4 Edge AI dependability - Trustworthiness extended properties. The progress made in developing standards and regulatory frameworks for AI and edge AI aims to ensure the responsible use of AI in various applications. The relevant standards for AI that can be applied to edge AI systems are ISO/IEC 42001 and ISO/IEC TR 24028:2020 that are described below. The ISO/IEC 42001 standard, a management system for AI, focuses on building trust and dependability in AI systems. It provides a framework to establish, implement, maintain, and continually improve their AI management systems, ensuring the responsible development and use of AI. The standard emphasises trustworthiness, fairness, transparency, and accountability in AI systems [43][44]. The ISO/IEC TR 24028:2020 standard addresses topics related to trustworthiness in AI systems, including approaches to establish trust in AI systems through transparency, explainability, controllability, etc.; engineering pitfalls and typical associated threats and risks to AI systems, along with possible mitigation techniques and methods; and approaches to assess and achieve availability, resiliency, reliability, accuracy, safety, security and privacy of AI systems [5]. Traditional V&V workflows, such as the V-model, are insufficient for ensuring the accuracy and reliability of AI and edge models. As a result, transformations of these workflows occurred to better serve edge AI applications. 1.2.1 Agentic AI and AI Agents The evolution of generative AI and the emergence of AI agents and agentic AI requires addressing them under the presentation of foundational concepts
1.2 Foundational Concepts and Edge AI Verification and Validation Taxonomy 13 and edge AI verification and validation taxonomy by defining the concepts and their specific characteristics. AI Agents can be defined as autonomous software entities engineered for goal-directed task execution within bounded digital environments. These agents are characterised by their ability to perceive structured or unstructured inputs, to reason over contextual information, and to initiate actions toward achieving specific objectives. The main characteristics of AI and edge AI agents are autonomy, task specificity, reactivity and adaptability, which enable the agents to operate as modular, lightweight interfaces between pre-trained AI models and domain-specific pipelines and workflows. AI agents are the concrete instantiations of the agentic AI paradigm. An AI agent is a specific software or hardware entity that embodies the principles of agentic AI. It is a tangible system equipped with sensors to perceive its environment and effectors to act upon it. While agentic AI is the “what,” the AI agent is the “how”, the actual implementation that performs tasks, makes decisions, and interacts with external environments. Agentic AI systems describe a paradigm shift from isolated AI agents to collaborative, multi-agent ecosystems capable of decomposing and executing complex goals [21]. These systems typically consist of orchestrated or communicating agents that interact via tools, APIs, and shared environments [23][14]. A key distinction between agentic AI and AI agents lies in their level of abstraction, as the agentic AI is a conceptual framework, whereas an AI agent is a functional system. An analogy can be drawn between the theory of computation and a physical computer. One provides the theoretical foundation and a model of what is possible, while the other is the practical machine that executes computations based on that theory. Agentic AI reflects a broad paradigm in AI and edge AI centred on creating systems that can perceive their environment, reason about their observations, and act autonomously to achieve specific goals. It is the underlying philosophy and set of principles that guide the development of intelligent, goal-oriented systems. This concept emphasises proactivity, reactivity, and social ability, defining the potential for AI to operate as an independent actor rather than a passive tool. Agentic AI systems introduce internal orchestration mechanisms and multi-agent collaboration frameworks. Agentic AI extends the foundational architecture to support complex, distributed, and adaptive behaviours by integrating components such as specialised agents, persistent memory, orchestration and advanced reasoning and planning. Agentic AI introduces novel memory integration, communication
14 Edge AI Systems Verification and Validation paradigms, and decentralised control, paving the way for the next generation of adaptive workflow automation in autonomous systems, swarm robotics, and autonomous vehicles with scalable, adaptive intelligence. In robotics and automation, agentic AI enables collaborative behaviour in multi-robot systems. Each robot operates as a task-specialised agent, such as a picker, transporter, or mapper, while an orchestrator supervises and adapts workflows. These architectures rely on shared spatial memory, real-time sensor fusion, and inter-agent synchronisation for coordinated physical actions. Use cases include warehouse automation, drone-based orchard inspection, and robotic harvesting [25]. Verification and validation of edge AI systems, based on AI agents and agentic AI components, must focus on ensuring the correctness, reliability, and robustness of autonomous decision-making in highly dynamic and constrained environments, which requires validating that the AI agents consistently perform their intended functions correctly under varying external environment conditions, including unexpected scenarios and adversarial inputs. Due to the limited computational resources typical of edge devices, V&V must also confirm that the system meets stringent real-time performance requirements, ensuring timely responses to critical events despite hardware and network limitations. Another aspect is assessing the resilience and safety of adaptive learning processes within these systems, particularly as they evolve in open environments. V&V efforts should capture how individual agents and collective multi-agent behaviours emerge and interact, verifying alignment with overall system objectives and preventing unsafe or unintended actions. Additionally, transparency and trustworthiness are key elements that enable human oversight, offering clear traceability and checkability of decisions made by autonomous components at the edge. 1.3 Defining Verification and Validation per Standard Several ISO standards offer consistent definitions for verification and validation, primarily within the context of quality management and systems/software engineering that can be applicable to AI and edge AI as presented below. ISO 9000:2015 (Quality management systems - Fundamentals and vocabulary) provides the definition for verification as the “confirmation,
1.3 Defining Verification and Validation per Standard 15 through the provision of objective evidence, that specified requirements have been fulfilled" [29]. The focus is on confirming that the system or component conforms to its design specifications and requirements [31]. It answers the question: “Did we build the product right?” [32]. Verification is often viewed as an internal process comparing the outputs of a development phase against the inputs [31]. Validation is defined as the “confirmation, through the provision of objective evidence, that the requirements for a specific intended use or application have been fulfilled” [29]. The focus shifts to ensuring the system meets the needs of the user and fulfils its intended purpose in the actual context of use. It answers the question: “Did we build the right product?” [32]. Validation often involves testing under real or simulated use conditions and considers stakeholder needs [31]. ISO 9001:2015 (Quality management systems - Requirements), states that validation activities ensure the resulting products/services meet requirements for the specified application or intended use. Validation often involves acceptance testing with end-users and assessing fitness for purpose, making it frequently an external process, whereas verification is more often internal. Both verification and validation are essential components of quality management and are necessary for ensuring a dependable system [30]. ISO/IEC/IEEE 15288:2015 (Systems and software engineering - System life cycle processes) standard integrates V&V into the system lifecycle and considers that the verification process has as purpose “to provide objective evidence that a system or system element fulfils its specified requirements and characteristics” [34]. It involves activities comparing the system or element against requirements, design descriptions, and other required characteristics, confirming it was “built right” [35], while the validation process has as purpose “to provide objective evidence that the system, when in use, fulfils its business or mission objectives and stakeholder requirements, achieving its intended use in its intended operational environment” [33]. This process confirms that stakeholder requirements are correctly defined, and that the system meets its intended purpose in the context where it will operate [33]. ISO/IEC 22989:2022 (Information technology - Artificial intelligence - Artificial intelligence concepts and terminology) AI-specific standard defines verification as “confirmation, through the provision of objective evidence, that specified requirements have been fulfilled,” noting it assures conformance to specification [9]. While not explicitly defining validation in the same way, it defines trustworthiness as the “ability to meet stakeholder
16 Edge AI Systems Verification and Validation expectations in a verifiable way” [38]. This definition links the core goal of validation (meeting stakeholder expectations/needs) directly to the concept of trustworthiness in AI. The standard also incorporates a “verification and validation” phase within its depiction of the AI system lifecycle [39]. ISO/IEC TR 24028:2020 (Information technology - Artificial intelligence - Overview of trustworthiness in Artificial Intelligence) technical report further reinforces the link between validation and trustworthiness and defines trustworthiness as the “ability to meet stakeholder expectations in a verifiable way” [40]. This aligns the concept of trustworthiness directly with the objective of validation – confirming that stakeholder needs and intended use requirements are met [42]. The report discusses assessing and achieving key characteristics like reliability, safety, security, and privacy, all crucial aspects evaluated during validation [41]. ISO/IEC 42001:2023 (AI Management System) standard specifies requirements for establishing, implementing, maintaining, and continually improving an AI Management System (AIMS) within an organization [43]. An AIMS provides a structured framework for responsible AI governance, risk management, and operational control throughout the AI lifecycle [44]. Verification activities are integral to an AIMS, supporting risk assessment, impact assessment, performance evaluation, and ensuring compliance with policies and objectives [44]. Notably, ISO/IEC 22989 (providing the core AI terminology) is a normative reference for ISO/IEC 42001, highlighting the foundational role of clear definitions [45]. IEEE 1012-2016 (IEEE Standard for System, Software, and Hardware Verification and Validation) standard applies to systems, software, and hardware being developed, maintained, or reused (legacy, commercial offthe-shelf [COTS], non-developmental items) [91]. The term “software” also includes firmware and microcode. Additionally, each of the terms “system,” “software,” and “hardware” encompasses documentation. V&V processes include the analysis, evaluation, review, inspection, assessment, and testing of products. V&V processes are used to determine whether the development products of a given activity conform to the requirements of that activity and whether the product satisfies its intended use and user needs. V&V lifecycle process requirements are specified for different integrity levels. The scope of V&V processes encompasses systems, software, hardware, and their interfaces.
1.4 Key Elements for Edge AI Verification and Validation 17 1.4 Key Elements for Edge AI Verification and Validation The elements for verifying and validating edge AI may encompass operational aspects, system integration, AI models and human-machine interaction. Verification and validation are important to ensure reliability, performance and accuracy of complex systems. Figure 1.5 Verification and validation. Edge AI system verification and validation refers to the processes and methodologies used to ensure that an edge AI system is dependable, performs as expected, and meets certain standards before it is deployed. These processes are crucial as the edge AI algorithms can work with highstakes decision-making, various sizes datasets, learn and evolve over time. The processes are needed for ensuring that the edge AI systems do what they are supposed to do, without unintended consequences, biases, or errors. AI and edge AI systems typically focus on the actual algorithms and models to ensure that they perform as intended under various conditions. In addition, edge AI systems focus on validating the systems performance on resource-constrained devices, network conditions and privacy in real-world scenarios. It is critical to distinguish verification from validation. While verification checks conformance to specifications (“Did we build the system right?”), validation confirms that the system meets the needs of the customer and other stakeholders and fulfils its intended purpose in its operational environment (“Did we build the right system?”) [30]. The introduction of AI in product and systems development has significantly increased the complexity of electronic components and systems (ECS), by integrating various technologies such as hardware, software, ML,
18 Edge AI Systems Verification and Validation DL, NNs, generative AI, and advanced data analytics. This complexity necessitates robust verification and validation frameworks and benchmarking to ensure these systems operate correctly and efficiently as illustrated in Figure 1.6. Complex edge AI models require verification and validation to ensure their predictions, decisions, and content generation outputs are reliable and accurate, which is critical for maintaining the trustworthiness of AI systems. Failures in edge AI-based ECS can have significant economic and business-critical consequences, including system failures, financial loss, and damage to infrastructure, making the dependability of edge AI systems paramount. In machine vision, specific verification concerns arise from the need to ensure reliable object detection, tracking, segmentation, or pose estimation across a wide range of dynamic conditions. For example, verification must confirm that visual inference results remain stable under varying lighting, occlusion, and motion blur, common challenges in edge deployments like factory floors or drones. Ensuring robustness and reproducibility in edge-based machine vision systems is inherently difficult due to the high variability and noise in visual data. Unlike structured tabular inputs, images and videos exhibit a vast range of intra-class variation—objects or actions belonging to the same class can appear drastically different depending on factors such as: • Lighting conditions (e.g., shadows, reflections). • Occlusions or partial views. • Background clutter. • Camera distortions, blur, or motion artifacts. • Variability in object shape, colour, texture, or viewpoint. A comprehensive V&V framework, presented in Figure 1.6, along with benchmarking of edge AI-based methods, frameworks, tools, and ECS, is essential to ensure performance and dependable system properties like security, reliability, robustness, and fairness. Verification ensures that edge AI-based methods, frameworks, tools, and electronic components and systems are built correctly and meet specifications, while validation confirms they perform as intended in real-world scenarios. In edge AI systems there is a need of creating a structured approach to defining and applying such a framework to edge AI-based tools and methods, ensuring ECS meet functional and non-functional requirements, quality, KPIs, and performance standards.
1.4 Key Elements for Edge AI Verification and Validation 25 emerging to help proactively identify, assess, and mitigate potential negative ethical and societal consequences before and during deployment [61]. A key focus is validating fairness and non-discrimination, moving beyond simple dataset metrics to assess the actual impact on different demographic groups in real-world deployment contexts [90]. This also involves considering broader societal implications related to employment, environmental sustainability, and the functioning of democratic processes [7]. 1.4.2.6 Data Quality and Suitability High-quality data ensures that models are trained effectively and can make accurate predictions in real-world scenarios [36]. As AI and edge AI systems become more complex and are deployed in diverse environments, the challenges associated with data quality and suitability have become increasingly significant. Considering the specific requirements for various AI and edge AI systems, challenges for data quality and suitability in AI and edge AI validation include: Relevance and Representativeness: the data used for training and validation is relevant and representative of the real-world environment in which the AI and edge AI systems operate. Data must reflect the diversity of conditions, contexts, and populations that the system will encounter. If the training data is biased or unrepresentative, the model’s performance may deteriorate when applied to actual situations. Volume and Availability: Considered very important, especially in scenarios where data may be generated at high velocity. Obtaining enough high-quality data for training and validation can be difficult. In many cases, developers may struggle to gather sufficient diverse data from edge devices, leading to models that are not well-trained for all possible situations they may encounter in deployment. Label Quality: important for supervised learning, as it directly impacts model accuracy. Inaccurate or inconsistent labelling can mislead the training process and result in poor performance in operational environments. Ensuring the reliability of labels, especially when data is labelled manually or derived from semi-automated processes, can be a significant extra work. Bias and Fairness: the biases in learning and training of data, can lead to outcomes that are unfair when models are deployed. AI and edge AI systems trained on biased data may perpetuate existing stereotypes or discriminate against certain classes and groups. Addressing data bias and ensuring fairness
26 Edge AI Systems Verification and Validation in model predictions is key to building trustworthy AI and edge systems that serve all stakeholders equitably. Data Drift: refers to the shifts in data distributions over time, which can degrade model performance. As the underlying data evolves, models may become less accurate or irrelevant. Ongoing monitoring and adaptation of models are necessary to mitigate the effects of data drift, making it a continuous challenge for AI and edge AI validation. Data Preprocessing: is a critical step in data management, particularly for edge AI systems with limited resources. Cleaning and transforming data into suitable formats can be challenging, when using with diverse data sources and formats. This preprocessing must be efficient to ensure real-time performance while maintaining data accuracy and integrity. Synthetic Data: helps augment training datasets and has several limitations. The effectiveness of synthetic data depends on its ability to mimic realworld scenarios accurately. If synthetic data does not accurately represent the complexities of real-world environments, it may lead to models that underperform when applied to actual data. Edge-Specific Challenges: these are related to data collection from distributed edge devices, considering elements like latency, bandwidth constraints, and intermittent connectivity, which can complicate the data validation process. Ensuring data quality in these scenarios requires innovative approaches to data management and model training. 1.5 The Edge AI Verification and Validation Lifecycle According to the OECD recommendation on artificial intelligence [10], an AI system is a machine-based framework that, driven by either explicit or implicit goals, deduces from the input it receives how to produce outputs such as predictions, content, recommendations, or decisions that may impact physical or virtual environments. The levels of autonomy and adaptability of different AI systems can vary after they are deployed. The lifecycle of an AI system generally encompasses multiple stages, which include planning and design; data collection and processing; model development and/or adaptation of existing models for specific tasks; testing, evaluation, verification, and validation; deployment for use; operation and monitoring; and retirement or decommissioning. These stages often occur iteratively and are not strictly
1.5 The Edge AI Verification and Validation Lifecycle 27 linear. The choice to retire an AI system can be made at any time during the operation and monitoring stage. AI and edge AI systems are distinct from other system types, which can influence the processes of the lifecycle model, such as [9]: • Most SW systems are designed to operate in exactly defined manners dictated by their requirements and specifications. In contrast, AI and edge AI systems that utilize ML rely on data-driven training and optimization techniques to address a wide range of inputs. • Traditional SW applications tend to be predictable, whereas this is less frequently true for AI and edge AI systems. • Additionally, traditional SW applications are generally verifiable, while evaluating the performance of AI and edge AI systems often necessitates statistical methods, making their verification more complex. • AI and edge AI systems usually require numerous iterations of enhancement to reach satisfactory performance levels. The edge AI development lifecycle outlines the stages involved in creating and operationalizing edge AI systems. It starts with problem definition, functional, non-functional requirements and data collection, followed by data preparation and feature engineering. Model selection and architecture design precede the training phase, where algorithms learn from the prepared dataset. Validation and testing ensure model performance and generalization. Iterative refinement optimizes the model based on results. Deployment integrates the AI system into production environments. Monitoring and maintenance track performance, address drift, and update the model as needed. Embedding AI and generative AI into the system design requires shifting from the current V-model, which addresses the HW and SW development cycle, to a W-model superimposed on the V-model to account for data and AI-specific artifacts. This includes the AI model development and data into the AI system’s lifecycle development, as illustrated in Figure 1.7 [92]. When superimposed on the traditional V-model used in HW and SW development, the W-model AI development lifecycle creates a comprehensive framework that addresses the distinct yet interrelated processes of AI data development and HW/SW engineering. This approach ensures that AI models and supporting systems are developed in a cohesive, iterative, and validated manner. This approach aligns existing tools and methods with AI technologies. The extension into the W-model structures represents the development workflow of AI systems comprising HW, SW, AI stack, and data components.
28 Edge AI Systems Verification and Validation Figure 1.7 Edge AI system W-Model (adapted from [92]) The inner W part of the model represents the AI-enabled processes and workflows integrated into the conventional model. The AI system Wmodel emphasises systematic validation and verification at each stage of the AI development process, helping ensure the robustness, reliability, and performance of AI systems. The W-model addresses the specific design and development requirements of edge AI systems, distinguishing them from traditional HW/SW and computing paradigms. The novelty in the model lies in the fact that the data required for development and AI, ML/DL, and generative AI model training is integrated into the development cycle, superimposed on the traditional V-model, and follows the algorithm selection and training in each lifecycle stage. The edge AI system W-model emphasises that AI and generative AI are integral to the lifecycle development processes of any AI-based product or service. As presented in the AI system W-model at the start of the development lifecycle, developers can utilise AI and generative AI to understand domain requirements and design architecture. The design captures both functional and non-functional requirements for embedded computing systems, such as those in automotive control or industrial units, considering hardware constraints and real-time performance needs. Challenges concerning edge AI requirements and AI requirements engineering are extensive and due in part to the practice by some to treat the AI element as a “black box”. Formal specification has been attempted and has proven to be difficult for tasks that are hard to formalise, requiring decisions on the use of quantitative
1.6 Failure Case Behaviour in Edge-based Machine Vision Systems 29 or Boolean specifications, as well as the incorporation of data and formal requirements. The challenge is to design effective methods for specifying both desired and undesired properties of systems that utilise AIor ML-based components. When considering the broader principles of agentic AI and AI agents, their development must be integrated into the edge AI system W-model as a specific part of the lifecycle development processes of any AI-based product or service. As a result, the agentic AI and AI agents V&V extends beyond the individual agent, focusing on validating the system’s autonomy and the ability to achieve long-term goals without unexpected consequences by assessing the alignment of the AI agent’s goals with the overall objectives of the edge AI system and ensuring that its learning and adaptation mechanisms do not lead to unsafe or undesirable states over time. 1.6 Failure Case Behaviour in Edge-based Machine Vision Systems In the context of edge-based machine vision systems, the study of failure case behaviour is a critical component of any robust verification and validation framework. These systems are increasingly deployed in real-world, safetycritical environments, ranging from autonomous vehicles and industrial robotics to surveillance and medical diagnostics, where failures can result in substantial consequences. While conventional validation focuses on averagecase performance metrics such as accuracy or mean average precision, these metrics often obscure rare but consequential failure modes. An edge AI system that performs well under ideal conditions may fail unexpectedly in the presence of visual distortions, environmental variability, or edge hardware constraints. Failures in machine vision models frequently arise in conditions that deviate from the data distribution seen during training. Examples include poor lighting, motion blur, occlusions, scale variation, or visual clutter. In edge deployments, such conditions are not only likely but expected, and the consequences of misclassification or missed detection can be severe. Furthermore, edge systems often operate with limited fallback options, and they must respond in real time, leaving little margin for error recovery. Understanding and characterizing these failure scenarios is therefore essential for both safety assurance and iterative model improvement. One important strategy for investigating failure modes involves deliberate stress testing through visual perturbations. By applying controlled
30 Edge AI Systems Verification and Validation transformations—such as adding noise, blurring, shifting brightness, or introducing occlusions—it becomes possible to evaluate how resilient a vision model is to real-world distortions. These tests often reveal brittle model behaviours that are not apparent during standard validation. In addition, targeted scenario-based testing using simulation tools or recorded video sequences enables systematic exploration of edge cases. This is especially valuable for applications involving dynamic environments, such as autonomous navigation, where rare events (e.g., unexpected pedestrian appearance or sensor occlusion) may not be captured adequately in available datasets. Scenario replay or simulation also supports reproducibility of observed failures, which is often a challenge in field deployments. Another aspect of failure case analysis is the examination of uncertainty and confidence levels in model predictions. Machine vision systems may produce incorrect predictions with unjustified confidence, especially when faced with unfamiliar or out-of-distribution inputs. Monitoring softmax confidence, prediction entropy, or Bayesian uncertainty estimates can help identify instances where the model is likely to fail. These signals may be used in runtime monitoring or to trigger fail-safe mechanisms. Post-hoc explainability methods, such as saliency maps or activation heatmaps, also play an important role in understanding failure behaviour. By visualising the regions of an input image that contributed most to a model’s prediction, one can diagnose whether a failure was due to the model focusing on irrelevant or misleading features. This insight often reveals underlying dataset biases or spurious correlations that were inadvertently learned. By combining scenario condition variables and the predictions as features in probabilistic frameworks (e.g., Bayesian networks) is also a method for modelling the uncertainty in predictions. Hardware-specific issues also need to be considered in failure analysis. For instance, the quantization of weights and activations required for execution on edge hardware (e.g., FPGAs or ASICs) can introduce numerical inaccuracies that degrade model performance in subtle ways. Testing the consistency between floating-point reference models and their hardware-deployed counterparts is essential to identify precision-induced errors. Similarly, real-time system profiling can reveal frame drops, synchronization mismatches, or input-output latency violations that lead to perceptual failures. Finally, insights gained from the analysis of failure cases should feed back into the design and development process. Difficult or misclassified examples can be incorporated into retraining pipelines, synthetic data can be generated
1.7 Research Challenges in Edge AI Verification and Validation 31 to increase robustness, and system architectures can be adapted to detect and respond to high-uncertainty inputs. Ideally, safety envelopes are defined at design time to formally capture the operational conditions under which the system is guaranteed to function correctly. This creates a closed-loop process that not only identifies but also mitigates and prevents known failure patterns. The systematic study of failure case behaviour is indispensable for building trustworthy machine vision systems on edge platforms. It enables developers to move beyond average-case performance toward comprehensive assurance of correctness, robustness, and safety under realistic and adverse conditions. 1.7 Research Challenges in Edge AI Verification and Validation Edge AI represents a cutting-edge computing approach that seeks to relocate the training and inference of ML models to the network’s edge [12]. However, implementing intelligence at the edge presents several significant challenges, such as the necessity to limit model architecture designs, ensuring the secure distribution and execution of the trained models, and managing the considerable network load needed to disseminate the models and the data gathered for training. Edge AI systems that incorporate continuous learning involve the gradual updating of the models within the systems during production and test runs operations [9]. The data input into a system during these operations, is not only evaluated to generate an output but is also concurrently utilized to modify the model, aiming to enhance it based on the production data. Depending on the design of the continuous learning system, certain human interventions may be necessary, such as data labelling, validating the application of specific incremental updates, or monitoring the performance of the edge AI system. Continuous learning can address the limitations of the initial training data and assist in managing data drift and concept drift, but it also presents significant challenges in ensuring the edge AI system operates correctly while learning. It is essential to verify the system in production and to capture the production data to be able to include them as part of the training dataset in future system updates. Catastrophic interference (and catastrophic unlearning) occurs when the training for new tasks disrupts the model’s comprehension of previous tasks [11]. As new information supersedes earlier learning, the model forfeits its
32 Edge AI Systems Verification and Validation capability to manage its initial tasks. Given the risk of catastrophic interference, continuous learning necessitates the capability to learn over time by integrating new observations from current data while preserving prior knowledge [9]. Numerous ML algorithms excel at learning tasks only when the data is provided in a single batch. As a model is trained on a specific task, its parameters are modified to effectively tackle that task. However, when new training data is introduced, the adjustments made for these new inputs can erase the knowledge the model had previously gained. In the context of neural networks, this occurrence is regarded as one of their key limitations. The combination of edge AI, IoT and Cyber-Physical Systems (CPS) marks a significant transformation in data processing by bringing it closer to the origin. This strategy minimizes latency, improves real-time decisionmaking, and lessens the load on centralized cloud resources. In CPS, control logic is utilized to process input from sensors, through actions of actuators and thus affecting processes occurring in the physical world [9]. This is particularly evident in robotics, where sensor data is directly employed to manage the robot’s operations and execute tasks in the physical world. Typically, robots are equipped with sensors at the edge to evaluate their current conditions, processors to facilitate control through analysis and action planning, and actuators to implement those actions. In contrast to industrial robots, which are consistently repeating the same trajectories and actions without deviations, service robots or collaborative robots must adapt to evolving situations and dynamic environments [9]. Programming this adaptability presents significant challenges due to the inherent variability. Components of edge AI systems can play a role in the control software and planning processes through the “Sense-Plan-Act” framework, allowing robots to modify their actions in response to obstacles or changes in the location of target objects. The integration of robotics and edge AI system components facilitates automated physical interactions with objects, environments, and individuals. In machine vision-based robotics, the visual processing pipeline itself must be verified and validated not only for accuracy but also for real-time responsiveness. Edge V&V must ensure that latency from image acquisition to action initiation does not exceed application-specific safety thresholds. Techniques like real-time trace logging and FPGA-based image path profiling can support this validation. The challenges and appropriate methodologies for AI verification are not uniform; they vary significantly depending on the type of AI model employed and the application domain’s risk profile.
1.7 Research Challenges in Edge AI Verification and Validation 33 Model-Specific Challenges and Verification Focus Deep Learning (DL) / Sub-Symbolic AI: •Challenges: The primary verification challenges stem from their inherent opacity (making internal logic inscrutable) [64], strong data dependency (performance tied to training data quality and representativeness) [76], difficulty in formal specification of complex learned behaviours [64], susceptibility to adversarial examples, and challenges in generalization beyond training data distributions [76]. Scalability of verification methods is a major bottleneck due to the vast number of parameters and high-dimensional inputs [52]. Non-determinism can also arise during training or inference [70]. •Verification Focus: Emphasis is placed on empirical performance evaluation using diverse test datasets, extensive robustness testing against perturbations and adversarial attacks, fairness audits to detect biases learned from data, applying explainable AI (XAI) techniques (like LIME, SHAP, saliency maps) and interpretable AI (IAI) to gain insights into model decisions [74][75] and, where feasible, formal verification of specific, localised properties such as robustness bounds around specific inputs [69]. Symbolic AI / Rule-Based Systems: •Challenges: These systems often suffer from brittleness, meaning they struggle to handle situations not explicitly covered by their predefined rules or knowledge base [73]. Creating and maintaining large, consistent, and complete knowledge bases can be labour-intensive and requires significant domain expertise [72]. They typically lack the ability to learn directly from raw, unstructured data. A computer vision model trained to detect stop signs may misclassify a slightly occluded or weathered sign because it hasn’t seen enough variation in training. Traditional software can also exhibit brittleness, i.e. they both struggle but in different forms. •Verification Focus: Verification centres on the logical integrity of the system. This includes checking the consistency of the rule set and knowledge base (absence of contradictions), analysing completeness (do the rules cover the intended domain?), formally verifying logical properties like soundness and validity of reasoning steps [64] and ensuring the traceability of outputs back to specific rules, which provides inherent explainability [72].
34 Edge AI Systems Verification and Validation Neuro-symbolic AI: •Challenges: This hybrid approach aims to combine the strengths of DL and symbolic AI but verifying the interaction and ensuring consistency between the neural (learning) and symbolic (reasoning) components is a key challenge [64]. Developing unified V&V frameworks that can handle both paradigms simultaneously is an active area of research [64]. •Verification Focus: Requires a multi-pronged approach: verifying the neural components using DL-specific techniques, verifying the symbolic components using logic-based methods, and crucially, verifying the interface and the correctness of the combined system’s behaviour. A major research direction involves leveraging the symbolic part to constrain, explain, or formally verify aspects of the neural part’s behaviour [64]. Domain-Specific Challenges and Verification Focus Safety-Critical Systems (e.g., Automotive, Aerospace, Medical, Industrial Control): •Requirements: These domains demand high levels of reliability, safety, robustness, and predictability [49]. System failures can have catastrophic consequences, including loss of life, severe injury, or significant environmental damage [49]. •Challenges: The need for provable guarantees clashes with the opacity and non-determinism of many AI components [65]. Meeting stringent regulatory standards (e.g., ISO 26262, IEC 62304, DO-178C) requires extensive evidence and documentation, which is difficult for AI/ML [68]. Managing the complexity of interaction with the physical world and ensuring safety across a vast range of operational scenarios is extremely challenging [71]. Exhaustive testing is typically infeasible due to the combinatorial explosion of possibilities [47]. Achieving deterministic replay for debugging and analysis is crucial but difficult [78]. •Verification Focus: Emphasis on rigorous methodologies, including formal methods where applicable, extensive simulation-based testing covering edge cases and failure modes, hardware-in-the-loop and realworld testing, fault tolerance analysis, adherence to domain-specific safety standards, meticulous documentation, and end-to-end requirements traceability [65]. Building a robust safety case with sufficient evidence is paramount [78].
1.9 Conclusion 41 for system behaviour. Research is actively exploring different integration architectures and their implications for validation. Agentic AI and AI agents brings new challenges required the advancements of research focusing on developing new V&V techniques tailored to the dynamic nature of agentic AI, including advancing methods in runtime monitoring and formal verification that can cope with learning-based components and non-determinism. Creating simulation platforms that can model complex, real-world physics and multi-agent interactions will be crucial for testing edge systems exhaustively before deployment. The use of digital twin and immersive triplet environments could enable the safe exploration of an agent’s behaviour under a wide range of standard and adverse conditions, helping to identify potential failure modes early. In this context, based on the technology trends research should address the system-level and collaborative aspects of agentic AI at the edge by creating frameworks for validating not only individual agents but also the collective, emergent behaviour of multiagent systems. Developing techniques to ensure that the goals of individual agents remain aligned with the overall system objectives, even as they adapt and learn, is paramount. Research into explainable XAI and IAI for edge devices is required, as it will enable human operators to understand, trust, and effectively manage the decisions of autonomous agents, ensuring safe and predictable operation in complex, real-world scenarios. 1.9 Conclusion The rapid advancement and deployment of edge AI necessitate a parallel evolution in the designers’ ability to ensure that edge AI systems are safe, reliable, fair, and aligned with human values. Verification, as defined by standards such as ISO/IEC 22989, is the assurance through objective evidence that specified requirements have been fulfilled, forming a cornerstone of building essential trust. It provides the rigorous checks needed to confirm that AI systems are built according to their intended design and specifications. The unique characteristics of AI, particularly its potential opacity, nondeterminism, complex data dependencies, and difficulty in formally specifying requirements for emergent behaviours, pose significant challenges to traditional V&V approaches. The black-box nature of many models hinders direct inspection, scalability limitations restrict the application of formal methods, and the dynamic nature of edge AI systems and their environments demands continuous evaluation beyond design-time checks. Addressing conceptual challenges related to fairness, value alignment, and
42 Edge AI Systems Verification and Validation adversarial robustness requires ongoing fundamental research, and significant progress is being made. International standards provide common terminology (ISO/IEC 22989), frameworks for trustworthiness (ISO/IEC TR 24028), and management systems for responsible AI governance (ISO/IEC 42001). Methodologically, an increasing number of researchers are adopting formal methods for specific AI and edge AI verification tasks, developing advanced testing techniques (e.g., metamorphic and adversarial testing). Metamorphic testing techniques are used to verify the behaviour of AI models, when predicting the exact output for a given input is challenging or impossible. The metamorphic testing techniques focus on identifying relationships between inputs and outputs, known as metamorphic relations, that act as logical rules or properties that should hold true when inputs are modified. Adversarial testing is a technique in which inputs are intentionally designed to expose weaknesses or flaws in a system, thereby identifying scenarios where the system produces harmful or biased outputs. This enables testers to identify vulnerabilities and ensure the system responds safely and effectively [37]. Generative AI excels at pattern recognition, classification, and predictive analytics, generates new patterns and multimodal content (e.g., text, sound, images) and plays a dual role in the verification and validation process, for example, as part of an edge AI system that has to be verified and validated and as a technology that supports the V&V processes by generating V&V requirements, specifications and automatically performing the V&V. The V&V of emerging edge AI agents face challenges arising from the inherent autonomy and the dynamic environments in which these agents operate. The agents can rely on machine learning models that can produce non-deterministic outputs, making the behaviour difficult to predict and formally verify. Continuous interaction with external environments introduces an extensive and unpredictable operational space, where unforeseen events can lead to emergent behaviours that were not anticipated during the design and testing phases, posing risks to safety and reliability. Further challenges for the V&V processes are the unique constraints of the edge environment itself. Edge AI systems must operate within the limitations of computational power, memory, and energy, which can impact the performance and consistency of their decision-making processes. Edge AI agents must make real-time decisions, where latency is a critical factor. Validating that an edge AI agent responds correctly and within strict time constraints, especially when facing intermittent connectivity or degraded
1.9 Conclusion 43 sensor input, is a significant hurdle that requires novel testing methodologies beyond traditional software V&V. The process of V&V agentic edge AI systems requires addressing the interaction between human and AI agent components to ensure that the agent’s behaviour is understandable and transparent to human users, which allows for efficient oversight and human intervention when necessary. The V&V process must confirm that the edge system can communicate its state and intentions evidently and that its autonomous actions are auditable, interpretable and explainable. This supports explainable AI techniques, implementing runtime verification for operational assurance, and opening the use of neuro-symbolic architectures to bridge the gap between learning and reasoning. Neuro-symbolic AI is a type of AI that integrates neural and symbolic AI architectures to address the weaknesses of each, providing a robust AI capable of reasoning, learning, and cognitive modelling. Continued research is essential to develop more scalable and robust verification techniques that can handle the complexity of new edge AI systems. Addressing foundational AI safety problems, enhancing automated and human-centric V&V approaches, and building comprehensive, trustworthy AI frameworks that integrate technical verification with ethical considerations and governance are key priorities. Achieving verifiably trustworthy AI requires a holistic perspective, acknowledging the interplay between hardware, software, AI models, data, systems, processes, and the technical, application, and environmental contexts [40]. Achieving the goal of trustworthy AI and edge AI systems that are demonstrably beneficial and responsibly integrated into society requires elevating validation beyond a mere technical, end-of-phase check. It demands a holistic, continuous, and lifecycle-integrated approach [54]. This approach must rigorously integrate technical validation (ensuring robustness, reliability, and security) with user-centric validation (confirming usability, fitness for purpose, meeting needs), ethical validation (assessing fairness, accountability, and value alignment), and real-world effectiveness monitoring [90]. Success requires multidisciplinary collaboration, bringing together AI/ML experts, software engineers, domain specialists, human factors engineers, ethicists, social scientists, legal experts, end-users, and regulators. International standard bodies like ISO and organisations like NIST provide essential frameworks, common terminology, and guidance (e.g., ISO 9000,
44 Edge AI Systems Verification and Validation ISO/IEC/IEEE 15288, ISO/IEC 22989, ISO/IEC TR 24028, ISO/IEC 42001, NIST AI RMF) [89]. The rapidly evolving nature of AI means that significant ongoing research and innovation in validation techniques are imperative to address the challenges effectively. Edge AI system validation is a dynamic and increasingly critical field. As AI capabilities continue to advance and these systems become more deeply embedded in our lives, the methods used to ensure they are fit for purpose, safe, and aligned with human values must also evolve. The focus is shifting from static, pre-deployment checks towards continuous, adaptive, and context-aware validation processes that span the entire edge AI lifecycle. Addressing the complex technical, ethical, and societal challenges associated with edge AI validation requires sustained research, multidisciplinary collaboration, and international cooperation. Continued innovation in validation methodologies and tools will be essential to harness the transformative potential of AI responsibly and build a future where edge AI systems are trustworthy and integrated into industrial and business processes. Acknowledgements This publication has received funding through the projects Chips JU EdgeAI and HE dAIEDGE. The Chips JU EdgeAI “Edge AI Technologies for Optimised Performance Embedded Processing” project is supported by the Chips Joint Undertaking and its members including top-up funding by Austria, Belgium, France, Greece, Italy, Latvia, Netherlands, and Norway under grant agreement No 101097300. The HE dAIEDGE “A network of excellence for distributed, trustworthy, efficient and scalable AI at the Edge” project is supported under grant agreement No 101120726. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the Chips Joint Undertaking. Neither the European Union nor the granting authority can be held responsible for them. References [1] A. Lavin et al., “Technology readiness levels for machine learning systems,” Nature Communications, vol. 13, no. 1, p. 6039, Oct. 2022, https://doi.org/10.1038/s41467-022-33128-9.
References 45 [2] S. Mahmud, S. Saisubramanian, and S. Zilberstein, “Verification and Validation of AI Systems Using Explanations,” Proceedings of the AAAI Symposium Series, vol. 4, no. 1, pp. 76–80, Nov. 2024, https://doi.org/ 10.1609/aaaiss.v4i1.31774. [3] “Verification and Validation of Systems in Which AI is a Key Element - SEBoK,” sebokwiki.org. https://sebokwiki.org/wiki/Verification_and _Validation_of_Systems_in_Which_AI_is_a_Key_Element. [4] “ISO/IEC TR 24028:2020 – Overview of trustworthiness in artificial intelligence,” BSI, 2020. https://www.bsigroup.com/en-IN/training-co urses/isoiec-tr-240282020--overview-of-trustworthiness-in-artificialintelligence/. [5] “ISO/IEC TR 24028:2020 - Information technology - Artificial intelligence - Overview of trustworthiness in artificial intelligence” ISO. https://www.iso.org/standard/77608.html. [6] NIST, “AI Risk Management Framework,” Artificial Intelligence Risk Management Framework (AI RMF 1.0), vol. 1, Jan. 2023, https://doi.or g/10.6028/nist.ai.100-1. [7] J. Jeon, “Standardization Trends on Safety and Trustworthiness Technology for Advanced AI”, 2024, https://arxiv.org/abs/2410.22151. [8] T. R. McIntosh, T. Susnjak, T. Liu, P. Watters, and M. N. Halgamuge, “Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence,” arXiv (Cornell University), Feb. 2024. Available at: https://doi.org/10.48550/arxiv.2402.09880 [9] “ISO/IEC 22989:2022 – Information technology – Artificial intelligence – Artificial intelligence concepts and terminology,” Edition 1, 2022. ht tps://www.iso.org/standard/74296.html [10] “Recommendation of the Council on Artificial Intelligence,” OECD Legal Instruments, 2025. https://legalinstruments.oecd.org/en/instr uments/oecd-legal-0449 [11] “What is catastrophic forgetting?” IBM, April 2025. https://www.ibm. com/think/topics/catastrophic-forgetting [12] T. Meuser, et.al. “Revisiting Edge AI: Opportunities and Challenges”. IEEE Internet Computing, vol. 28, July-August 2024. https://www.co mputer.org/csdl/magazine/ic/2024/04/10621659/1Z5lGDb639C [13] Z. Ren and C. J. Anumba, “Multi-agent systems in construction–state of the art and prospects,” Automation in Construction, vol. 13, no. 3, pp. 421–434, 2004, https://doi.org/10.1016/j.autcon.2003.12.002. [14] G. Papagni, J. de Pagter, S. Zafari, M. Filzmoser, and S. T. Koeszegi, “Artificial agents’ explainability to support trust: considerations on
46 Edge AI Systems Verification and Validation timing and context,” AI & Society, Vol. 38, No. 2, pp. 947–960, 2023. https://doi.org/10.1007/s00146-022-01462-7 [15] P. Wang and H. Ding, “The rationality of explanation or human capacity? Understanding the impact of explainable artificial intelligence on human-AI trust and decision performance,” Information Processing & Management, Vol. 61, No. 4, p. 103732, 2024. https://doi.org/10.1016/ j.ipm.2024.103732. [16] F. Sado, C. K. Loo, W. S. Liew, M. Kerzel, and S. Wermter, “Explainable Goal-driven Agents and Robots - A Comprehensive Review”, ACM Computing Surveys, Volume 55, Issue 10, pp. 1–41, 2023. https://do i.org/10.1145/3564240. [17] R. Sapkota, K. I. Roumeliotis, and M. Karkee, “AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenge,” arXiv.org, 2025. https://arxiv.org/abs/2505.10468. [18] C. Riedl and D. De Cremer, “AI for collective intelligence,” Collective Intelligence, vol. 4, no. 2, Apr. 2025, https://doi.org/10.1177/26339137 251328909. [19] F. Piccialli, D. Chiaro, S. Sarwar, D. Cerciello, P. Qi, and V. Mele, “AgentAI: A comprehensive survey on autonomous agents in distributed AI for industry 4.0,” Expert Systems with Applications, vol. 291, p. 128404, Oct. 2025, https://doi.org/10.1016/j.eswa.2025.128404. [20] W. Xu, Z. Liang, K. Mei, H. Gao, J. Tan, and Y. Zhang, “A-MEM: Agentic Memory for LLM Agents,” arXiv.org, 2025. https://arxiv.or g/abs/2502.12110. [21] D. B. Acharya, K. Kuppan and B. Divya, “Agentic AI: Autonomous Intelligence for Complex Goals—A Comprehensive Survey,” in IEEE Access, vol. 13, pp. 18912-18936, 2025, https://www.doi.org/10.1109/ ACCESS.2025.3532853. [22] R. Zhang et al., “Toward Agentic AI: Generative Information Retrieval Inspired Intelligent Communications and Networking,” arXiv.org, 2025. https://arxiv.org/abs/2502.16866. [23] M. Gridach, J. Nanavati, K. Zine, L. Mendes, and C. Mack, “Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions,” arXiv.org, 2025. https://arxiv.org/abs/2503.08979. [24] E. Miehling et al., “Agentic AI Needs a Systems Theory,” arXiv.org, 2025. https://arxiv.org/abs/2503.00237. [25] S. Hong et al., “MetaGPT: Meta Programming for Multi-Agent Collaborative Framework,” arXiv.org, Aug. 07, 2023. https://arxiv.org/abs/23 08.00352.
References 47 [26] U. M. Borghoff, P. Bottoni, and R. Pareschi, “Human-artificial interaction in the age of agentic AI: a system-theoretical approach,” Frontiers in Human Dynamics, vol. 7, May 2025, https://doi.org/10.3389/fhumd. 2025.1579166. [27] J. Heer, “Agency plus automation: Designing artificial intelligence into interactive systems,” Proceedings of the National Academy of Sciences, Vol. 116, No. 6, pp. 1844–1850, 2019. https://doi.org/10.1073/pnas.180 7184115. [28] E. Oliveira, K. Fischer, and O. Stepankova, “Multi-agent systems: which research for which applications,” Robotics and Autonomous Systems, vol. 27, no. 1-2, pp. 91–106, 1999, https://doi.org/10.1016/S0921-8890 (98)00085-2. [29] Validation - Glossary | CSRC - NIST Computer Security Resource Center, https://csrc.nist.gov/glossary/term/validation [30] Verification and validation - Wikipedia, https://en.wikipedia.org/wiki/ Verification_and_validation [31] Design Review, Verification and Validation - Quality Gurus, https://ww w.qualitygurus.com/design-review-verification-and-validation/ [32] Verification Versus Validation – What’s the Difference? - Climedo, http s://climedo.de/en/blog/verification-versus-validation-whats-the-differe nce/ [33] System Validation - SEBoK, https://sebokwiki.org/wiki/System_Valid ation [34] Implementing ISO 15288 V&V Processes using the V&V Studio - The Reuse Company, https://www.reusecompany.com/wp-content/uploads/ 2021/02/VV-Studio-Webinar-Jan-2021.pdf [35] Verification (glossary) - SEBoK, https://sebokwiki.org/wiki/Verificatio n_(glossary) [36] “Sapien’s AI Glossary of Data Terms, Definitions & Insights,” Sapien.io, 2025. https://www.sapien.io/glossary/all [37] A. Pande, “Metamorphic and adversarial strategies for testing AI systems,” Ministry of Testing, Jan 14, 2025. https://www.ministryoftestin g.com/articles/metamorphic-and-adversarial-strategies-for-testing-ai-s ystems [38] ISO and IEC Make Foundational Standard on Artificial Intelligence Publicly Available, https://www.holisticai.com/news/iso-iec-22989foundational-standard-on-ai-open-source [39] The foundational standards for AI | JTC 1, https://jtc1info.org/wp-cont ent/uploads/2022/06/03_08_Paul_Milan_Wei_The-foundational-stan dards-for-AI-20220525-ww-mp.pdf
48 Edge AI Systems Verification and Validation [40] E. Manziuk, O. Barmak, I. Krak, O. Mazurets, and T. Skrypnyk, “Formal Model of Trustworthy Artificial Intelligence Based on Standardization,” CEUR-WS.org, https://ceur-ws.org/Vol-2853/short18.pdf [41] ISO/IEC TR 24028:2020 - Information technology - Artificial intelligence - OECD.AI, https://oecd.ai/en/catalogue/tools/isoiec-tr-2402820 20-information-technology-artificial-intelligence-overview-of-trustwo rthiness-in-artificial-intelligence [42] Exploring the landscape of trustworthy artificial intelligence: Status and challenges, https://content.iospress.com/articles/intelligent-decision-t echnologies/idt240366 [43] ISO 42001 Artificial Intelligence Management System - Amazon Web Services (AWS), https://aws.amazon.com/compliance/iso-42001-faqs/ [44] ISO 42001 - AI Management System - BSI, https://www.bsigroup.com /en-US/products-and-services/standards/iso-42001-ai-management-s ystem/ [45] ISO/IEC 22989:2023 Understanding AI Concepts and Definitions Training Course | BSI, https://www.bsigroup.com/en-ID/trainingcourses/isoiec-229892023-understanding-ai-concepts-and-definitions -training-course/ [46] AI Compliance Audit: Step-by-Step Guide - Dialzara, https://dialzara.c om/blog/ai-compliance-audit-step-by-step-guide/ [47] R. Prieto, “Verification and Validation of Project Management Artificial Intelligence Key Points,” Jun. 2020. https://www.researchgate.net/publi cation/342452507_Verification_and_Validation_of_Project_Managem ent_Artificial_Intelligence_Key_Points [48] AI audit checklist (updated 2025) | Complete AI audit procedures | Technical evaluation framework | System reliability guide | Compliance checklist | Lumenalta, https://lumenalta.com/insights/ai-audit-checklis t-updated-2025 [49] Y. Wang and S. H. Chung. “Artificial intelligence in safety-critical systems: a systematic review,” Industrial Management & Data Systems,| Emerald Insight, Dec. 2021. https://www.emerald.com/insight/content/ doi/10.1108/imds-07-2021-0419/full/html [50] Trustworthy AI - AI@UCSF - University of California San Francisco, https://ai.ucsf.edu/trustworthy [51] AI Risks and Trustworthiness - NIST AIRC - National Institute of Standards and Technology, https://airc.nist.gov/airmf-resources/ai rmf/3-sec-characteristics/
References 49 [52] S. A. Seshia, D. Sadigh, and S. Shankar Sastry. Toward Verified Artificial Intelligence. Communications of the ACM, July 2022. https: //cacm.acm.org/research/toward-verified-artificial-intelligence/ [53] AI, Opacity, and Personal Autonomy, https://d-nb.info/1275205275/34 [54] Measure - NIST AIRC - National Institute of Standards and Technology, https://airc.nist.gov/airmf-resources/playbook/measure/ [55] Understanding the NIST AI RMF: What It Is and How to Put It Into Practice - Secureframe, https://secureframe.com/blog/nist-ai-rmfy [56] AI Life Cycle Core Principles - CodeX - Stanford Law School, https: //law.stanford.edu/2023/03/17/ai-life-cycle-core-principles/ [57] Ethical and societal implications of algorithms, data, and artificial intelligence: a roadmap for research - Nuffield Foundation, https://www.nu ffieldfoundation.org/sites/default/files/files/Ethical-and-Societal-Impl ications-of-Data-and-AI-report-Nuffield-Foundat.pdf [58] Messages on “When using AI systems, what are some best practices for ensuring the results you receive are accurate, relevant, and aligned with your original goals?” - ProjectManagement.com, https://www.projectm anagement.com/discussion-topic/203772/when-using-ai-systems--wha t-are-some-best-practices-for-ensuring-the-results-you-receive-are-a ccurate--relevant--and-aligned-with-your-original-goals-?sort=asc&p ageNum=39 [59] A Framework for the Verification and Validation of Artificial Intelligence Machine Learning Systems - JagWorks@USA - University of South Alabama, https://jagworks.southalabama.edu/theses_diss/137/y [60] Trustworthy AI - The Data Science Institute at Columbia University, https://datascience.columbia.edu/news/2020/trustworthy-ai/ [61] NIST launches ARIA program to assess societal impacts, ensure trustworthy AI systems, https://industrialcyber.co/ai/nist-launches-aria-pro gram-to-assess-societal-impacts-ensure-trustworthy-ai-systems/ [62] User Acceptance Testing (UAT): Definition, Process, and Tools - LambdaTest, https://www.lambdatest.com/learning-hub/user-acceptance-test ing [63] Human-AI Interaction and User Satisfaction: Empirical Evidence from Online Reviews of AI Products - ResearchGate, https://www.research gate.net/publication/390142284_Human-AI_Interaction_and_User_Sat isfaction_Empirical_Evidence_from_Online_Reviews_of_AI_Product s/download
50 Edge AI Systems Verification and Validation [64] J. Renkhoff, K. Feng, M. Meier-Doernberg, A. Velasquez, and H. H. Song, “A Survey on Verification and Validation, Testing and Evaluations of Neurosymbolic Artificial Intelligence,” IEEE transactions on artificial intelligence, pp. 1–15, Jan. 2024, https://doi.org/10.1109/tai.2024.335 1798. [65] A. E. Goodloe, “Assuring Safety-Critical Machine Learning-Enabled Systems: Challenges and Promise,” Computer, vol. 56, no. 9, pp. 83–88, Sep. 2023, https://doi.org/10.1109/mc.2023.3266860 [66] A. Woodie, “Top 10 Challenges to GenAI Success,” BigDATAwire, Jan. 22, 2024. https://www.bigdatawire.com/2024/01/22/top-10-challenges -to-genai-success/ [67] C. Bronsdon, “AI Safety Metrics: How to Ensure Secure and Reliable AI Applications” - Galileo AI, 2025, https://www.galileo.ai/blog/introd uction-to-ai-safety [68] R. Camacho, “A Practical Guide for AI in Safety-Critical Embedded Systems - Parasoft, 2025, https://www.parasoft.com/blog/ai-in-safety-c ritical-embedded-systems/ [69] Y. Y. Elboher et., al “Formal Verification of Deep Neural Networks for Object Detection,” Arxiv.org, 2023. https://arxiv.org/html/2407.01295 [70] What are non-deterministic AI outputs? - Statsig, 2024, https://www.st atsig.com/perspectives/what-are-non-deterministic-ai-outputs- [71] K. Leahy et al., “Grand Challenges in the Verification of Autonomous Systems,” arXiv (Cornell University), Nov. 2024, https://doi.org/10.485 50/arxiv.2411.14155. [72] Symbolic AI vs. Deep Learning: Key Differences and Their Roles in AI Development, https://smythos.com/ai-agents/agent-architectures/symb olic-ai-vs-deep-learning/ [73] Symbolic AI vs. Machine Learning: A Comprehensive Guide - SmythOS, https://smythos.com/ai-agents/ai-tutorials/symbolic-ai-v s-machine-learning/ [74] O. Vermesan, V. Piuri, F. Scotti, A. Genovese, R. D. Labati, and P. Coscia, “Explainability and Interpretability Concepts for Edge AI Systems,” River Publishers eBooks, pp. 197–227, Feb. 2024, https://doi.org/10.1 201/9781003478713-9. [75] What Is Explainable AI (XAI)? Palo Alto Networks, https://www.palo altonetworks.com/cyberpedia/explainable-ai [76] Deep Learning’s Challenges and Neurosymbolic AI’s Solutions - AskUI, 2024. https://www.askui.com/blog-posts/deep-learnings-ch allenges-and-neurosymbolic-ais-solutions
2.3 Evaluation Methodology and Metrics 57 Figure 2.2 Vertical Hybrid Federated Learning Architecture Figure 2.3 Horizontal Hybrid Federated Learning Architecture equivalent to a pure FL approach. We investigate the impact of varying rhconfigurations in the overall DL performance in Section 4. We show in Figure 2.3 an overview on the conceptual architecture of Vertical Hybrid FL. We further highlight that the hybridization rates determine the amount of required privacy preservation. The maximum value (1) corresponds to when all clients share their local data with the server. Vice versa, the minimum value (0) forbids any transmission of raw data outside the clients’ devices. 2.3 Evaluation Methodology and Metrics In this section, we describe the experiments that we performed to investigate the performances of the two hybrid approaches explained in the previous
58 Pioneering the Hybridization of Federated Learning section. First, we will describe the methodology used to test the effectiveness of HFL across two datasets: FEMNIST and UCI HAR. The widely recognized University of California Irvine (UCI) HAR dataset [8] was created using data from smartphone accelerometer and gyroscope sensors, which were used to classify six types of human activities. It was gathered from 30 volunteers, each carrying a smartphone while performing six distinct activities: walking, walking upstairs, walking downstairs, sitting, standing, and lying down. The dataset consists of time-series data captured across the three axes of both sensors, along with the corresponding activity labels. It has been extensively used in research for developing and assessing HAR models using machine learning techniques. Each record in the dataset contains a vector of 561 features, derived from both time and frequency domain calculations. The FEMNIST dataset [9] is an adaptation of the extended version of the MNIST dataset that has been modified to be suitable for FL tasks. The MNIST dataset contains 28 by 28 pixels images of handwritten digits and characters (62 classes in total), and the goal of the DL model is to guess the actual character represented. The FEMNIST dataset groups the elements on top of the user that actually performed the handwriting, producing a number of sub-datasets each of them with a different style of writing. The number of users of the FEMNIST dataset is 3500, however, for the purpose of our experiment, we considered only 30 users, in order to make experiments comparable between the two datasets. For UCI HAR we employed, as a base local model, a simple feed-forward neural network, while for FEMNIST we adopted a convolutional network. Each of the local sub-datasets is split into training and test set using a stratified split with a 70%-30% ratio. We performed federated classification experiments by employing 6 epochs and 20 rounds of federation, recording the accuracy score at the end of the last round. We tested both vertical and horizontal hybridization, by setting alternately rvand rhto values spanning from 0% to 100% with a 10% step. The experiments were implemented in Python using the Flower framework (https://flower.readthedocs.io/en/latest/). Since Flower does not support the implementation of hybridization, we adopted the two following methods to simulate the two hybridization methods (as Figure 2.4 suggests): • Vertical Hybridization was simulated by aggregating all the clients that share their whole dataset instead of the weights into a single client.
2.4 Evaluation Results 59 Figure 2.4 Implementation of the Vertical Hybridization in Flower • Horizontal Hybridization was simulated by extracting from each client the portion of the training dataset that they aim to share and assigning it to a new “sink” client. Each experiment was then repeated 20 times, by randomizing the clients or the dataset portions to be shared. This ensures scientific rigor and smooths out certain corner situations that may arise. 2.4 Evaluation Results This section presents the outcome of the experiments presented in the previous section. The results are aimed at evaluating the HFL approaches on the datasets. We investigate how different levels of data sharing in vertical
60 Pioneering the Hybridization of Federated Learning and horizontal hybridization affect model performance. We first discuss the results for the UCI HAR dataset, followed by FEMNIST, to uncover any dataset-specific trends and performance differences. A general overview of the results shows, as expected, an overall improvement in model performance as the degree of hybridization increases, which is notable for both hybrid approaches. The UCI HAR dataset performed best. Even with the relatively small model size, it consistently achieved excellent results, maintaining accuracy above 90% and reaching almost 95% in experiments with higher levels of hybridisation. This is shown in Figures 2.5 and 2.6, where the increase in model performance as the level of hybridisation increases is evident, with an increase of 2 percentage points already at a low level of horizontal hybridization (20%). The FEMNIST is the dataset where the performance improvements from sharing data is most noticeable. As we can observe in the Figures 2.7 and 2.8, the sharing of a small number of data points could lead to a significant improvement in performance. Specifically, sharing 10% of the data led to an improvement of approximately 5% in accuracy, while sharing 20% led to an additional improvement of approximately 2/3, resulting in a final performance of 88%. This is comparable to the 90% accuracy obtained through centralized training. Beyond a data sharing rate of 30%, the improvements Figure 2.5 Horizontal Hybridization Results for UCI HAR
2.4 Evaluation Results 61 Figure 2.6 Vertical Hybridization Results for UCI HAR Figure 2.7 Horizontal Hybridization Results for FEMNIST obtained are increasingly marginal, with a maximum of 1%. This suggests that a data sharing rate of 20% represents an optimal balance between performance and data sharing. The results demonstrate that the sharing of a portion of the dataset has a significant positive impact on the model’s performance. This effect was
62 Pioneering the Hybridization of Federated Learning Figure 2.8 Vertical Hybridization Results for FEMNIST observed across two distinct datasets, UCI HAR and FEMNIST, indicating that the improvement is not specific to any problem or model. The performance enhancement was especially evident in the case of the FEMNIST dataset. 2.5 Conclusion and Future Works In this paper we examined the effects of hybridization for Federated Learning scenarios. We specifically directed our research towards Human Activity Recognition, imagining scenarios in which certain clients would be willing to share (part of) their data with the central server for a reward, penalizing their privacy to an extent. Results showed that a minimal amount of hybridization does provide an increase in the performance. The extent to which the privacy of the users is compromised by this is a future work. We aim to study how to select carefully data in a way in which the privacy is minimally affected, as well as to blend the two hybridization techniques, to select the best configuration. Acknowledgements This research is funded by the “Progetto Casa delle Tecnologie Emergenti” - Comune di Bologna - PSC MISE 2014-2020.
References 63 References [1] K. S. Awaisi, Q. Ye and S. Sampalli, “A Survey of Industrial AIoT: Opportunities, Challenges, and Directions,” in IEEE Access, vol. 12, pp. 96946-96996, 2024. https://doi.org/10.1109/ACCESS.2024.3426279 [2] S. Ankalaki, “Simple to Complex, Single to Concurrent SensorBased Human Activity Recognition: Perception and Open Challenges,” in IEEE Access, vol. 12, pp. 93450-93486, 2024. https://doi.org/10.110 9/ACCESS.2024.3422831 [3] H. Brendan McMahan, E. Moore, D. Ramage, S. Hampson, B. Agüera y Arcas, “Communication-Efficient Learning of Deep Networks from Decentralized Data”, arXvi, https://arxiv.org/abs/1602.05629 [4] J. Cui, H. Zhu, H. Deng, Z. Chen, and D. Liu, “Fearh: Federated machine learning with anonymous random hybridization on electronic medical records,” Journal of Biomedical Informatics, vol. 117, p. 103735, 2021. https://doi.org/10.1016/j.jbi.2021.103735 [5] M. Moshawrab, M. Adda, A. Bouzouane, H. Ibrahim, and A. Raad, “Reviewing federated learning aggregation algorithms; strategies, contributions, limitations and future perspectives,” Electronics, vol. 12, no. 10, p. 2287, May 2023. https://doi.org/10.3390/electronics12102287 [6] H. Lee and D. Seo, “FedLC: Optimizing Federated Learning in Non-IID Data via Label-Wise Clustering,” in IEEE Access, vol. 11, pp. 4208242095, 2023. https://doi.org/10.1109/ACCESS.2023.3271517 [7] D. Ibdah, N. Lachtar, S. M. Raparthi and A. Bacha, “Why Should I Read the Privacy Policy, I Just Need the Service”: A Study on Attitudes and Perceptions Toward Privacy Policies,” in IEEE Access, vol. 9, pp. 166465-166487, 2021. https://doi.org/10.1109/ACCESS.2021.3130086 [8] Anguita, Davide, Alessandro Ghio, Luca Oneto, Xavier Parra, and Jorge Luis Reyes-Ortiz. “A public domain dataset for human activity recognition using smartphones.” In Esann, vol. 3, p. 3. 2013. https: //www.esann.org/sites/default/files/proceedings/legacy/es2013-84.pdf [9] Caldas, Sebastian, Sai Meher Karthik Duddu, Peter Wu, Tian Li, Jakub Koneˇ cný, H. Brendan McMahan, Virginia Smith, and Ameet Talwalkar. “Leaf: A benchmark for federated settings.” arXiv preprint arXiv:1812.01097. https://arxiv.org/abs/1812.01097
3 Edge Intelligence Architecture for Distributed and Federated Learning Systems Pierluigi Dell’Acqua, Lorenzo Carnevale, and Massimo Villari University of Messina, Italy Abstract In recent years, deep neural networks have achieved success in Electric Vehicles (EVs) monitoring, primarily due to their scalability with large-scale data and numerous model parameters. However, EVs rely on resource-constrained edge devices that struggle with complex models, and data privacy concerns prevent sharing data outside the owning device. Federated Learning (FL) and Knowledge Distillation (KD) have emerged as key solutions, enabling model simplification and distributed training on private data. FL allows models to be trained locally on edge devices, addressing privacy concerns while keeping data decentralized and avoiding central server dependencies. This approach requires lightweight models optimized for edge intelligence deployment. To address this challenge, we propose architectural solutions leveraging FL, KD and model compression techniques to create simplified Artificial Neural Networks (ANNs) suitable for edge devices in EVs. The proposed architecture integrates these methods into a federated environment, ensuring distributed training while maintaining computational efficiency for EV monitoring and predictive maintenance applications. By combining FL, KD, and model compression, our approach enables efficient and privacypreserving Machine Learning (ML) models, enhancing Edge Intelligence (EI) for EV monitoring in resource-constrained settings. 65
66 Edge Intelligence Architecture for Distributed and Federated Learning Systems Keywords: knowledge distillation, federated learning, edge computing, IoT, edge intelligence. 3.1 Introduction The automotive industry has witnessed a substantial expansion in both the scale and intricacy of electrical and electronic system architectures. In this regard, the EVs production is becoming increasingly prevalent in the field. Therefore, the challenge of predicting diagnosing faults and improving the LI-ion Battery (LIB) lifetime in EVs becomes progressively more demanding. Modern monitoring systems approach the battery state parameters maintenance, such as the State of Charge (SoC), State of Health (SoH), State of Power, Remaining Useful Life, within safe limits, safeguarding the battery’s safety. These are based on ML and Artificial Intelligence (AI) models, which can adapt the analysis to the specific technology. Furthermore, the computational trend is moving the data elaboration to the edge, with no requiring EVs producers to share their own diagnostic data as training datasets, preserving the industrial secrets and users’ privacy. Actually, edge devices are characterized by low capacity and low performance that generally do not allow complex operations. For this reason, model simplification becomes necessary. This scenario can be intended as a specific EI scenario, where AI models and algorithms are deployed and executed on resource-constrained edge devices. EI refers to the integration of AI capabilities directly on edge devices, enabling real-time data processing, decision-making and autonomy at the edge of the network. In such contexts, the need to balance computational demands with limited hardware resources is critical. Therefore, FL strategies, that aim to train models in a distributed manner and keep data locally on users’ devices to preserve the privacy, can be adopted. The basic idea of FL unfolds in several key stages: i) a model based on an ANN is centrally initialized and subsequently disseminated to various peripheral devices; ii) these devices independently train the model using their locally available data, sending back for aggregation [17] to the central server only the outcomes of this localized training, such as the model weights. This work aims to define a comprehensive and unified EI architecture designed to the specific demands of EVs monitoring systems. The goal is to leverage the potential of edge computing and distributed AI strategies to
3.3 Use Case 73 FL meets limitations and several KD frameworks are designed to overcome them. For example, the [9] improves communication efficiency reducing communication among clients, [5] adopts a grouping strategy to group clients that share homogeneous resources improving communication efficiency and balancing computing resources, [1], [2], [16] focus mainly on handling model and data heterogeneity. 3.2.4 Beyond the State of the Art Actually, the scientific literature is challenging proposing a standard architecture for EI. Many methodologies may be involved both for training and inference. FL, KD, quantization and pruning are examples of enabling key technologies available for the exploitation of AI models into the edge of the network. This work aims to propose an architecture that includes all the mentioned methodologies on a simple workflow that optimizes the deployment and execution of EI solutions for EV. 3.3 Use Case The use case focuses on the development of health monitoring systems for LIBs lifetime and contextual risk assessment in EVs. Among the various monitoring strategies, ML-based methods provide high accuracy, although they require large datasets for effective training [14]. Due to the complexity and critical nature of managing LIBs in EVs, deploying ML systems on resource-constrained microcontroller-based platforms is crucial. This integration enables real-time monitoring and predictive maintenance, optimizing battery performance and extending operational life. Although cloud computing offers incomparable performance that can be leveraged for centralized activities (e.g., initial training, FL central aggregation), three main reasons lead to the need to push AI computation to the edge: •Sensing data: the sensing layer within EVs produces a continuous and rich flow of data that cannot be transferred to the cloud to prevent bandwidth saturation. •Real-time response: in a monitoring context, it is desirable to have realtime alerts rather than waiting for a stable connection with the central processor.
74 Edge Intelligence Architecture for Distributed and Federated Learning Systems •Privacy concerns: manufacturers are not inclined to share their own diagnostic data with cloud data centers where big data are collected and processed. The basic flow begins with defining and training an AI model in the cloud, potentially utilizing a more complex model in a KD scenario [10]. Once trained, the model may undergo optimization techniques such as quantization or pruning to reduce its size and improve inference time. After these optimizations, the model is converted into formats compatible with microcontroller-based devices (e.g., ONNX, TinyML), enabling deployment on distributed clients within EVs. Once each client runs its own optimized AI model, realtime monitoring takes place locally on the EV. However, as new data continuously flows from the vehicle’s sensing layer, the model can be further refined. These improvements can be shared with the central cloud, as well as with other clients, when stable connection is available, in a FL environment. This decentralized learning process allows each client to contribute to the overall model improvement without the need to share raw data, thus preserving privacy while enabling continual learning and adaptation (Figure 3.4). This use case can be classified as level 4 within the EI framework proposed by [28], where both cloud and edge devices collaborate to perform training and inference tasks. However, in scenarios where the cloud is unable to handle training (e.g., due to the lack of a training dataset), the use case Figure 3.4 Use case scenario.
3.4 Architecture Proposal 75 shifts to a level 5, where the edge devices will be engaged for both training and inference. 3.4 Architecture Proposal In the context of EVs monitoring, leveraging data-driven approaches based on ML and AI algorithms is a critical point because of the limited computational and energy capacity of the machines. In the following sections, we will build the final architecture step-by-step, describing each involved component. 3.4.1 Assumption The proposed architecture shares computation responsibility between cloud and edge computing, exploiting EI methodologies for training and inference, such as the FL, KD, quantization and pruning. The cloud is adopted as much as possible for high-computation activities, such as serving as i) complex AI model training with existing datasets and ii) FL central node aggregator. Most of the inference and a relevant part of the training is intended to be executed on the edge. The EVs are equipped with low-performance edge devices capable of handling models that are generally not too complex. These devices will perform monitoring and diagnostic tasks on their own sensing data, thereby preserving user privacy constraints. Moreover, since clients do not maintain a continuous connection either with the central cloud or among themselves, communication efficiency should be ensured. Clients will be responsible for performing training and sharing their model parameters to contribute to global model improvements. 3.4.2 Cluster Aggregator The Cluster Aggregator, depicted in figure 3.5, is deployed in the cloud and primarily functions as the central aggregator for the FL system. This role is facilitated by several key modules responsible for specific tasks to ensure smooth and efficient operations within the federated framework. 1) Model Aggregation Module: This module is going to execute the FL core function. It generates the global model by aggregating (e.g., Federated Average [17]) trained models from the peripheral devices. This process ensures that the central model continuously improves by integrating
76 Edge Intelligence Architecture for Distributed and Federated Learning Systems Figure 3.5 Cluster Aggregator schema designed to handle FL central aggregator tasks and to implement a distillation framework adaptable during the training process. knowledge from distributed nodes without directly accessing their raw data, preserving privacy. This module is designed to handle different forms of model updates and can incorporate advanced techniques, such as weighted averaging, depending on the specific characteristics of the distributed models. 2) FL State Manager: It plays a critical role in maintaining synchronization between the cloud-based aggregator and peripheral devices. It tracks the status of each client, ensuring that the aggregation process considers only those clients that have successfully completed their local model training. The state manager keeps a record of which devices are actively participating in each round of FL, their connectivity status, and whether their contributions are valid for aggregation. This component also manages potential failures or delays in communication, ensuring the system can handle interruptions and continue functioning smoothly. Its role becomes even more significant in EVs scenario, where each EV changes its position very frequently and a stable connection cannot be guaranteed.
3.4 Architecture Proposal 77 3) System Configuration Handler: The handler is tasked with managing the configuration of the entire system. This module ensures that the software environment is correctly set up, with all dependencies and configurations aligned for optimal performance. Additionally, it handles dynamic updates to system settings, such as changing communication protocols or modifying the aggregation frequency. Moreover, it guarantees that all components are correctly initialized and maintained throughout the lifecycle of the FL process. 4) Communication Handler: It manages the communication between the cloud-based aggregator and peripheral devices. It sets up and oversees the data transfer channels, ensuring that the communication is both efficient and secure. Given the distributed nature of FL, reliable communication is crucial for transmitting model updates, hyperparameters, and any other necessary metadata between clients and the cloud. The Communication Handler also implements protocols to minimize latency, reduce communication overhead, and ensure data integrity during transfer. 5) Training and Model Distillation Modules: The Training Module on the cloud side is activated when the global model is initialized and trained using an existing dataset. Once this initial training phase is completed, the global model is shared with the peripheral clients to begin the federated learning process. The clients use the global model as a starting point, performing local training on their own data and subsequently sharing their model updates with the central aggregator. In addition to the standard training workflow, the Model Distillation Module plays a crucial role by implementing one or more KD strategies [7], [10], [19] during the training process. These strategies can be employed for several reasons: • When the global model does not achieve an acceptable level of performance, KD can be used to refine it further by leveraging smaller, more efficient models that capture the key patterns of the original data • Mainly in regression problem, different KD strategies make up for the lack of the dataset used for training • [11], [27], allowing the transfer of knowledge from the pre-trained model to the global model without needing access to the original data
78 Edge Intelligence Architecture for Distributed and Federated Learning Systems • When the FL process is subject to constraints, such as heterogeneous models and/or non-IID data across clients, KD can help align the learning processes [18]. By integrating KD into the training pipeline, the system can enhance the robustness and flexibility of the FL process, improving model performance in scenarios where traditional FL might face limitations. All Together, the modules described above form a robust infrastructure that supports efficient and secure FL and their combination enables distributed training and aggregation while preserving user privacy, ensuring system reliability and maintaining overall system integrity. 3.4.3 Cloud Components Although the Cluster Aggregator is the main component deployed on the Cloud, other elements need to be integrated to simplify model training and ANN-based model simplification, making them deployable on resourceconstrained devices embedded in EVs (Figure 3.6). 1) Compression Server: It is introduced to reduce model complexity and size, addressing the challenges posed by the low computational performance and limited storage capacity of edge devices. It utilizes various handlers (i.e., software components capable of managing specific functionalities) to apply common compression techniques such as quantization, pruning and sparsification. These techniques are crucial for optimizing models to efficiently run on edge devices with constrained resources. Quantization reduces the precision of the numerical values used to represent the model’s parameters, thereby decreasing both the model’s size and its computational requirements. This allows edge devices to process models more efficiently without compromising significant accuracy. Pruning involves removing redundant or non-contributing weights from the model, which not only reduces its complexity but can also enhance performance by simplifying the model’s structure. This results in a leaner, faster model that is more suitable for deployment in resourcelimited environments. Sparsification, on the other hand, introduces sparsity into the model by setting insignificant weights to zero. This sparsity can be leveraged by specialized hardware to accelerate computations, further enhancing the model’s performance on edge devices.
3.4 Architecture Proposal 79 Figure 3.6 Software components deployed in the Cloud. Together, these compression techniques enable the deployment of complex models on edge devices, ensuring efficient operation while maintaining the balance between performance and resource utilization. 2) Release Server: The server is responsible for preparing ANN-based models in specific formats, such as ONNX, to enable seamless integration and deployment across a wide range of platforms and devices. The ONNX format, in particular, is highly valued for its interoperability between different machine learning frameworks, allowing models to be trained in one framework and deployed in another with minimal
80 Edge Intelligence Architecture for Distributed and Federated Learning Systems conversion effort. This flexibility is crucial in environments where multiple frameworks are in use, ensuring that models can be efficiently transferred and utilized without compatibility issues. Additionally, the Release Server plays a critical role in the context of TinyML, where models must be optimized for execution on ultra-low-power devices such as microcontrollers and sensors. In this scenario, the server supports the conversion of models into highly compact formats suitable for deployment on resource-constrained edge devices. By integrating model compression techniques and optimizing for reduced memory and power consumption, the Release Server ensures that even complex ANN-based models can run efficiently in embedded systems. 3) Deploy Server: AI models will be deployed leveraging the Over-the-air (OTA) protocol that allows remote updates on microcontrollers without requiring physical access or direct connections. This approach is particularly beneficial for IoT and embedded systems where devices are often distributed in locations that are difficult to reach. The OTA deployment process begins with a central server preparing the new firmware or software update. The microcontroller periodically checks for updates via a secure wireless communication channel, such as Wi-Fi or cellular networks. When an update is available, the microcontroller downloads the update package and performs integrity checks. If the update passes the verification process, it is stored in a dedicated memory partition on the device. Finally, the microcontroller reboots and switches to the new firmware version, ensuring minimal downtime and continuous operation. Security plays a critical role in the OTA update process. To prevent unauthorized or malicious updates, encryption methods, authentication protocols, and secure boot mechanisms are often employed to ensure the integrity and authenticity of the update process. 4) Data Storage Server: Storage space provides crucial functionality. As its name implies, this server is responsible for storing large datasets required for the training process. It ensures that data is readily accessible and managed efficiently, supporting the extensive data requirements of modern machine learning algorithms. The Data Storage Server also handles data preprocessing and augmentation tasks, preparing the data in a suitable format for training. These integrated components work in tandem to create a robust and efficient pipeline for deploying ANN-based models on resource-constrained
3.4 Architecture Proposal 81 devices embedded in EVs. By addressing the challenges of model size, complexity, and data management, the system ensures that high-performance models can operate effectively even in environments with limited computational resources. 3.4.4 Distributed Agent The FL framework is intrinsically a distributed framework. In the proposed architecture (figure 3.7), a Distributed Agent is hosted by each peripheral client, which resides on an edge device within an EV. 1) FL Client Module: Within the agent, this module is responsible for establishing communication with the central cluster, receiving the global model’s weights, and sending back the updated weights after performing local training on its local dataset. This module plays a crucial role in the federated learning process, ensuring that the client’s contributions are incorporated into the global model. It is important to emphasize the complexity of the task managed by the Communication Handler. This component must work in coordination with the FL Participation Handler to address the asynchronicity of communication. Due to the intermittent nature of connectivity between each EV and the central aggregator, as well as among the EVs themselves, there is no guarantee of continuous communication. This sporadic connectivity necessitates robust mechanisms to ensure that updates are transmitted accurately and efficiently whenever a connection becomes available, thereby maintaining the integrity and effectiveness of the federated learning process. 2) Inference Module: The local ANN-based model will be utilized in inference tasks to implement the monitoring process. This module takes the trained model and applies it to real-time data gathered from the EV, enabling functions such as predictive maintenance, performance optimization, and anomaly detection. By leveraging the local model, the EV can make intelligent decisions without relying on constant cloud connectivity, thus enhancing the system’s reliability and responsiveness. Additionally, the architecture ensures data privacy and security, as the FL approach allows data to remain on the edge device. Only the model updates, which are less sensitive than raw data, are shared with the central cluster. This decentralized approach not only enhances privacy but also reduces the bandwidth required for data transmission, which is critical in mobile and resource-constrained environments like EVs.
82 Edge Intelligence Architecture for Distributed and Federated Learning Systems Figure 3.7 The final architecture includes a Cluster Aggregator, deployed in the Cloud, and Distributed Agents, deployed on resourceconstrained edge devices.
4 Challenges and Performance of SLAM Algorithms on Resource-constrained Devices Calvin Galagain1, 2, Martyna Poreba1, and François Goulette2 1Université Paris-Saclay, CEA-List, France 2ENSTA Paris, France Abstract Evaluating the performance of Simultaneous Localization and Mapping (SLAM) algorithms is essential for the progress of robotic systems. However, conducting a comprehensive assessment of SLAM systems in the context of recent advancements is challenging due to the wide variety of hardware platforms, algorithm configurations, and datasets available. This study aims to test SLAM algorithms on resource-constrained devices such as the NVIDIA Jetson AGX Orin 64GB. Experiments are conducted with various visualbased localization algorithms that either leverage deep learning models for specific tasks within the SLAM process or are learned end-to-end to estimate camera pose. The evaluation focuses on the following systems: RDS-SLAM and VDO-SLAM, which utilize semantic information to achieve precise motion estimation; TSformer-VO, an end-to-end Transformer-based model designed for monocular visual odometry; and DeepVO, which based on recurrent neural networks. The systems are evaluated using several metrics, including ATE and RPE to assess pose accuracy and rotational drift, respectively, alongside runtime, energy consumption, and resource usage to gauge their efficiency and practicality for real-world applications. 89
90 Challenges and Performance of SLAM Algorithms Keywords: SLAM, localization, benchmarking, computational efficiency, system efficiency, performance. 4.1 Introduction and Background Simultaneous Localization and Mapping (SLAM) is a crucial technology that enables autonomous systems, such as robots, drones, and AR/VR devices, to navigate and understand their environments without relying on external reference systems. SLAM algorithms operate by simultaneously constructing a map of an unknown environment while tracking the system’s position within it. They typically utilise a combination of sensors, including cameras, LiDAR, and inertial measurement units (IMUs), to gather data about the surrounding area. Integrating these sensors improves the accuracy and robustness of SLAM but also introduces a higher computational burden, posing a substantial challenge to achieving real-time performance. To maintain operational efficiency, optimizations such as simplifying algorithms, reducing the number of processed features, and leveraging parallel processing capabilities are essential. When deployed on embedded systems, SLAM encounters unique challenges due to resource constraints, including limited processing power, memory, and energy consumption. These limitations necessitate the development of highly efficient algorithms capable of performing complex tasks in real time, such as image processing, sensor fusion, loop closure detection, and optimization. Despite advances in the field, numerous challenges persist. Ensuring robustness against sensor noise, managing varying environmental conditions, and scaling algorithms to accommodate different map sizes and complexities are critical areas of ongoing research. Furthermore, the demand for lightweight implementations that do not compromise performance underscores the continuous evolution of SLAM technologies. The primary aim of this paper is to benchmark specific SLAM methods on the NVIDIA Jetson AGX Orin 64GB, a robust embedded platform tailored for real-time processing in autonomous systems. Specifically, we perform a comprehensive performance analysis, comparing RDS-SLAM [1] and VDO-SLAM [2], two semantic SLAM techniques, against the Visual Odometry (VO) performance of CNN-based methods trained in an end-to-end manner. This benchmarking aims to assess the performance of these SLAM algorithms under constrained resource conditions, with a specific focus on two key aspects:
4.2 Related Work 91 •Performance Comparison: Evaluating the ability of each algorithm to accurately localize in dynamic environments. •Resource Utilisation Assessment: Analysing how each SLAM method leverages the resources of the Jetson AGX Orin, specifically regarding CPU and GPU performance, memory usage, and energy consumption Finally, the study provides recommendations for optimizing SLAM algorithms on embedded platforms and identifies key areas for future research and development. 4.2 Related Work SLAM has attracted considerable attention over the past few decades, leading to the development of numerous approaches [1-7]. A thorough review of the existing literature on SLAM methods reveals a diverse array of algorithms tailored for different applications and hardware platforms. For example, ORB-SLAM [3] has been widely recognized for its efficiency and robustness on conventional CPUs, demonstrating good performance across various environments. Similarly, VINS-Mono [4] and VINS-RGBD [5] are noted for their effectiveness in combining visual and inertial data. Recent studies have highlighted the potential of DNN-based SLAM systems, such as RDS-SLAM [1], VDO-SLAM [2], DF-SLAM [6], Dyna-SLAM [7], and DeepFactors [8], which leverage deep learning techniques to improve feature extraction, pose estimation, or environment understanding. End-to-end deep learning methods like DeepVO [9], TS-Former [10], and DROID-SLAM [11] also mark a significant shift in visual localization systems by using neural networks to directly learn the entire process from raw sensor data to pose estimation and map generation, bypassing traditional hand-crafted feature extraction and geometric modeling. However, these methods frequently encounter limitations compared to traditional SLAM techniques, such as lower accuracy and significant dependence on large training datasets. Neural Radiance Fields (NeRF)- based SLAM [12–16] offers a novel solution by incorporating NeRF models into SLAM systems to enhance the representation of 3D environments. In contrast to traditional SLAM methods that rely on discrete points or sparse features, NeRF-enhanced approaches produce continuous volumetric fields to create highly detailed and realistic 3D reconstructions. Through the use of neural networks, these methods can model entire scenes and facilitate photorealistic rendering by understanding the interactions of
92 Challenges and Performance of SLAM Algorithms light with surfaces. However, NeRF-based SLAM methods face challenges on embedded systems due to their high computational requirements, energy consumption, complex integration, and limited generalisation, underscoring the need for more efficient solutions. A growing number of embedded computing platforms now feature NPU/GPU units, enabling lightweight deep learning networks to function in real time. Many researchers have worked to modify SLAM algorithms for low-power embedded platforms, reengineering them to ensure compatibility with these devices. Despite these advancements in SLAM technology, implementing these algorithms on resource constrained platforms still faces considerable challenges, largely determined by the distinct characteristics of both the algorithms and the underlying embedded architectures. Although some researchers [17, 18] have explored VO systems resulting in a decreased accuracy, others [19–22] have successfully developed keyframe-based SLAM systems. Various approaches have been explored to optimize computation and power overhead in visual-inertial odometry (VIO), particularly through hardware acceleration using FPGAs [23–25]. Although advancements have been made, keyframe-based SLAM systems still face challenges in achieving an optimal balance between efficiency and accuracy for mobile robot applications. One of the most recent developments, Dynamic-VINS [26], an enhanced iteration of VINS-Mono and VINS-RGBD, showcases remarkable performance on resource-constrained platforms such as the HUAWEI Atlas200 DK and NVIDIA Jetson AGX Xavier. Finally, a hardware-software co-design approach is proposed to optimize latency, power consumption, and tracking speed in VIO systems by incorporating the Optical Flow (OF) estimation on the sensor [27]. In the VINS-Mono pipeline, feature tracking was substituted with an OF camera that employs an ASIC-based accelerator, while the other components of the VIO pipeline operate on the main processor of a Raspberry Pi Compute Module4. Assessing the performance of SLAM algorithms is essential for both researchers and users of robotic systems. Comprehensive benchmarking enables in-depth evaluations, helping to identify the most effective SLAM algorithms, and laying the groundwork for future enhancements and innovations in the field. The wide variety of hardware configurations, algorithm settings, and datasets complicates thorough comparisons across the stateof-the-art. There is a notable scarcity of research focused on benchmarking SLAM algorithms on embedded systems, highlighting significant gaps in the existing literature that this study aims to address. For example, the research
4.3 Methodology 93 presented in [28] assesses power consumption, accuracy metrics, and processing frame rates for ORB-SLAM and OpenVSLAM [29] on NVIDIA Jetson embedded systems, particularly the Jetson Nano, Jetson TX2, and Jetson Xavier. The SLAM Hive Benchmarking Suite [30] has recently addressed this challenge by offering a scalable solution that leverages container technology and cloud deployment to analyze thousands of SLAM executions. 4.3 Methodology For this comparison study, a diverse range of approaches have been selected, including semantic geometric SLAM methods as well as VO techniques derived from deep learning models trained in an end-to-end manner. This enables a comprehensive evaluation of various strategies to identify those most suitable for embedded systems, focusing on balancing performance and computational efficiency. The systems are assessed based on several metrics such as pose accuracy and rotational drift, along with an analysis of runtime, energy consumption, and GPU and CPU usage to determine their efficiency and suitability for real-world applications. 4.3.1 Selected systems We examine the following systems: RDS-SLAM, which enhances the localization process with semantic segmentation; VDO-SLAM, which utilizes semantic information for accurate motion estimation and tracking of dynamic rigid objects; TSformer-VO, an end-to-end Transformer-based model for monocular visual odometry that learns motion estimation directly from raw images; and DeepVO, which leverages recurrent networks to capture temporal dependencies. Unlike traditional SLAM algorithms, which assume a static scene, RDSSLAM (see Figure 4.1) detects and excludes dynamic objects to enhance the robustness of tracking and mapping. The algorithm extends the base framework of ORB-SLAM3 by introducing two parallel threads: a semantic segmentation thread and a semantic-based optimization thread. These threads enable the segmentation of images into static and dynamic objects using methods such as Mask R-CNN [31] or SegNet [32], while optimizing the tracking data in real time without blocking the process. The semantic thread is executed selectively to keyframes, rather than every frame. The semantic information is then propagated across the global map, where each map
94 Challenges and Performance of SLAM Algorithms Figure 4.1 Overview of RDS-SLAM [1]. point is assigned a moving probability. This probability is updated as new keyframes are processed and is used to classify map points as dynamic, static, or unknown. The algorithm identifies dynamic objects using segmentation masks from semantic models, assuming that classes such as people and vehicles are likely dynamic. Points classified as dynamic are excluded from the tracking process to avoid introducing errors in the camera pose estimation. In contrast, static points are used to improve the accuracy of the tracking. This approach allows RDS-SLAM to achieve precise tracking and robust mapping, even in environments with dynamic objects that typically pose challenges for traditional SLAM algorithms. The VDO-SLAM system (Figure 4.2) is also designed to handle dynamic environments. Before the execution of the SLAM algorithm, two crucial pre-processing steps are applied. First, Mask R-CNN is used for instancelevel semantic segmentation, which allows for the identification of both static and dynamic objects, such as vehicles and pedestrians, by generating object masks. Second, PWC-Net [33], a state-of-the-art optical flow network, is applied to estimate the dense pixel motion between consecutive frames. Using these data, VDO-SLAM can estimate the full SE(3) motion of dynamic objects, including both their linear velocity and rotational movement, while also refining its own camera pose. This integration allows
4.3 Methodology 95 Figure 4.2 Overview of VDO-SLAM [2]. for robust navigation in dynamic environments where traditional SLAM methods, which assume static surroundings, would fail. Visual odometry, a key component of the tracking phase in SLAM, can also be accomplished using methods based on Convolutional Neural Networks (CNNs). The goal is to enable the system to directly learn motion estimation from raw image inputs, eliminating the need for traditional feature extraction and matching techniques. DeepVO combines Convolutional Neural Networks (CNNs) to extract visual features from images with Long Short-Term Memory networks (LSTMs) to capture and model temporal relationships between consecutive images, enabling the prediction of camera movements. On the other hand, TSformer-VO takes a distinct approach by utilizing transformers to extract spatio-temporal features from video sequences. Unlike DeepVO, which relies on recurrent mechanisms, TSformer-VO utilizes spatio-temporal attention to capture interactions between images across both spatial and temporal dimensions. This allows for a more holistic understanding of the scene, leading to enhanced precision in estimating the camera’s 6-DoF poses. By processing long-range dependencies within the video data, TSformer-VO reduces pose drift and improves robustness in dynamic environments. Additionally, the model’s end-to-end learning framework enables it to adaptively optimize feature representations, making it a powerful alternative to traditional visual odometry methods. 4.3.2 Selected systems For the benchmarking setup, we used the NVIDIA Jetson AGX Orin 64GB, a high-performance platform specifically designed for real-time AI and embedded applications. The Jetson AGX Orin features a 2048-core NVIDIA Ampere architecture GPU with 64 Tensor Cores and a 12-core ARM CortexA78AE CPU, running at 2.2 GHz. The system is equipped with 64GB of LPDDR5 memory and provides 275 TOPS (INT8) of AI performance.
96 Challenges and Performance of SLAM Algorithms The Jetson platform allows for efficient execution of SLAM algorithms, balancing high computational power and energy efficiency, making it ideal for real-time applications in resource-constrained embedded environments. On the software side, all algorithms were deployed using Docker to guarantee consistent and efficient testing across various approaches and datasets. This containerization enables performance and results to be compared under the same conditions, thus streamlining the testing process. 4.3.3 Evaluation metrics The evaluation of SLAM algorithms is based on a set of well-defined metrics designed to assess both accuracy and computational efficiency. These metrics provide quantitative insight into the global alignment and local consistency of the estimated trajectories. Specifically, we utilize Absolute Trajectory Error (ATE) and Relative Pose Error (RPE) to measure the discrepancy between the estimated and ground truth trajectories, capturing both overall accuracy and the drift over time. Additionally, Frames Per Second (FPS) is employed to evaluate the real-time performance of the system. To further assess computational efficiency, we monitor system resource usage, including CPU and GPU utilization, memory footprint, and power consumption. Detailed definitions and the methodology for computing these metrics can be found in the Appendix. 4.3.4 Dataset For evaluation, we used the TUM RGB-D dataset [34], a widely recognized benchmark for SLAM systems. This dataset provides various indoor sequences captured using RGB-D cameras, which include both static and dynamic scenes. These sequences are ideal for testing the performance of SLAM algorithms in diverse and challenging real-world scenarios. To ensure diversity in our evaluation, we selected a specific subset of sequences that represent a wide range of environments, motions, and dynamics. The chosen sequences are as follows: •freiburg1_desk: This sequence captures normal movements within a static office environment, making it suitable for evaluating the performance of SLAM algorithms in controlled, steady indoor settings. •freiburg1_xyz: In this sequence, the camera undergoes translations along all three axes (X, Y, Z) while the environment remains static.
4.4 Experimentation 97 It evaluates the algorithm’s ability to handle structured translational movements. •freiburg2_xyz: This sequence involves rapid translations in a static indoor setting, challenging SLAM systems with faster motions, and requiring precise tracking in environments with minimal changes. •freiburg2_rpi: The camera performs quick rotations around the roll, pitch, and yaw axes within an indoor space. This sequence stresses the system’s ability to manage sudden rotational movements. •freiburg3_long_office_household: A longer sequence set in a domestic environment with various objects and changing lighting conditions. This provides a complex, real-world scenario with longer-term tracking requirements and varying conditions. •freiburg3_walking_xyz: This sequence captures rapid movements with significant translations and human dynamics, simulating more unpredictable and dynamic real-world conditions. •freiburg3_walking_static: Featuring fast movements within a static environment, this sequence tests the robustness of SLAM algorithms when confronted with high-speed camera motion while the scene remains unchanged. 4.4 Experimentation The experiments were conducted to analyze the performance of the selected SLAM algorithms under realistic conditions, with a focus on their accuracy, efficiency, and suitability for embedded platforms. 4.4.1 Performance evaluation Our study begins with a visual comparison of the accuracy in trajectory estimation across the selected localization systems, using sequences from the TUM RGB-D dataset. Figure 4.3 shows the trajectories for the freiburg3_structure_texture_far. For TSformer, three configurations are used, considering 1, 2, or 3 images to predict the camera motion. Table 4.1 & Table 4.2 summarize the average performance on selected sequences. Inference for models, whether trained end-to-end or integrated as a semantic thread, is performed using PyTorch (FP32), without utilizing TensorRT for performance optimization on NVIDIA hardware. VDO-SLAM demonstrates a balanced approach, achieving the best performance among
98 Challenges and Performance of SLAM Algorithms Figure 4.3 Trajectory predictions: each color denotes a different tested system. Table 4.1 Performance Metrics: overall localization accuracy (ATE), Error between successive poses (RPE) and Inference Time Methods ATE (cm) RPE (cm) FPS DeepVO 402.5 1.42 9.6 TSformer-VO-1 135.2 0.95 7.1 TSformer-VO-2 172.9 0.91 5.0 TSformer-VO-3 152.2 0.92 3.8 RDS SLAM (Mask RCNN) 3.4 0.96 3.2 RDS SLAM (SegNet) 3.3 1.00 7.5 VDO SLAM 3.4 1.00 8.1 the two SLAM systems tested, with an FPS of 8.1, reasonable energy consumption (∼10.85W), and a competitive ATE of 3.4 cm. However, a key factor driving VDO-SLAM’s efficiency is that semantic segmentation is handled in a pre-processing step (0% in the SLAM pipeline itself). This approach allows the system to focus the bulk of its computational
4.6 Appendix 105 [28] T. Peng, D. Zhang, D. L. N. Hettiarachchi, and J. Loomis, “An Evaluation of Embedded GPU Systems for Visual SLAM Algorithms,” Electronic Imaging, vol. 2020, no. 6, pp. 325–1325–6, Jan. 2020, doi: https: //doi.org/10.2352/issn.2470-1173.2020.6.iriacv-074. [29] S. Sumikura, M. Shibuya, and K. Sakurada, “OpenVSLAM: A Versatile Visual SLAM Framework,” doi: https://doi.org/10.1145/3343031.3350 539. [30] X. Liu, Y. Yang, B. Xu, and S. Schwertfeger, “Benchmarking SLAM Algorithms in the Cloud: The SLAM Hive System,” arXiv.org, 2024. https://arxiv.org/abs/2406.17586. [31] K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask RCNN,” arXiv.org, 2017. https://arxiv.org/abs/1703.06870. [32] V. Badrinarayanan, A. Kendall, and R. Cipolla, “SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation,” arXiv:1511.00561 [cs], Oct. 2016, Available: https://arxiv.org/ abs/1511.00561 [33] D. Sun, X. Yang, M.-Y. Liu, and J. Kautz, “PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2018, doi: https://doi.org/10.1109/cvpr.2018.00931. [34] J. Sturm, N. Engelhard, F. Endres, W. Burgard, and D. Cremers, “A benchmark for the evaluation of RGB-D SLAM systems,” 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, Oct. 2012, doi: https://doi.org/10.1109/iros.2012.6385773. [35] Z. Shan, R. Li, and S. Schwertfeger, “RGBD-Inertial Trajectory Estimation and Mapping for Ground Robots,” Sensors, vol. 19, no. 10, p. 2251, May 2019, doi: https://doi.org/10.3390/s19102251. 4.6 Appendix 4.6.1 Calculation of the metrics used in evaluation In order to evaluate the accuracy and performance of SLAM algorithms, we employ several key metrics that assess both the global alignment and local consistency of the estimated trajectories. The following metrics provide a comprehensive understanding of the system’s behavior, including Absolute Trajectory Error (ATE), Relative Pose Error (RPE), and Frames Per Second (FPS). In addition, we monitor system resource usage using hardware
106 Challenges and Performance of SLAM Algorithms statistics on CPU, GPU, memory, and power consumption to evaluate the efficiency on embedded platforms. 4.6.1.1 Absolute Trajectory Error (ATE) The Absolute Trajectory Error (ATE) is used to measure the global accuracy of the SLAM system by evaluating the difference between the predicted trajectory and the ground-truth trajectory after aligning them. Let the groundtruth poses be represented as a series of homogeneous transformation matrices Pgt i∈SE(3) where iindexes the sequence frames, and let the estimated poses be Pest i. The ATE is computed as the root mean square error (RMSE) between the positions of the estimated trajectory and the ground truth after aligning them through a rigid body transformation. The ATE is mathematically defined as: ATERMSE =v u u t1 n n X i=1 ∥tgt i−test i∥2,(4.1) where tgt iand test irepresent the translation vectors (positions) extracted from the ground-truth and estimated poses Pgt iand Pest i, respectively. The RMSE measures the Euclidean distance between corresponding positions in both trajectories, giving a global measure of accuracy. 4.6.1.2 Relative Pose Error (RPE) The Relative Pose Error (RPE) evaluates the local consistency of the trajectory by comparing the relative motion between consecutive poses in the estimated trajectory to that in the ground-truth trajectory. This metric is essential for assessing the short-term accuracy of the system in tracking small movements, which is critical in dynamic environments. Given two consecutive ground-truth poses Pgt iand Pgt i+1, the relative transformation between these poses is: Tgt i=Pgt i−1Pgt i+1.(4.2) Similarly, the relative transformation between consecutive estimated poses is: Test i=Pest i−1Pest i+1.(4.3) The RPE is then computed as the difference between the relative transformations of the ground-truth and estimated poses. The translational RPE is
4.6 Appendix 107 defined as the root mean square error (RMSE) of the translation vectors from the relative transformations: RPEtrans =v u u t1 n n X i=1 ∥tgt i−tgt i+△−test i−test i+△∥2,(4.4) where tgt iand test iare the translation components of the relative transformations Tgt iand Test iand △is the time interval. For the rotational RPE, we measure the angular difference between the relative rotation matrices Rgt iand Rest i,which can be quantified using the angle-axis representation. The rotational error is defined as: RPErot =1 nXn i=1 arccos tr Rgt i−Rgt i+△ TTRest i−Rest i+△ T−1 2 , (4.5) where Rgt iand Rest irepresent the rotation matrices of the ground-truth and estimated poses, tr denotes the trace of a matrix, and △is the time interval. This metric provides the average rotational error over the trajectory. 4.6.1.3 Frames Per Second (FPS) In real-time applications, the Frames Per Second (FPS) is a critical metric that measures the number of frames processed by the SLAM system per second. High FPS is essential for ensuring that the SLAM algorithm operates efficiently and in real time, especially in embedded or constrained environments where computational resources are limited. FPS can be computed as: FPS = number of frames total processing time.(4.6) This metric helps assess the speed and responsiveness of the SLAM system. 4.6.2 Alignment methods In this section, we describe four different trajectory alignment methods commonly used to compare predicted poses against ground-truth data in odometry and SLAM systems: scale, 6DOF, 7DOF, and 7DOF with scale. Each method focuses on optimizing different parameters (scale, rotation, and translation) to minimize the alignment error.
108 Challenges and Performance of SLAM Algorithms 4.6.2.1 Scale alignment (scale) The scale alignment adjusts only the global scale factor cbetween the predicted trajectory Xand the ground-truth trajectory Y, without modifying the rotations or translations. Method: • The objective is to minimize the distance between Xand Yby adjusting only the scale. • The optimal scale factor c is computed as: c=P(X . Y ) PX2.(4.7) • The aligned trajectory is then scaled as: Xaligned =c . X. (4.8) This method is particularly useful for monocular systems, where the scale is typically unknown and must be estimated post-facto. 4.6.2.2 6 Degress of Freedom (6DOF) The 6DOF alignment adjusts both the rotation and the translation between the predicted trajectory Xand the ground-truth Y, but keeps the scale fixed. This allows the correction of orientation and position errors. Method: • The six degrees of freedom include three for rotation and three for translation. • The alignment is based on the Singular Value Decomposition (SVD) of the covariance matrix between Xand. • The covariance matrix covxy is computed as: covxy =1 n n X i=1 (yi−meany)(xi−meanx)T.(4.9) • The SVD of covxy gives: covxy =UDV T.(4.10) where Uand Vare orthogonal matrices representing the principal directions of Yand X, respectively, and Dis a diagonal matrix of singular values.
4.6 Appendix 109 • The rotation matrix Ris computed as: R = UXVT.(4.11) where Σis a diagonal matrix ensuring a proper right-handed coordinate system. • The translation vector tis calculated as: t = meany−R.meanx.(4.12) • The final aligned trajectory is: Xaligned = R .X+t.(4.13) This method is ideal for LiDAR or stereo systems where scale is fixed but orientation and position errors need correction. 4.6.2.3 7 Degress of Freedom (7DOF) The 7DOF alignment extends the 6DOF method by also adjusting the scale factor c, in addition to the rotation and translation. This allows for simultaneous correction of scale, rotation, and translation errors. Method: • In addition to the rotation Rand translation t, the scale factor cis computed to minimize the alignment error. • The scale factor cis given by: c = 1 σx . tr(DΣ),(4.14) where σxis the variance of the points in X, and tr(DΣ) is the trace of the product of the singular values and the diagonal matrix Σ. • The transformed trajectory becomes: Xaligned = c .R.X+t.(4.15) This approach is beneficial in systems where the scale might not be exactly known, such as stereo or some lidar systems. 4.6.2.4 Scale + 7 Degrees of Freedom The scale_7DOF alignment combines the benefits of the 7DOF alignment with an additional scale optimization step. After applying the rotation and
110 Challenges and Performance of SLAM Algorithms translation, the scale factor is optimized separately to further minimize the final trajectory error. Method: • First, perform a 6DOF alignment to adjust the rotation Rand translation t. • Then, optimize the scale factor cindependently using the formula: c = P(Xaligned ·Y) P(Xaligned2).(4.16) • The final trajectory becomes: Xfinal =c·Xaligned.(4.17) This method provides a finer correction of scale after the rotational and translational alignment, making it useful when significant scale variations exist between predicted and ground-truth trajectories.
5 Designing Accelerated Edge AI Systems with Model Based Methodology Petri Solanti1and Russell Klein2 1Siemens EDA, Germany 2Siemens EDA, USA Abstract Deploying AI at the edge can be challenging. AI algorithms are very compute intensive. In the data centre, multiple large, power-hungry GPUs are often employed. However, edge systems typically have constrained compute capabilities and limited power. Further, many systems need to deal with size, weight, cost, thermal, and other limitations. Successfully deploying AI while meeting these limitations requires a holistic analysis of the system. A Model-Based Cybertronic System Engineering (MBCSE) methodology enables modelling and analysis of complex systems at a high abstraction level. It can be used to analytically find an optimal system architecture and hardware/software partitioning. Meeting the computational requirements may call for the development of bespoke machine learning accelerators. These are complex dedicated compute resources that deliver parallel computation, local data buffers, and some level of programmability. Designing an optimal accelerator architecture can be accomplished with an AI assisted High-Level Synthesis (HLS) process to efficiently explore the design space. This paper describes a Model Based Cybertronic System Engineering (MBCSE) methodology that can be used to craft a combined hardware/software implementation of an inferencing algorithm, balancing performance, power, cost, and other key design metrics. It begins with an algorithmic analysis, determining areas of significant complexity. This is followed by allocation of functions to physical computation elements, targeting 111
112 Designing Accelerated Edge AI Systems with Model Based Methodology both board-level and chip-level placement. During the allocation phase, complex algorithms may be mapped to bespoke accelerators that will be synthesized from the algorithmic description using high-level synthesis. Finally, an analysis of design is performed to ensure that all design metrics are met. Keywords: edge AI, system design, system optimization, high-level synthesis. 5.1 Introduction and Background Edge systems are often limited in several dimensions. Power consumption and compute capabilities are often limited in support of form-factor, weight, cost, mobility, and other requirements. This makes deploying AI on these systems challenging, as AI algorithms typically consume significant compute resource and power. One way to mitigate this is to architect the system with the compute resources that meet, but do not materially exceed, the requirements for the AI processing needed. AI processing can be performed on different types of compute resources. For example, inferences can be run on general-purpose processors, graphics processing units (GPUs), arrays of multiply/accumulate processing elements, FPGA fabric, or bespoke hardware accelerators. Each of these represents compromises for the system designer. For example, deploying AI on a general-purpose processor typically will require around 100 watts of power or more. This results in a heavy battery, limited battery life, and potential cooling issues. But it significantly eases the development efforts, as the same software and machine learning frameworks can be used in the edge design as was used by the data scientists in the data centre. A bespoke hardware accelerator will be orders of magnitude more efficient but requires a custom IC development effort. Developers need a way to understand the impact of their design tradeoffs early in the design cycle to find the optimal architecture for the system that addresses the myriads of design constraints imposed by the business, regulatory, competitive, and other forces influencing the design of the system. Yet, finding a suitable methodology that can address all the design needs is tricky. Model-based Systems Engineering has been used for software development for 25 years. Unified Modelling Language (UML) [1] was the first standardized graphical modelling language for specification, construction, documentation and visualization of software intensive systems. Because
5.2 Model Based Cybertronic Systems Engineering 113 UML was originally developed for modelling SW systems, it was not suitable for general system modelling. System Modelling Language (SysML) was based on UML but targeted more to the needs of general systems engineering. It is widely used for modelling system functionality but has severe capacity limitations. Systems with functionality implemented in both SW and bespoke HW blocks are called Cybertronics Systems [2]. These are difficult to model with UML or SysML, which are targeted to single-domain modelling. Finding the optimal HW/SW partitioning needs an extensive design space exploration and capability to analyse different metrics like processor and bus utilization, task latency, power consumption or network load. The modelling methodology must support clear separation of function from structure, function allocation to structural elements and mapping of structural elements to any target technology. One of the methodologies developed for this purpose is Architecture Analysis and Design Language (AADL) [3]. It is used to model the SW and HW architectures of embedded real-time systems. AADL is a useful methodology for HW/SW system modelling and analysis, but it is lacking capabilities to the operational and functional analysis and modelling of cyber-physical systems. 5.2 Model Based Cybertronic Systems Engineering Challenges of the cybertronics systems engineering are diverse. It begins with the system context that can be a network, a computing enclosure with multiple Printed Circuit Boards (PCB), a single PCB, a System-on-Chip (SoC), a Field Programmable Gate Array (FPGA) or embedded SW. The physical system sets the functional constraints that define the requirements for the cybertronics subsystems. Furthermore, the individual algorithms communicating with each other and consuming computing resources, can be allocated to different processing elements. The design space is huge and finding the optimal architecture is difficult. Another challenge is the variety of the design domains that are usually tightly coupled. An architectural component like PCB contains other architectural components that are cybertronics subsystems themselves such as 3DIC or SoC. SW functionality can be allocated onto multiple processors that are potentially in different subsystems. In such cases the network of data exchanges between the functions must be allocated to physical interconnect that can be a network segment, platform bus like PCIe, network-on-chip, or something similar, depending on the implementation technology.
114 Designing Accelerated Edge AI Systems with Model Based Methodology Third challenge is the fragmented organization of the domain specific development processes. Communication between the teams is negligible because the different terminologies and specifications are interpreted differently. This makes maintaining the integrity of the system difficult. Model Based Cybertronic Systems Engineering is a new model-based methodology developed for modelling and analysis of cybertronics systems. It borrows concepts from multiple modern systems engineering methodologies. The main target of this new methodology is to unify the modelling methodologies on different system levels and enable seamless communication between the architects in the different implementation domains and the design teams. The first methodology is called HW/SW co-architecting [4]. It is based on C++ function tree breakdown, where the functions are allocated to SW that is executed on one or more processors or HW accelerators that are implemented by using HLS. Because the C++ source code can be used both as SW code and input for the HLS, the mapping decisions can be done late in the design cycle. This approach works nicely for algorithms that are available, or can be translated into, C++, like streaming video or AI algorithms. The second methodology used in MBCSE is ARChitecture Analysis and Design Integrated Approach, ARCADIA [5]. It is a very simple static information model methodology for system modelling according to the INCOSE SE handbook. Arcadia can be used to model the functional, logical, and physical architecture of the system using standardized artifacts. The functions and structures are kept separated, but the allocation of functions, exchanges, and components to structures is supported. Arcadia methodology has no technology binding nor simulation capabilities itself, so the designer can specify any target technology or simulation by using properties. It also has capability to transition a component to a subsystem in a separate project. These three capabilities make Arcadia a perfect methodology for cybertronics system modelling. Yet, a comprehensive system modelling methodology must enable design space exploration with different types of simulations. Property Model Methodology [6] is a dynamic, simulation-based methodology for requirements driven system modelling and analysis. Simulations are needed to validate the functional correctness of the system model, to analyse the system performance of different architecture options and verify the implementation. PMM and Arcadia complement each other and form a basis for a cybertronics system modelling and validation methodology.
References 217 [53] M. Nagel, R. A. Amjad, M. Van Baalen, C. Louizos, and T. Blankevoort, “Up or down? adaptive rounding for post-training quantization,” in International Conference on Machine Learning. PMLR, 2020, pp. 7197– 7206. [54] M. Horowitz, “1.1 computing’s energy problem (and what we can do about it),” in 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC). IEEE, 2014, pp. 10–14. [55] J. Johnson, “Rethinking floating point for deep learning,” arXiv preprint arXiv:1811.01721, 2018. [56] D. L. N. Hettiarachchi, V. S. P. Davuluru, and E. J. Balster, “Integer vs. floating-point processing on modern fpga technology,” in 2020 10th Annual Computing and Communication Workshop and Conference (CCWC). IEEE, 2020, pp. 0606–0612. [57] D. Zhang, J. Yang, D. Ye, and G. Hua, “Lq-nets: Learned quantization for highly accurate and compact deep neural networks,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 365– 382. [58] E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh, “Gptq: Accurate post-training quantization for generative pre-trained transformers,” arXiv preprint arXiv:2210.17323, 2022. [59] B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, “Quantization and training of neural networks for efficient integer-arithmetic-only inference,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 2704– 2713. [60] A. Polino, R. Pascanu, and D. Alistarh, “Model compression via distillation and quantization,” arXiv preprint arXiv:1802.05668, 2018. [61] J. H. Lee, S. Ha, S. Choi, W.-J. Lee, and S. Lee, “Quantization for rapid deployment of deep neural networks,” arXiv preprint arXiv:1810.05488, 2018. [62] R. Banner, Y. Nahshan, and D. Soudry, “Post training 4-bit quantization of convolutional networks for rapid-deployment,” Advances in Neural Information Processing Systems, vol. 32, 2019. [63] M. G. d. Nascimento, R. Fawcett, and V. A. Prisacariu, “Dsconv: efficient convolution operator,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 5148–5157. [64] B. Darvish Rouhani, D. Lo, R. Zhao, M. Liu, J. Fowers, K. Ovtcharov, A. Vinogradsky, S. Massengill, L. Yang, R. Bittner et al., “Pushing the limits of narrow precision inferencing at cloud scale with microsoft
218 Recent Trends in Edge AI: Efficient Design, Training and Deployment floating point,” Advances in neural information processing systems, vol. 33, pp. 10 271–10 281, 2020. [65] G. Xiao, J. Lin, M. Seznec, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” arXiv preprint arXiv:2211.10438, 2022. [66] Y. Choukroun, E. Kravchik, F. Yang, and P. Kisilev, “Low-bit quantization of neural networks for efficient inference,” in 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW). IEEE, 2019, pp. 3009–3018. [67] P. Wang, Q. Chen, X. He, and J. Cheng, “Towards accurate posttraining network quantization via bit-split and stitching,” in International Conference on Machine Learning. PMLR, 2020, pp. 9847–9856. [68] Z. Yao, R. Yazdani Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He, “Zeroquant: Efficient and affordable post-training quantization for largescale transformers,” Advances in Neural Information Processing Systems, vol. 35, pp. 27 168–27 183, 2022. [69] Y. Bondarenko, M. Nagel, and T. Blankevoort, “Understanding and overcoming the challenges of efficient transformer quantization,” arXiv preprint arXiv:2109.12948, 2021. [70] Y. Cai, Z. Yao, Z. Dong, A. Gholami, M. W. Mahoney, and K. Keutzer, “Zeroq: A novel zero shot quantization framework,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 13 169–13 178. [71] J. Choi, Z. Wang, S. Venkataramani, P. I.-J. Chuang, V. Srinivasan, and K. Gopalakrishnan, “Pact: Parameterized clipping activation for quantized neural networks,” arXiv preprint arXiv:1805.06085, 2018. E. Park, S. Yoo, and P. Vajda, “Value-aware quantization for training and inference of neural networks,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 580–595. [72] M. Nagel, M. Fournarakis, R. A. Amjad, Y. Bondarenko, M. Van Baalen, and T. Blankevoort, “A white paper on neural network quantization,” arXiv preprint arXiv:2106.08295, 2021. [73] Y. Bengio, N. Lt’eonard, and A. Courville, “Estimating or propagating gradients through stochastic neurons for conditional computation,” arXiv preprint arXiv:1308.3432, 2013. [74] P. Yin, J. Lyu, S. Zhang, S. Osher, Y. Qi, and J. Xin, “Understanding straight-through estimator in training activation quantized neural nets,” arXiv preprint arXiv:1903.05662, 2019.
References 219 [75] I. Hubara, Y. Nahshan, Y. Hanani, R. Banner, and D. Soudry, “Improving post training neural quantization: Layer-wise calibration and integer programming,” arXiv preprint arXiv:2006.10518, 2020. [76] M. Nagel, M. v. Baalen, T. Blankevoort, and M. Welling, “Datafree quantization through weight equalization and bias correction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 1325–1334. [77] A. Finkelstein, U. Almog, and M. Grobman, “Fighting quantization bias with bias,” arXiv preprint arXiv:1906.03193, 2019. [78] E. Meller, A. Finkelstein, U. Almog, and M. Grobman, “Same, same but different: Recovering neural network quantization error through weight factorization,” in International Conference on Machine Learning. PMLR, 2019, pp. 4486–4495. [79] C. N. Silla and A. A. Freitas, “A survey of hierarchical classification across different application domains,” Data Min Knowl Disc, vol. 22, pp. 31–72, 2011. [80] S. Kiritchenko and F. Famili, “Functional annotation of genes using hierarchical text categorization,” Proceedings of BioLink SIG, ISMB, 01 2005. [81] K. Goetschalckx, B. Moons, S. Lauwereins, M. Andraud, and M. Verhelst, “Optimized hierarchical cascaded processing; optimized hierarchical cascaded processing,” IEEE Journal on Emerging and Selected Topics in Circuits and Systems, vol. 8, 2018. [Online]. Available: http: //www.ieee.org/publicationsstandards/publications/rights/index.html [82] S. Adams, R. Meekins, P. A. Beling, K. Farinholt, N. Brown, S. Polter, and Q. Dong, “Hierarchical fault classification for resource constrained systems,” Mechanical Systems and Signal Processing, vol. 134, p. 106266, Dec. 2019. [Online]. Available: https://linkinghub.elsevier. com/retrieve/pii/S0888327019304819 [83] S. Adams, T. Cody, P. A. Beling, S. Polter, and K. Farinholt, “Hierarchical classification for unknown faults; hierarchical classification for unknown faults,” 2020 IEEE International Conference on Prognostics and Health Management (ICPHM), 2020. [84] J. Wissing, S. Scheele, A. Mohammed, D. Kolossa, and U. Schmid, “Himledge – energy-aware optimization for hierarchical machine learning,” in Advanced Research in Technologies, Information, Innovation and Sustainability, T. Guarda, F. Portela, and M. F. Augusto, Eds. Cham: Springer Nature Switzerland, 2022, pp. 15–29.
220 Recent Trends in Edge AI: Efficient Design, Training and Deployment [85] A. Thomas, Y. Guo, Y. Kim, B. Aksanli, A. Kumar, T. S. Rosing, and U. S. Diego, “Hierarchical and distributed machine learning inference beyond the edge; hierarchical and distributed machine learning inference beyond the edge,” 2019 IEEE 16th International Conference on Networking, Sensing and Control (ICNSC), 2019. [86] F. Samie, L. Bauer, and J. Henkel, “Hierarchical classification for constrained iot devices: A case study on human activity recognition,” IEEE Internet of Things Journal, vol. 7, pp. 8287–8295, 9 2020, [87] J. Karjee, K. Anand, V. N. Bhargav, P. S. Naik, R. B. V. Dabbiru, and N. Srinidhi, “Split computing: Dynamic partitioning and reliable communications in iot-edge for 6g vision.” IEEE, 8 2021, pp. 233–240. [Online]. Available: https://ieeexplore.ieee.org/document/9590311/ [88] S. Teerapittayanon, B. McDanel, and H. T. Kung, “Branchynet: Fast inference via early exiting from deep neural networks,” vol. 0. Institute of Electrical and Electronics Engineers Inc., 1 2016, pp. 2464–2469. [89] T. Bolukbasi, J. Wang, O. Dekel, and V. Saligrama, “Adaptive neural networks for efficient inference,” in Proceedings of the 34th International Conference on Machine Learning - Volume 70, ser. ICML’17. JMLR.org, 2017, p. 527–536. [90] G. Huang, D. Chen, T. Li, F. Wu, L. van der Maaten, and K. Q. Weinberger, “Multi-scale dense networks for resource efficient image classification,” in International Conference on Learning Representations, 2017. [Online]. Available: https://api.semanticscholar.org/Co rpusID:3475998 [91] I. Leontiadis, S. Laskaridis, S. I. Venieris, and N. D. Lane, “It’s always personal: Using early exits for efficient on-device cnn personalisation.” Association for Computing Machinery, Inc, 2 2021, pp. 15–21. [92] Y. Li, Y. Wu, X. Zhang, J. Hu, and I. Lee, “Energy-aware adaptive multiexit neural network inference implementation for a millimeter-scale sensing system,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 30, pp. 849–859, 7 2022.
10 Scalable Sensor Fusion for Motion Localization in Large RF Sensing Networks Fetze Pijlman1,2 1Signify, The Netherlands 2Eindhoven University of Technology, The Netherlands Abstract RF sensing in wireless communication networks is a novel approach for motion detection, but it faces challenges in accurately localising motion which is crucial for confinement in lighting control use cases. A probabilistic model enables motion localisation through sensor fusion. However, in probabilistic models the posterior estimations do not scale well with large networks as likelihoods of all possible system states need to be computed. It will be demonstrated that variational Bayesian techniques offer attractive approximations to the posteriors where the approximations require computational resources that scale with the number of nodes. This method is of general interest to large networks as it models nonlocal effects through localised updates. Keywords: RF sensing, sensor fusion, probabilistic models, variational inference, confinement. 10.1 Motivation In a smart lighting system, one can control the lighting by sensing motion through the wirelessly connected lights (nodes) which is also known as RF sensing. For some background on RF sensing the reader is referred to Liu [1] and Wu [2]. Fluctuations in RSSI values, Received Signal Strength 221
222 Scalable Sensor Fusion for Motion Localization in Large RF Sensing Networks Indicator, between nodes are strong indicators of nearby motion. For lighting control, the system should be able to detect human arm motions and steps as specified by the NEN norms [3]. As major motions generally lead to large RSSI fluctuations, it becomes a challenge for lighting control to sense minor motions within a room while not being sensitive to major motions just outside the room. An example situation is given in Figure 10.1. Figure 10.1 Downlights and indicated node pairs that are monitored for RSSI fluctuations. On the left a person working in an office and on the right a person walking on a corridor. Probabilistic hidden Markov models are popular for motion sensing for lighting control. Typically, parameters such as motion rate during presence, average duration of presence, and probability of entering a room are introduced (see for example Papatsimpa [4]). These parameters can be used for propagating states and updating them with observations. For modelling the signal variations due to motions in the surroundings, it is natural to extend the probabilistic models by modelling crosstalk. Motion directly underneath a node pair is visible, but a fraction of said motion is also visible to nearby other node pairs. The introduction of crosstalk complicates Bayesian inference as the number of different presence states now scale with 2Nwhere Nis the number of presence areas. This essentially blocks the application of Bayesian inference for large networks. Another hurdle for large networks is the limited bandwidth for communication. In a mesh ZigBee network, the communication budget is 30-50 bytes per second per node, which limits the network size to about 50 nodes. The goal of this paper is to apply a probabilistic model for motion/presence localisation by monitoring fluctuations in RSSI values between nearby node pairs. Using a variational Bayesian method, one can approximate posteriors by maximising the so-called free energy. As we will see, the maximisation is an iterative process that only involves local interactions. This leads to a scalable approach in which the computational power
10.2 Spensor Fusion via a Probabilistic Model 223 and memory scale with the number of nodes. Moreover, the calculations can be distributed over the nodes themselves, enabling a truly decentralised approach. 10.2 Spensor Fusion via a Probabilistic Model For localising motion, a probabilistic model will be constructed (see Bishop [5], Murphy [6] and Morey [7] for an introduction to probabilistic models). The probabilistic model has the fluctuation of RSSI values of the node pairs as observational data X. By monitoring the fluctuations in RSSI values of a node pair one effectively obtains a motion sensor. In the remainder of the article node pairs will be referred to as sensors. The probabilistic model contains hidden states of the physics in the various areas. These hidden states describe at each time instant tfor each area: the presence s(t)and the motion m(t).In addition, at each time instant the visibility of a motion from an area ito a particular sensor jis described by Cij(t). In order to simplify the calculation, one would like to make the Markov assumption which influences the choice for the value of the time-step. A simple model can be obtained by choosing the time-step to be larger than 1 second in which case motion at time tonly depends on the presence state at tbut not on a previous motion state. As the RSSI fluctuations X(t) only depend on the m(t) and the probability that a motion is visible Cto a sensor, one can write the distribution as P(X(t),m(t),s(t),s(t−1) |C(t)) =P(X(t)|m(t), C(t)) Y i P(mi(t)|si(t))P(si(t)|si(t−1)) P(si(t−1)). Note that in the equation above it was assumed for simplicity that the presences in each area are uncorrelated. In Figure 10.2 an example of the interactions is given at a time instant. With the Markov assumption, the posterior can now be determined by Bayesian inference P(s(t),m(t), C (t)|X(:t)) (10.1) =P(X(t)|s(t),m(t), C (t)) P(X(t)) P(s(t),m(t), C (t)|X(:t−1)) ,
224 Scalable Sensor Fusion for Motion Localization in Large RF Sensing Networks Figure 10.2 Bayesian network at a time instant showing the dependence between the various states and the sensor observations. where X (: t) denotes all RSSI fluctuations up and including t, and X(t) denotes RSSI fluctuations at time t. Assuming the independence of presences in areas one finds for each area i P(si(t),mi(t)|X(: t−1)) = P(mi(t)|si(t)) P(si(t)|X(: t−1)) P(si(t)|X(: t−1)) = X si(t−1) P(si(t)|si(t−1))P(si(t−1) |X(: t−1)) . Note that the propagation of presence states involves a transition matrix P(si(t)|si(t−1)). The posterior in Equation 10.1 will be computed at each time-step using variational methods (see also Attias [8] and Jordan [9]). By maximising the free energy one can approximate the posterior P(s(t),m(t), C (t)|X(t)) by a probabilitty distribution qin terms of the KL-divergence DKL(P||q). This yields P(s(t),m(t), C (t)|X(t)) ≈argmax qX s(t),m(t),C(t) q(s(t),m(t), C (t))log P(X(t),s(t),m(t), C (t)) q(s(t),m(t), C (t)) . (10.2) For notational simplicity the explicit time-dependence will be dropped in the remainder.
10.3 Update Equations 225 In order to solve the equation above one often makes additional assumptions on the probability distribution q. However, the mean-field approximation such as q(s(t),m(t), C (t)) = q(s(t)) q(m(t)) q(C(t)) is blocked as probability for motion during absence is zero. Zero probabilities lead to problems with the logarithms. Instead, we approximate qby q(s,m, C) = Q i,j q(Cij|mi)q(mi|si)q(si),(10.3) where q(mi= 1|si= 0) = 0 and jrefers to the sensor index. For notational simplicity the difference between the q’s such as q(s1) and q(s2)has been made implicit. Note that this approximation contains more distributions compared to the mean-field approximation; q(mi|si= 0) has no relation with q(mi|si= 1). 10.3 Update Equations Equation 10.2 can be solved by iteratively maximizing the lower bound evidence. With each update for one of the q′sthe lower bound evidence gets increased. Please see Figure 10.3 for a subset of the network that is relevant for determining the states. Please note that moij is the visible motion towards sensor jexcluding any motion from mi. Figure 10.3 Isolated part of the network that is relevant for determining the states.
226 Scalable Sensor Fusion for Motion Localization in Large RF Sensing Networks In the following subsections the update equations for the various factors will be determined by plugging Equation 10.3 in Equation 10.2. In the derivation the following relation will be often employed argmaxq(a)P a q(a) (log P(X, a)−log q(a)) = P(X|a)P(a) Pa′PX|a′′P(a′). (10.4) The relation can be proven by using the method of Lagrange multipliers and treating Paq(a) = 1 as a constraint (see also Arfken [10]). It is instructive to work out the update equation in case a mean-field approximation would apply. Consider the example in Figure 10.4 and write q(s, t, u)as q(s)q(t)q(u). In this example the update equation for q(t) would be q(t) = argmaxq(t)P s,t,u q(s)q(t)q(u) log P(X,s,t,u) q(s)q(t)q(u) = argmaxq(t)P s,t,u q(s)q(t)q(u) logP(s|t)P(t|u)) q(s)q(t)q(u) = argmaxq(t)P t q(t) log Q s P(s|t)q(s)Q u P(t|u)q(u)−log q(t) =QsP(s|t)q(s)QuP(t|u)q(u) Pt′QsPs|t′′ ′q(s)QuP(t′|u)q(u) (10.5) Figure 10.4 Example network with a mean-field assumption.
11 Multi-Step Object Re-Identification on Edge Devices: A Pipeline for Vehicle Re-Identification Tomass Zutis, Peteris Racinskis, Anzelika Bureka, Janis Judvaitis, Janis Arents, and Modris Greitans Institute of Electronics and Computer Science (EDI), Latvia Abstract In modern computer vision tasks, the ability to identify and track objects across different scenes and environments has become important for numerous applications, especially in transportation. Inspired by this need, we propose a method that leverages a multi-step process focused on extracting and using object features for object re-identification.The proposed pipeline includes the following steps: detecting an object, converting its features into a vector embedding, storing this embedding in a vector database, and then querying the database to find the same or similar objects based on their feature embeddings. This approach enables us to identify the same object across different images or cameras, even in varying locations. This is essential in scenarios like Vehicle Re-Identification. For such scenario, implementing this process on edge devices is crucial. Therefore, ways to tailor the pipeline and its outputs for edge devices are outlined. The paper details the pipeline’s structure along with the experimental setup demonstrating its application, particularly in vehicle re-identification. The pipeline achieves 70-80% reidentification precision when dealing with vehicle images from our network cameras and above 70% Rank-1 accuracy when dealing with a CityFlow video track scenario. Keywords: vehicle re-identification, feature extraction, re-identification pipeline, computer vision, traffic monitoring. 233
234 Multi-Step Object Re-Identification on Edge Devices: A Pipeline for Vehicle 11.1 Introduction Object recognition in photos and videos has long been a key area of research, with significant advancement driven by computer vision [1]. Initially, image classification addressed the question, “What is in the image?” followed by object detection answering, “Where and what are the objects?”—largely thanks to Convolutional Neural Networks (CNNs) and their variants [2][3]. This paper focuses on object re-identification, which is a sub-problem of image retrieval. Object re-identification aims to distinguish instances of the same class, such as vehicles and persons (as illustrated in Figure 11.1), across different scenes despite changes in conditions like scene, lighting , or object pose [4][5]. However, truly impactful re-identification research in 2024 must support edge computing. Because of the increase in processing latency and big data, edge computing is becoming essential for real-time re-identification, especially in applications like traffic monitoring, where network cameras should transmit only processed results [6][7]. The volume of available video footage, and the influx of sensory data have made the large-scale accumulation of big data inevitable [8]. This is why fully automated systems are needed for processing data and re-identifying objects in smart cities. Manual processing by humans is not feasible, especially when real-time decisions are required. We propose a pipeline to handle these tasks of efficiently processing live video feeds and identifying objects across multiple scenes. The novelty of this research lies in the integration of state-of-the-art methods into a unified pipeline, tested on Figure 11.1 Re-identifiable classes in smart city environments.
11.2 Related work and state of the art 235 real-world vehicle re-identification scenarios that fit into smart city initiatives and traffic management using edge computing solutions. 11.2 Related work and state of the art 11.2.1 Object detection CNNs have been incredibly useful in computer vision tasks, including object detection. YOLO - the “You Only Look Once” model is one of the best performing and regularly updated choices. The latest YOLO v8 version has shown significant improvements in accuracy and speed [9], which is crucial for real-time applications like ours. 11.2.2 Object feature extraction Feature extraction maps an image from its colour space to a higherdimensional feature space [10]. Before feature extraction, multiple preprocessing stages are usually employed: normalization, thresholding, binarization, resizing and others. We can expect a model to extract colour, texture, shape, motion and localization features. A model learns intra-class variations as features when training a model on objects of one class like face features in the case of facial recognition. As of 2024, CNNs became the predominant choice in object feature extraction thanks to their strong representation power and their ability to learn deep invariant embeddings [11]. 11.2.3 Vehicle re-identification There are multiple benchmarks for Vehicle Re-ID. MBR4B-LAI model [12] tops the VeRi-776 benchmark, “A strong baseline” model [13] tops the CityFlow benchmark and the VehicleNet model [14] is best at the VeRi benchmark. We pay particular interest to [14] by Zheng et al. because of the baseline model that is applicable to all types of object re-identification and feature extraction. In [15] the same author first introduces us to their baseline model and its architecture and demonstrates its capabilities specifically in pedestrian re-identification. In further papers, however, Zheng et al. demonstrate tailoring of this baseline to vehicle re-identification [16] and person re-identification [11]. We underline this baseline models usefulness by the versatility of its use in publications, its open code base and customizability and its entry into most benchmarks. The model is successful in Rank1
236 Multi-Step Object Re-Identification on Edge Devices: A Pipeline for Vehicle precision, according to the benchmark results on Papers with Code [17]. The model is implemented in Pytorch and is based on ResNet50 pre-trained on ImageNet, although this backbone is customizable. We use the surrounding project code for this model available on [18]. The author has provisioned tools to train, finetune, test and visualize the results of the inference process. We further refer to this model as “baseline model”. 11.2.4 Available datasets The VehicleID introduced in [19] has each image associated with a vehicle ID. It is the dataset with one of the biggest unique ID collections and features pictures with different resolutions and quality of visibility as well as vehicles in motion state. Vehicles are mostly seen, however, from the front and the back only. The VeRi-776 was introduced in [20]. Each image is attached with vehicle ID, bounding box, type, colour, brand. The dataset has a smaller set of unique id’s compared to the VehicleID, but has better quality pictures, better visibility and vehicles from all angles not just back and front. The CityFlow dataset introduced in [21] is a traffic camera dataset consisting of synchronized HD videos from 40 cameras. The quality of images is high, and vehicles can be seen from many angles. The difference between this and the previous two widely used datasets can be seen in Figure 11.2. Figure 11.2 The difference between the distribution of vehicles for the (from the left) VehicleID, VeRi and CityFlow datasets.
11.3 Proposed methodology 237 The VehicleX synthetic data can supplement these three datasets. It contains generated images with domain adaptation from VehicleID, VeRi-776 and CityFlow [22]. 11.2.5 Edge implementation Requirements for a real-time system include implementing a network that follows the edge-computing paradigm. The edge-computing paradigm means that the video analytics run directly on the device and only the processed results and analytics are transmitted. [6] have developed a pilot project where they use mobility trackers using live CCTV feeds, with twenty sensors deployed over the city with the objective of citywide traffic monitoring in real-time. The devices had the ability to transmit the outputs either over Ethernet or LoRaWAN networks and had two main components: 1.) an NVIDIA Jetson TX2 high performance and power efficient embedded computing device with special units for accelerating neural network computations used for image processing and running Ubuntu 16.04 LTS and 2.) a Pycom LoPy 4 module handling the LoRaWAN communications. 11.3 Proposed methodology We propose a sequence of processes for object re-identification in the context of a smart city environment. It takes in video frames from camera x, detects the vehicles in the frames, crops images of the vehicles and saves them, turns the images into feature embeddings and saves them in a vector database. The same process is repeated for a different camera y ... z, so that vehicles can be re-identified from camera xto y ... z or vice versa. This has been illustrated in Figure 11.3. Figure 11.3 The proposed structure of the re-identification pipeline.
238 Multi-Step Object Re-Identification on Edge Devices: A Pipeline for Vehicle During development, we split the pipeline into two parts – Vehicle counting and tracking and Vehicle Re-Identification. 11.3.1 Vehicle detection, tracking and counting First step in the larger pipeline is object detection. We receive the videos from a network camera, detect the vehicles, get their bounding boxes and start tracking them. We use the YOLO v8 model for detection and the ByteTrack tracking package [23] to assign IDs to vehicles and track them through consecutive frames. Further we establish counting criteria for incoming and outgoing cars (for example entries and exits in an intersection). This can be done with the builtin functions of ByteTrack, for example, drawLine - counting when a car drives past a drawn line as seen in Figure 11.4. The count, entry and exit times of cars should be continuously logged. Before testing re-identification, verifying the accuracy of the vehicle counting step is essential. However, as this falls outside the paper’s scope, we have intentionally excluded these sections. Figure 11.4 The implementation of the counting lines and ByteTrack in cameras (from the left, upper row) 1.,3. and (lower row) 2. 11.3.2 Vehicle feature extraction and storage Consequently vehicles should be cropped out of the frame. Then the reidentification model will turn the object features into n-dimensional vectors
11.3 Proposed methodology 239 that consist of natural, real or complex numbers, where one number represents a feature or a part of a feature [24]. We require a Vector database, so vectors can be stored and queried efficiently [25]. The database must contain cosine similarity search complemented by metadata filters. The cosine similarity function is widely used and requires an input of at least two unit-length normalized vector inputs to output a vector distance [24]. We aim to store data in the form of Key:Value = ObjectId:ObjectFeaturesVector while more fields should be easy to add. For this we have chosen LanceDB – an opensource database for vector search, built for efficiency in handling vector data and integration with Python [26]. It is flexible in saving and querying data. We also introduce the following points of action: 11.3.2.1 Datasets We create our own “real-world” dataset for testing from footage we have gathered from our network cameras. In our custom dataset we aim to collect images of as many vehicles as the limited service-road traffic flow allows us to. We also aim to have a similar number of images per vehicle (i.e. 4-6 images, not more or less). These shots should be evenly distributed between far, medium and close distance and low, medium and high-resolution images respectively. We will also use the three widely used and public datasets that we’ve discussed in 2.4, to see which fits best for our real-world data and then combine the best performing standalone dataset with the Vehicle X synthetic data. 11.3.2.2 Training hyper-parameters The training parameters and their default values are found in [27], including Backbone, Learning rate, Warm epochs, Batch size, and Erasing probability. To further explore the model’s performance, we conducted experiments by varying additional hyperparameters, specifically: • Colour jitter: enabled or disabled; • Size of the last linear layer: 256, 512, or 1024; • Cosine learning rate: enabled or disabled; • Stride: 1, 2, or 3.
240 Multi-Step Object Re-Identification on Edge Devices: A Pipeline for Vehicle These variations were designed to evaluate the potential impact of these hyperparameters on training outcomes. 11.3.3 Edge device considerations While the pipeline has been tested on a desktop computer with an 8GB GPU, this setup provides a rough estimate of the computational load and performance we might expect on edge devices like the NVIDIA Jetson series [28]. Current work suggests that despite differences in power consumption and architecture, the constraints with desktop testing can offer insights for edge deployment that we wish to implement in the future. 11.4 Experimental settings 11.4.1 Receiving video from a Network camera Figure 11.5 The same car visible in our network cameras (from the left) 1., 2. and 3., respectively. To record “real-world” footage for the dataset we’ve envisioned in section 3.2.1, we will use 3 AXIS P1427-LE Network cameras [29] that record footage from the same service road and a parking lot. We are receiving the video in 1280×960 resolution with ∼2 fps. The view from the cameras illustrated in Figure 11.5. are summarized in Table 11.1. When a vehicle is at the gates, camera 1 and 2 will see the car from the opposite sides (front and rear). Camera 3 is located deeper into the territory and generally sees the path of a vehicle driving down the service road with the gate in a far distance.
11.4 Experimental settings 241 Table 11.1 The 3 network cameras used # Altitude (Above ground) Optimal Focal zone (Position, distance from cam.) Direction Sides of car seen 1. ∼3m centre of the frame, 4m →Front, Back, Sides and roof (for lower cars) 2. ∼3m further up the road from centre, 5-6 m ←More from the front and the back, but skewed sides and roof are visible 3. 6-8 m Wider area around centre, 6-8 m ←front/back and roof of the car visible well, sides in poor quality 11.4.2 Vehicle re-identification 11.4.2.1 Testing and data annotation We have chosen the following datasets to conduct our experiments on. •Benchmark datasets. By using the three benchmark datasets we discussed in section 2.4, we can access reliable test data, standardize our test metrics and evaluate the re-identification model itself. •CityFlow test track video. We test the whole re-identification part of the pipeline and simulate an intersection scenario where we are reidentifying vehicles with CityFlow test tracks. We will be using the scenario Nr. 1 (intersection S01) in this dataset, to re-identify vehicles from camera 1 to camera 4 [21]. There are around 2000 frames in each video and 91 unique vehicles seen. Both cameras point to the same intersection but from vastly different locations. •Custom test data. We have recorded footage from 3 of our cameras. All of them cover overlapping sections of a service road inside a closed territory. The vehicles were cropped from these videos and saved into 3 folders, each for its own camera. This dataset contains 70-100 images from each camera, with ∼25 unique vehicle identities. We will use this data to, first and foremost, test the generalisation of our trained models to the actual data that we will use this pipeline on. 11.4.2.2 Saving the feature extractions We will experiment with four methods (See their comparison in Table 11.2) for dealing with cropping vehicles from the frame and saving them into the database:
242 Multi-Step Object Re-Identification on Edge Devices: A Pipeline for Vehicle 1. Basic Frame-by-Frame Saving: In this method, a feature extraction of a vehicle is saved in every frame a vehicle is detected. These vectors are saved separately under the same vehicle ID in the database. Hence, there are multiple feature embeddings for the same vehicle. 2. Vector Summing: Instead of storing every feature vector separately, we maintain a single vector per vehicle. Each new embedding for a vehicle is summed with the existing vector, and the result is divided by the total number of updates, averaging the embeddings over time. This process keeps track of how many times the vector has been updated by adding an additional field in the database. 3. Zone-Based Saving: The frame is divided into zones using a grid, and a vehicle’s feature embedding is saved once per zone it passes through. The zones can be seen illustrated blue in Figure 11.6. Typically, this results in 4-6 saved vectors per vehicle, depending on its trajectory as opposed to many more vectors when saving in each frame. Each instance of feature extraction is stored as a separate vector under the same vehicle ID. 4. Zone-Based with Vector Summing: Similar to the previous method, but here we apply vector summing. The vehicle’s vector is summed for every frame in which a vehicle has changed zones. Figure 11.6 The implementation of the saving zones (drawn as blue rectangles) as seen on the CityFlow test track video.
11.7 Conclusion 249 Table 11.9 Testing methods of the re-identification pipeline on CityFlow video tracks Test scenario Basic Frame-by-Frame Saving: Vector Summing Zone-Based Saving Zone-Based with Vector Summing: Rank-1 accuracy in % 65.52 62.78 72.90 74.13 Test duration in seconds 1107.80 1085.30 330.12 328.90 as close to real world as possible we will use the Rank-1 accuracy to measure the accuracy of the pipeline, since in a scenario like this we simply care for whether the vehicle has been re-identified correctly or not. As mentioned in section 4.2.2 , we test the 4 approaches of capturing the feature embeddings and saving them into a database. Overall, the tests on CityFlow video tracks show that capturing vehicles in distinct zones improves both accuracy and efficiency, as clearly illustrated in Table 11.7. 11.6 Future research Although this paper proposes a complete pipeline for object re-identification on edge devices, multiple areas remain for future research. First, optimizing the pipeline for specific edge devices like the Nvidia Jetson by profiling performance and applying techniques such as model pruning, quantization, or TensorRT for efficient inference. Additionally, fine-tuning models trained on public datasets with our in-house data could reduce distribution shifts and improve real-world performance. Finally, a more in-depth analysis of feature vectors is needed to better understand which components are relevant for re-identification accuracy and which are redundant. 11.7 Conclusion In this paper, we presented a multi-step pipeline for object re-identification, focusing on real-time applications using edge devices. The pipeline handles object detection, feature extraction, and matching through a vector database, demonstrating reliable vehicle re-identification across various scenes. Performance dynamics were evaluated by comparing different datasets, models, pipeline processes. The authors observed the position and angle of cameras significantly influencing the accuracy of vehicle re-identification, with higher
250 Multi-Step Object Re-Identification on EdgeDevices: A Pipeline for Vehicle or wider vantage points producing better feature extractions for generalization across scenes. The findings demonstrate that while we successfully constructed an object re-identification pipeline by combining state-of-the-art methods, its effectiveness and constraints are highly dependent on the specific application scenario. Re-identification performance is influenced by factors such as the suitability of training data for real-world generalization, the model training approach, and the careful tuning of hyperparameters. For vehicle re-identification, it is crucial to consider where and when the vehicles are detected, the characteristics of the road or intersection, and the specific features of the camera recording the footage. From this, we can conclude that achieving high performance in vehicle re-identification requires not only advanced methodologies but also scenario-specific adaptations in data preparation, model optimization, and detection configuration. We also conclude that obtaining multiple embeddings of the same object in different poses and locations help re-identifying it later. Acknowledgements This work was supported by Chips Joint Undertaking EdgeAI project. The project EdgeAI “Edge AI Technologies for Optimised Performance Embedded Processing” is supported by the Chips Joint Undertaking and its members including top-up funding by Austria, Belgium, France, Greece, Italy, Latvia, Netherlands, and Norway under grant agreement No 101097300. References [1] S. J. Prince, Computer vision: models, learning, and inference. Cambridge University Press, 2012. [2] D. Lu and Q. Weng, “A survey of image classification methods and techniques for improving classification performance,” International Journal of Remote Sensing, vol. 28, no. 5, pp. 823–870, Mar. 2007, doi:https: //doi.org/10.1080/01431160600746456. [3] J. Du, “Understanding of Object Detection Based on CNN Family and YOLO,” Journal of Physics: Conference Series, vol. 1004, p. 012029, Apr. 2018, doi:https://doi.org/10.1088/1742-6596/1004/1/012029. [4] X. Li and Z. Zhou, “Object Re-Identification Based on Deep Learning,” IntechOpen eBooks, Jul. 2019, doi:https://doi.org/10.5772/intechopen.8 6564.
References 251 [5] R. Kuma, E. Weill, P. Sriram, and F. Aghdasi, “Vehicle re-identification: an efficient baseline using triplet embedding,” presented at the 2019 International Joint Conference on Neural Networks (IJCNN), IEEE, Jul. 2019, pp. 1–9. [6] J. Barthélemy, N. Verstaevel, H. Forehead, and P. Perez, “EdgeComputing Video Analytics for Real-Time Traffic Monitoring in a Smart City,” Sensors, vol. 19, no. 9, p. 2048, May 2019, doi: https: //doi.org/10.3390/s19092048. [7] C. Wang, Y. Yang, M. Qi, and H. Ma, “Efficient Cloud-edge Collaborative Inference for Object Re-identification,” arXiv preprint, vol. arXiv:2401.02041, 2024. [8] K. Cao, Y. Liu, G. Meng, and Q. Sun, “An Overview on Edge Computing Research,” IEEE Access, vol. 8, no. 1, pp. 85714–85728, 2020, doi: http s://doi.org/10.1109/access.2020.2991734. [9] “YOLOv8 Ultralytics: State-of-the-Art YOLO Models,” learnopencv.com, Jan. 10, 2023. https://learnopencv.com/ultralyt ics-yolov8/#YOLOv8-vs-YOLOv5 (accessed Sep. 23, 2024). [10] A. O. Salau and S. Jain, “Feature Extraction: A Survey of the Types, Techniques, Applications,” IEEE Xplore, Mar. 01, 2019. https://ieeexp lore.ieee.org/abstract/document/8938371 (accessed Sep. 23, 2024). [11] Z. Zheng, X. Yang, Z. Yu, L. Zheng, Y. Yang, and J. Kautz, “Joint Discriminative and Generative Learning for Person Re-identification.” Accessed: Jan. 10, 2025. [Online]. Available: https://openaccess.thecvf. com/content_CVPR_2019/papers/Zheng_Joint_Discriminative_and_G enerative_Learning_for_Person_Re-Identification_CVPR_2019_paper .pdf [12] E. Almeida, B. Silva, and J. Batista, “Strength in Diversity: MultiBranch Representation Learning for Vehicle Re-Identification,” arXiv (Cornell University), Jan. 2023, doi: https://doi.org/10.48550/arxiv.231 0.01129. [13] S. V. Huynh, N. H. Nguyen, N. T. Nguyen, Vinh Tq. Nguyen, C. Huynh, and C. Nguyen, “A Strong Baseline for Vehicle Re-Identification,” arXiv (Cornell University), Jun. 2021, doi: https://doi.org/10.1109/cvprw530 98.2021.00468. [14] Z. Zheng, T. Ruan, Y. Wei, and Y. Yang, “VehicleNet: Learning Robust Feature Representation for Vehicle Re-identification,” Computer Vision and Pattern Recognition, pp. 1–4, Jan. 2019.
252 Multi-Step Object Re-Identification on EdgeDevices: A Pipeline for Vehicle [15] Z. Zheng, L. Zheng, and Y. Yang, “A Discriminatively Learned CNN Embedding for Person Reidentification,” ACM Transactions on Multimedia Computing, Communications, and Applications, vol. 14, no. 1, pp. 1–20, Dec. 2017, doi:https://doi.org/10.1145/3159171. [16] Z. Zheng, T. Ruan, Y. Wei, Y. Yang, and T. Mei, “VehicleNet: Learning Robust Visual Representation for Vehicle Re-identification,” Apr. 2020. [17] “Papers with Code - Vehicle Re-Identification,” Paperswithcode.com, 2022. https://paperswithcode.com/task/vehicle-re-identification#b enchmarks (accessed Sept. 23, 2024). [18] layumi, “GitHub - layumi/Person_reID_baseline_pytorch: Pytorch ReID: A tiny, friendly, strong pytorch implement of person re-id / vehicle re-id baseline.,” GitHub, Jul. 25, 2022. https://github.com/l ayumi/Person_reID_baseline_pytorch/ (accessed Sept. 24, 2024). [19] H. Liu, Y. Tian, Y. Wang, L. Pang, and T. Huang, “Deep Relative Distance Learning: Tell the Difference between Similar Vehicles,” IEEE Xplore, Jun. 01, 2016. https://ieeexplore.ieee.org/document/7780607 [20] X. Liu, W. Liu, T. Mei, and H. Ma, “A Deep Learning-Based Approach to Progressive Vehicle Re-identification for Urban Surveillance,” Computer Vision – ECCV 2016, pp. 869–884, 2016, doi:https://doi.org/10.1 007/978-3-319-46475-6_53. [21] Z. Tang et al., “CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-Identification,” arXiv (Cornell University), Jun. 2019, doi:https://doi.org/10.1109/cvpr.2019.00900. [22] Y. Yao, L. Zheng, X. Yang, Milind Naphade, and T. Gedeon, “Simulating Content Consistent Vehicle Datasets with Attribute Descent,” Lecture notes in computer science, pp. 775–791, Jan. 2020, doi: https: //doi.org/10.1007/978-3-030-58539-6_46. [23] Y. Zhang et al., “ByteTrack: Multi-object Tracking by Associating Every Detection Box,” pp. 1–21, Oct. 2021, doi:https://doi.org/10.1 007/978-3-031-20047-2_1. [24] R. J. Bayardo, Y. Ma, and Ramakrishnan Srikant, “Scaling up all pairs similarity search,” The Web Conference, May 2007, doi:https://doi.org/ 10.1145/1242572.1242591. [25] T. Taipalus, “Vector database management systems: Fundamental concepts, use-cases, and current challenges,” Cognitive Systems Research, vol. 85, p. 101216, Jun. 2024, doi:https://doi.org/10.1016/j.cogsys.202 4.101216. [26] lancedb, “GitHub - lancedb/lancedb: Developer-friendly, serverless vector database for AI applications. Easily add long-term memory to your
References 253 LLM apps!,” GitHub, Sep. 24, 2024. https://github.com/lancedb/lance db (accessed Oct. 03, 2024). [27] regob, “GitHub - regob/vehicle_reid: Vehicle Re-identification,” GitHub, 2022. https://github.com/regob/vehicle_reid/tree/maste r(accessed Oct. 14, 2024). [28] S. Valladares, M. Toscano, R. Tufiño, P. Morillo, and D. Vallejo-Huanga, “Performance Evaluation of the Nvidia Jetson Nano Through a RealTime Machine Learning Application,” Advances in Intelligent Systems and Computing, pp. 343–349, 2021, doi: https://doi.org/10.1007/978-3030-68017-6_51. [29] “AXIS P1427-LE Network Camera - Product support | Axis Communications,” Axis.com, 2023. https://www.axis.com/products/axis-p1427-l e/support (accessed: September 26, 2024). [30] V. Shankar, R. Roelofs, H. Mania, A. Fang, B. Recht, and L. Schmidt, “Evaluating Machine Accuracy on ImageNet,” PMLR, pp. 8634–8644, Nov. 2020, Accessed: Sept. 26, 2024. [Online]. Available: http://procee dings.mlr.press/v119/shankar20c.html?ref=https://githubhelp.com [31] Mian Muhammad Talha, Hikmat Ullah Khan, S. Iqbal, M. Alghobiri, T. Iqbal, and M. Fayyaz, “Deep learning in news recommender systems: A comprehensive survey, challenges and future trends,” Neurocomputing, vol. 562, pp. 126881–126881, Dec. 2023, doi: https://doi.org/10.1016/j. neucom.2023.126881. [32] Jean-Charles Lamirel, Maha Ghribi, and Pascal Cuxac, “Unsupervised recall and precision measures: a step towards new efficient clustering quality indexes,” Aug. 2010.
12 A TinyMLOps Framework for Real-world Applications Mattia Antonini, Massimo Vecchio, and Fabio Antonelli Fondazione Bruno Kessler, Italy Abstract As devices become smarter, embedding intelligence in microcontrollers and constrained environments is critical. Optimising machine learning models for these tiny devices requires balancing software efficiency, such as accuracy, with hardware constraints like memory and power. We introduce a TinyMLOps-based framework for optimising models across the cloud-todevice continuum. In our approach, cloud resources handle heavy tasks like data labelling and model training, while microcontrollers gather real-time metrics on efficiency and hardware utilisation. Then, some repositories manage models and metadata identified during the optimisation phase, including performance metrics collected directly from the target devices, thus ensuring an accurate exploration of the model space in real-world conditions. Using tools such as MicroPython and MLFlow, our framework enables seamless AI deployment on resource-constrained devices, providing a scalable solution for the future of edge AI. Keywords: MLOps, edge AI, edge computing, model deployment, AI workflow, real-time metrics. 12.1 Introduction Artificial intelligence (AI) is becoming an invaluable companion in everyday life, present in wearables, smartphones, cars, and homes. We use AI to 255
256 A TinyMLOps Framework for Real-world Applications write, compose music, and edit pictures, making access almost effortless. However, this can obscure the real challenges of deploying AI models, especially on devices at the edge of the network [1]. These devices include single-board computers (e.g., Raspberry Pi, Coral Dev Kit, NVIDIA Jetson boards), which support full-fledged operating systems (OS) like Linux but are costly and energy-intensive. In contrast, microcontrollers (MCUs) offer lower computational power but are significantly cheaper and less demanding regarding energy and hardware resources. MCUs are ideal for specialised, real-time tasks in IoT environments but require adapted methodologies for their management [2]. The absence of a full-fledged OS and the constrained computational environment demand adjustments in orchestrating AI workflows on such devices. In this context, the TinyMLOps methodology presented in [3] can be fundamental in designing, deploying, and monitoring AI capabilities on constrained devices. From an operational standpoint, US-based companies like Roboflow [4], Edge Impulse [5], and Neuton.ai [6] provide platforms for designing, optimising, deploying, and executing AI pipelines at the network edge. While Roboflow focuses on computer vision problems, primarily supporting single-board computers, Edge Impulse specialises in creating and deploying machine learning models on resource-constrained devices like MCUs, making it ideal for IoT applications. However, it requires hardcoding models into firmware, updated via Over-the-Air (OTA) updates, which reduces flexibility in real-time scenarios. Similarly, the Neuton TinyML platform, provided by Neuton.ai, leverages a no-code approach to generate ultra-compact neural networks, optimising resource use for constrained devices. Its patented framework incrementally builds neural networks neuron by neuron, resulting in models significantly smaller than those from other frameworks. It enables deployment on devices with as little as 8-bit capacity, making it highly suitable for diverse IoT applications. This paper presents a TinyMLOps framework architecture designed to streamline AI workflows on MCUs, enabling seamless model deployment without embedding the model in the firmware. We describe the framework’s components, their role in supporting the methodology, and the various actors involved, from data labelling to model design, optimisation, firmware engineering, and operations. Furthermore, we provide a set of candidate stateof-the-art technologies that can be adopted to implement the framework and demonstrate its real-world feasibility. The rest of the paper is structured as follows. Section 12.2 gives an overview of the TinyMLOps methodology, then Section 12.3 presents the
12.2 TinyMLOps methodology 257 framework architecture by describing all the components. Section 12.4 provides a technological landscape for the framework. Finally, Section 12.5 concludes the paper. 12.2 TinyMLOps methodology The TinyMLOps methodology [3] evolves the MLOps methodology [7] to bring AI workflows and models to the edge and far edge of the network. Figure 12.1 The TinyMLOps Loop [3]. TinyMLOps particularly tackles the adaptation, deployment, and monitoring of models even on low-end MCUs (e.g., Raspberry Pi Pico, ESP32, Arduino, and STM32 families). The TinyMLOps approach is illustrated as a simple loop, as depicted in Figure 12.1. In more detail, the TinyMLOps loop starts with the TinyML circle (on the left side, highlighted in blue, Figure 12.1). It proceeds through a series of iterative steps involving collaboration between data scientists and operations engineers (on the right side highlighted in green in the Figure 12.1) to optimise and deploy AI models on edge devices. The steps are briefly introduced as follows. •Plan: data exploration and feature engineering to prepare and optimise data for model development; •Create: model training and feature adaptation tailored to the problem space; •Adapt & Optimize: the trained model is adapted and optimised for the target computing platform, typically through techniques such as quantisation and pruning;
258 A TinyMLOps Framework for Real-world Applications •Verify: ensures model compatibility and executability on the target platform, confirming that all operations are supported; •Package: the model is compiled or converted into a deployable format (e.g., TFLite) suitable for execution on edge devices; •Release: the model is deployed into the device memory, making it ready for inference; •Configure: run-time parameters may be adjusted based on platformand application-specific requirements (e.g., detection threshold); •Monitor: ongoing real-time surveillance of the model’s performance and health, ensuring its reliability post-deployment. If the model behaves unexpectedly, it triggers a new iteration of the entire loop. This closed-loop process facilitates seamless adaptation, deployment, and monitoring of TinyML models, ensuring efficient execution on resourceconstrained edge devices. 12.3 A TinyMLOps framework architecture As discussed above, the TinyMLOps methodology encompasses multiple abstract phases to deliver an efficient pipeline for deploying and maintaining machine learning models on edge devices. To implement these phases in realworld applications, we need to translate them into concrete components and pipelines within a software framework architecture. Figure 12.2 illustrates this architecture, which is divided into two sections corresponding to the two circles of the TinyMLOps loop (Figure 12.1). The left side (in blue) represents the components required for handling the TinyML-specific tasks, while the right side (in green) focuses on the components involved in the methodology’s operations phase. Figure 12.2 TinyMLOps Framework Architecture.
References 265 [7] M. M. John, H. H. Olsson, and J. Bosch, “Towards MLOps: A Framework and Maturity Model,” in 2021 47th Euromicro Conference on Software Engineering and Advanced Applications (SEAA), Palermo, Italy: IEEE, Sep. 2021, pp. 1–8. https://doi.org/10.1109/SEAA53835.2021.00050. [8] R. David et al., “TensorFlow Lite Micro: Embedded Machine Learning on TinyML Systems,” 2020, arXiv. https://doi.org/10.48550/ARXIV.201 0.08678. [9] M. Antonini, M. Pincheira, M. Vecchio, and F. Antonelli, “A TinyML approach to non-repudiable anomaly detection in extreme industrial environments,” in 2022 IEEE International Workshop on Metrology for Industry 4.0 & IoT (MetroInd4.0&IoT), Trento, Italy: IEEE, Jun. 2022, pp. 397–402. https://doi.org/10.1109/MetroInd4.0IoT54413. 2022.9831517.
13 Transfer and Self-learning in Probabilistic Models Fetze Pijlman1,2 1Signify, The Netherlands 2Eindhoven University of Technology, The Netherlands Abstract Transfer learning and self-learning are well known techniques that can separately improve probabilistic models. In this article it is investigated how to combine both methods in a single approach enabling transfer learning over different environments in which self-learning models are deployed. It is found that such learnings can be enabled through prior optimisation. Keywords: probabilistic models, Bayesian, online learning, self-learning, prior optimisation, transfer learning. 13.1 Motivation Probabilistic models are of interest for constructing classifiers or estimating parameters (see e.g. Murphy [1]). Examples of such applications are image classification in cameras and noise level estimation in radar sensors. These sensors are often deployed in diverse environments. This means that the underlying model could benefit from self-learning. Baye’s rule offers such a method P(θ|X) = P(X|θ)P(θ) P(X),(13.1) where θare the model parameters, X are the observations, and P(θ) is the prior distribution. Note that the observations X may also contain feedback 267
268 Transfer and Self-learning in Probabilistic Models Figure 13.1 On the left-hand side a prior was sent to various diverse environments where in each environment a posterior was estimated. On the right-hand side the question is how to choose the prior for a new unknown environment when sensor events from other environments are known. and interactions from users. A self-learning example is a motion sensor that self-learns its false positive rate. The prior distribution is our initial expectation on the parameters prior to the observation of new data. When no historical data is available, an expert choice should be made for the prior. However, in some situations data may be available from other diverse environments. The question here is how to choose a prior for a new unseen environment when data is available from other diverse environments. This problem is illustrated in Figure 13.1. In our motion sensor example, it may be believed that sensors in some applications have a false positive rate of once per week while in other applications the false positive rate is once per month. Note that for a new sensor deployment one cannot simply take an average of the posteriors from different environments as each posterior represents a different deployment: some posteriors may be very concentrated around a particular value which may slow down adaptation to a new environment and posteriors for different environments may have seen different amount of data (how to average that?). Moreover, averaging different posteriors is fundamentally wrong; similarly, one cannot take the average of an apple and an orange. To solve this problem Xuan [2] discusses a graphical model containing common and custom nodes. A challenge with such an approach is that the graph itself will also change as more data becomes available. Similarly, Suder [3] splits up the parameter space into common and custom parameters. A challenge with this approach is how to decide which parameters are common or which ones are not. Pautrat [4] discusses the grouping of similar environments, each having their own prior. A challenge with such an approach is
13.2 Prior Optimisation 269 how to decide on the number of groups and when to decide on forming a new group when a new environment is substantially deviating. Although the posteriors are fundamentally different for different environments, knowledge from different environments is still useful for determining an optimal prior for a new environment. Although the environments may be substantially different there is one aspect that they share in common which is the prior from which the probabilistic model could have started. This commonality offers an opportunity for transfer learning and it will be studied in this article. 13.2 Prior Optimisation A choice for prior is similar to a choice for model. When having a set of models Mithen the probability for a model to hold given the data is given by (see e.g. Murphy [1]) P(Mi|X) =P(X|Mi)P(Mi) P(X).(13.2) If a prior is parametrized through a hyperparameter µthen the best hyperparameter can be selected through argmaxµP(X|µ)P(µ).(13.3) In some cases, there may be no direct expectations for P(µ)as given in Equation 1.3 but there may be prior expectation for parameter θ. In those cases, an effective prior on µcan be derived. Prior to the knowledge on the prior P(θ)we take P(µ)to be uniform (first line equation below right-handside). The probabilities for observing a distribution P(θ)from µis simply the product of P(θi|µ)weighted by the amount P(θi) ∆θ(second line involves a product integral) P(µ|P[θ]) ∝P(P[θ]|µ) = lim ∆θ→0Y i P(θi|µ)P(θi)∆θ = exp ZdθP(θ) log P(θ|µ),(13.4) where in the last step the Geometric integral relation from Volterra was used. Note that the same result can be obtained by using variational Bayesian
270 Transfer and Self-learning in Probabilistic Models approximation. A variational approximation of a posterior is given by P(µ|X)≈argmaxq(µ)Rdµ q (µ)logP(X|µ)−log q(µ) P(µ). (13.5) Using the prior P(θ)for updating our knowledge on P(µ)gives P(µ|P[θ]) ≈argmax q(µ)Zdµq(µ)ZdθP (θ) log P(θ|µ)−log q(µ) P(µ) ∝exp ZdθP (θ) log P(θ|µ),(13.6) where in the last line we assumed the initial prior P(µ)to be flat. Equation 1.4 gives for the effective prior P(µ) P(µ|P[θ]) = exp RdθP(θ) log P(θ|µ) Rdµ′exp RdθP (θ) log P(θ|µ′)(13.7) In Equation 1.3, Xcontains all of the available data. It will be assumed that captured diversity in environments is representative for the underlying population of environments. The known environments are allowed to have different amounts of observations. As the known environments are assumed to be representative, an optimal prior for known environments will be an optimal prior for a new unknown environment. Taking the logarithm of Equation 1.3 and labelling data from environment iby Xione finds µ=argmaxµ[PilogP(Xi|µ) +logP(µ)] .(13.8) Using logP(X) = Rdθ q (θ)hlogP(X|θ)−logP(θ|X) P(θ)i,(13.9) with Rdθ q (θ) = 1,(13.10) one finds µ≈argmax µ"X iZdθiP(θi|Xi, µ)log P(Xi|θi) −log P(θi|Xi, µ) P(θi|µ)+ log P(µ).(13.11)
13.3 Example Categorical Distribution 271 In case the posterior for an environment cannot be exactly determined one can approximate by (see Attias [5], Morey [6], and Tran [7]) µ≈argmax µ max q(θi)"X iZdθiq(θi)log P(Xi|θi)−log q(θi) P(θi|µ) + log P(µ)#(13.12) Note that q(θi)approximates the posterior for each environment iin terms of the reversed KL-divergence. The reversed KL-divergence may lead to modes in q(θi)introducing errors. For this reason the posterior q(θi) should have sufficient freedom to describe all typical groups of customers. It is of interest to study the optimal prior in case of a single deployment (or in the limit that all deployments have the same environment) and for P(µ) being constant. If one has conjugate priors then one can iteratively solve Equation 1.11 by choosing at each iteration the prior P(θi|µ)to be equal to the posterior from the previous iteration P(θi|Xi, µ). In the limit of convergence, the logarithm containing ratio of posterior over prior will vanish and the posterior P(θi|X, µ)will strongly peak around µ= argmaxθiP(X|θi). In other words, in absence of any prior belief then the best prior for a new environment that is the same as the existing environment is a strongly peaked function around the most likely parameters (as expected). Note that strongly concentrated priors delay learnings from new data. Peaking of priors is reduced when a belief for P(µ)is included and when known environments are diverse. 13.3 Example Categorical Distribution For categorical distributions the conjugate prior is a Dirichlet distribution of which its hyperparameter is often denoted by αinstead of µthat was used in this paper upto now. Given applications isome observed categories j and assuming a flat prior such that the last term logP(µ)in Equation 1.11 disappears one obtains α= argmaxαPilogΓ(Pjαj)QjΓ(αj+nij ) Γ(Pjαj+nij )QjΓ(αj).(13.13) Consider the following example. In a first application one encounters 1000 items of category 0 and 2 items of category 1 while in a second
272 Transfer and Self-learning in Probabilistic Models application one encounters 10 items of category 0 and 2 items of category 1. Using the equation above one finds the optimal prior to be a Dirichlet distribution with α= (5.8,0.41). If the previous two applications are random draws from an underlying population of applications, then the obtained Dirichlet distribution would be optimal for a new application. The ratio of the hyperparameters indicates the expected ratio of classes, the sum of the hyperparameters is an indicator for the knowledge strength. Consider now another example in which one encounters in the first application 50 items of category 0 and 3 items of category 1, and in the second application 5 items of category 0 and 20 items of category 1. In this example the optimal hyperparameters for the Dirichlet distribution are α= (0.87,0.60). The low numbers indicate a lack of knowledge which is expected given the large differences in observations between the two applications. 13.4 Conclusions and Discussion Self-learning probabilistic models are useful for learning context and thereby self-learning increases performance in diverse environments. Similarly, transfer learning can help to use learned knowledge from one environment for another environment. In this article these two aspects are integrated into a single approach by optimising the prior. It was found that an optimal prior can be obtained by maximising a sum of evidences over environments. The amount of data per environment is allowed to vary. In order to optimise the prior, observations from different environments need to be shared. For cameras this means sharing of images, for radar sensors it means sharing of radar signals. Sharing of sensor events may involve privacy aspects. Although the approach leads to the best (or most optimal) prior one does need to realize that this approach carries the risk of overfitting. The situation is similar to variational Bayesian methods that yield the "best approximated" posterior which in reality may be far away from the true posterior. Acknowledgements EdgeAI “Edge AI Technologies for Optimised Performance Embedded Processing” project has received funding from Key Digital Technologies Joint Undertaking (KDT JU) under grant agreement No 101097300. The KDT
References 273 JU receives support from the European Union’s Horizon Europe research and innovation program and Austria, Belgium, France, Greece, Italy, Latvia, Luxembourg, Netherlands, Norway. References [1] K.P. Murphy, “Machine learning - a probabilistic perspective”, The MIT press (2012) [2] J. Xuan, J. Lu, G. Zhang, “Bayesian Transfer Learning: An Overview of Probabilistic Graphical Models for Transfer Learning”. arXiv:2109.13233 (2021) [3] P.M. Suder, J. Xu, and D.B. Dunson, “Bayesian Transfer Learning”. arXiv:2312.13484 (2023) [4] R. Pautrat, K. Chatzilygeroudis and J.B. Mouret, “Bayesian Optimization with Automatic Prior Selection for Data-Efficient Direct Policy Search”. arXiv 1709.06919 (2017) [5] H. Attias. “A variational bayesian framework for graphical models”. neurIPS (1999). [6] R. D. Morey, J.W. Romeijn, J.N. Rouder, The philosophy of Bayes factors and the quantification of statistical evidence, Journal of Mathematical Psychology (2016) [7] M.N. Tran, Trong-Nghia Nguyen, Viet-Hung Dao, A practical tutorial on Variational Bayes, arXiv: 2103.01327 (2021)
14.3 Hicnn Approach 281 problem which eventually also contributes to reducing energy consumption. HiCNN improves adaptability to varying machine conditions and eliminates reliance on manual feature engineering. This is achieved by defining and training three individual CNN architectures node-wise within a hierarchical tree structure. During the training, learned features from the classifiers in the initial hierarchy are forwarded to the lower classifiers. This approach ensures that the relevant feature vector is forwarded to the appropriate fault mode, thus preserving overall classification accuracy/performance. The CNN architectures of HiCNN are trained on raw data, and the deeper architectures are trained on features produced by CNN architectures lying in the earlier part of the hierarchy (Figure 14.1). As part of our empirical findings, a filter size of 5 in the initial conv layer of classifier C2 (Figure 14.2) ensures the availability of enough samples to train the classifiers located lower in the hierarchy. This type of feature forwarding of HiCNN helps us lower the total computations required by HiCNN during classification as we eliminate the need for more conv layers as classification progresses to a deeper hierarchy. 14.3.2 Feature forwarding As part of the HiCNN approach, anomaly detection functions as an event trigger within the HiCNN framework, prioritizing energy efficiency throughout its hierarchical structure by initiating classifications solely upon successful fault identification. The anomaly node architecture consists of four layers: one Conv1d (Convolutional 1-dimensional) layer, one max pooling layer, one batch normalization layer, and one full-connected layer. However, due to simpler task complexity, a classifier with one Conv1D layer is sufficient to perform the classification. This Conv1d layer consists of only one filter and a kernel size of 80. HiCNN architecture prioritizes anomaly detection by placing it as the parent classifier at the top of the hierarchical structure. This prioritization stems from data-dependent task complexity. [11]. Consequently, the HiCNN architecture employs a simple, single-layer CNN classifier for anomaly detection. This classifier acts as a gatekeeper, triggering further classifications only when an anomaly is identified. This approach is crucial for optimizing computational efficiency. By identifying anomalies early on, the HiCNN architecture avoids redundant computations in the lower levels of the hierarchy. In contrast, the baseline classifier continuously performs computations even in the state of no anomaly. This approach leads to unnecessary calculations and inefficient use of resources. This helps us
282 A Novel Hierarchical Approach to Perform On-device Energy Efficient Fault activate only a leg of the HiCNN architecture (Figure 14.2) ultimately helping to avoid using all the weights of trained architecture during inference as is the case with baseline CNN architecture. 14.3.3 Baseline CNN and Hierarchical CNN This subsection differentiates Baseline and HiCNN algorithms. A flat classification approach is represented by a single complex multilabel classifier responsible for classifying all classes at once. We constructed a baseline convolutional neural network (CNN) with a flat classification structure. This model is iteratively fine-tuned to achieve robust performance, yielding approximately 97 % accuracy on the test dataset. Subsequently, this baseline classifier serves as a benchmark for comparison against a hierarchical classification approach. The baseline model (Figure 14.2) requires a computationally intensive architecture to accommodate the one-hot encoded representation of 28 classes. This single architecture incorporates six convolutional layers and five dense layers, independently managing the classification task while demonstrating noteworthy performance. However, the baseline approach does not optimize for computational efficiency. Figure 14.2 Baseline Architecture vs HiCNN architecture, indicating towards flexible architectures and number of layers used. The HiCNN architectures are used with different filter size and filter numbers. On the other hand, as illustrated in Figure 14.1, if the initial mode corresponds to M1, only classifier M1 would be activated. Subsequent severity class determinations follow the same principle. This inherent flexibility enables the hierarchical model to manage complex classification problems by using specialized classifiers at each node. The training process within the HiCNN follows a node-wise approach, beginning with the Anomaly
14.4 Evaluation 283 node (Figure 14.1). This hierarchical strategy involves the fine-tuning of classifiers at each level of the HiCNN architecture after each training epoch. Employing specialized architectures for each classification task, every component is individually trained and fine-tuned. The initial architecture focuses on fault detection, followed by mode partitioning, fault classification, and severity classification. HiCNN promotes energy efficiency while preserving flexibility by adjusting layerspecific hyper-parameters, such as filter count and dimension. By controlling these parameters, we can regulate the volume of data propagated to the severity node, thus influencing the overall computational complexity of the HiCNN model. [13] To propagate relevant features to deeper nodes (i.e., severity classification), the mode partition CNN classifier at hierarchy level two is designed with two outputs. During training, a custom loss function outputting zero is employed for the first convolutional layer of mode classifiers (Figure 14.1) responsible for feature output. The overall accuracy of the HiCNN model is heavily influenced by the quantity of data samples provided to the Mode partition node. For training, individual architectures are connected in a distributed, treelike structure (Figure 14.1) using conditional logic (if-else) and then used to perform final inference on the test dataset. Beginning with the first instance of the test dataset, HiCNN dynamically selects the subsequent classifier based on predictions from the preceding stage. This process continues until a leaf node is reached, signaling the termination of the classification process. Edge devices offer limited memory space and the size of trained tensorflow models is large enough not to be able to be deployed on them. Tensorflow-Lite (TFLite) is a library offered by tensorflow to downscale the model’s optimum to be deployed on edge devices. Therefore, to perform inference on edge using HiCNN, initial thirty-one classifiers were converted to TF-Lite format to downscale. Conditional logic (if-else) is again utilized to run hierarchical inference on the edge device. 14.4 Evaluation In this section, contents are divided into two following sub-sections to discuss experimental setup and measurement. 14.4.1 Experimental setup As part of the data pre-processing pipeline, a custom data generator class is created to partition the dataset into 60:20:20 ratios for training, validation, and
284 A Novel Hierarchical Approach to Perform On-device Energy Efficient Fault testing. Before evaluation, input features were scaled to have zero mean and unit variance. A 512-sample window for training the HiCNN is determined to be optimal for balancing classification accuracy and latency, an important hyper-parameter for balancing energy efficiency and performance within the HiCNN architecture. Table 14.1 Comparison between Baseline and HiCNN Algorithm Accuracy Current drawn(mA) Baseline 98% 368.8 HiCNN 95% 50.8 The window size used to train the HiCNN is critical for model performance optimization, as an inappropriate window size could adversely affect the generalization capabilities of classifiers in HiCNN due to the smaller number of data samples available for model learning leading to underfitting. Notably, while reducing the window size to 256 samples does impact HiCNN’s classification accuracy (approximately 70 % performance), its energy efficiency remains the same. The smaller size of input samples to the initial architecture means the input fed to the following architectures would be scaled down proportionally due to the smaller kernel size of feature forwarding conv output layers, leading to a proportional reduction in performance due to underfitting in deeper models of hierarchy but no change in energy consumption while running inference. 14.4.2 Measurement All the classifiers of HiCNN are trained utilizing TensorFlow 2 libraries. Subsequently, the trained models are deployed to a Raspberry Pi using the Tensor-Flow lite library [12] for inference. A dedicated hardware setup is employed for precise energy consumption measurements, incorporating a data-logging multimeter. This configuration allowed for the ongoing measurement of electrical current consumption during the classification process recorded at 100 millisecond intervals. To account for the influence of background processes and sensors, the baseline power consumption of Raspberry Pi-2b is established during idle operation. This baseline power is then subtracted from the power consumption recorded during inference, yielding energy in Joules per inference, expressed in Joules. The Raspberry Pi is powered by a 5V DC power supply, with the multimeter interfaced for data acquisition (Figure 14.3). LabVIEW orchestrated the triggering of
14.5 Conclusion and Future work 285 Figure 14.3 Energy measurements setup using Raspberry pi connected to multimeter. classifications while simultaneously capturing current readings (in amperes) and corresponding timestamps (in milliseconds). The hyper-parameter window size 512 in HiCNN further guarantees HiCNN’s energy efficiency regardless of the model’s generalization capability. This could be beneficial for transfer learning applications. As a result, energy consumption stays constant with these variations. This occurs because the baseline algorithm performs unnecessary computations for non-anomaly cases, unlike HiCNN. The HiCNN architecture achieves a test accuracy of approximately 95%, slightly lower than the baseline model’s 98 %. However, HiCNN exhibits significant efficiency gains. TensorFlow Lite conversion and evaluation on desktop and Raspberry Pi reveal considerably faster inference times for HiCNN (380 microseconds vs. 1980 microseconds). Additionally, HiCNN demonstrates lower current draw (0.38 amperes vs. 0.46 amperes), (Figure 14.4 and Table 14.1) and reduced inference time (12 seconds vs. 86 seconds) for the classification task, consuming 197.8 Joules with baseline algorithm vs 22.8 Joules for HiCNN. This, along with lower energy per decision (0.002 joules vs. 0.0455 joules), translates to lower current draw and latency for HiCNN during every inference.
286 A Novel Hierarchical Approach to Perform On-device Energy Efficient Fault Figure 14.4 Current consumption during inference for Baseline algorithm vs Hierarchical algorithm indicating towards low latency inference by HiCNN. 14.5 Conclusion and Future work The HiCNN framework demonstrates promising results in attaining energy efficiency. It achieves up to 80% reduction in computational overhead compared to the baseline classifier according to per-decision energy consumption. This efficiency gain is attributed to controlled computations, ensuring decreased energy consumption with architectures devised per task complexity. These results emphasize the potential of algorithmically optimized approaches for embedded systems, enabling near-real-time monitoring with improvement in response latency. HiCNN is more suited to hierarchical datasets and even though HiCNN requires more memory, the reduced matrix operations could be accredited to the selective portions of the HiCNN network active during inference. A potential limitation arises when low data volume is propagated to lower hierarchical levels. This can lead to insufficient training data for severity classification, increasing the risk of model overfitting or underfitting in subsequent models due to inherent model complexity. Whereas HiCNN energy efficiency is completely independent of sample size, the energy consumption of the baseline model is independent of the
References 287 window size till a threshold. The Baseline algorithm’s performance remains relatively stable despite changes in the number of input neurons unless it drops significantly below a threshold, such as 64 neurons. Future research directions should include a comparison with manual feature engineering. Investigating the performance of HiCNN against methods employing manual feature extraction could provide insights into error propagation trade-offs. Another strategy could be hybrid cloud/edge deployment, which could help explore the potential benefits of partial algorithm/architecture deployment across cloud and on-device environments. References [1] “SIGCOMM Comput. Commun. Rev.,” vol. 13–14, no. 5–1, 1984. [2] D. Becking, “Compressed neural networks,” 2022. [3] E. Denton, “Networks for Efficient Evaluation L1-Norm Batch Normalization for Efficient Training of Deep Neural Networks,” n.d. [4] X. Jie, “Predictive Exit: Prediction of Fine-Grained Early Exits for Computationand Energy-Efficient Inference,” 2021. [5] I. Jungla, “Computation reduction in CNN inference by exploiting clustering and sparsity,” 2022. [6] J. Lin, “On-Device Training Under 256KB Memory,” 2022. [7] C. T. Liu, “Computation-performance optimization of convolutional networks with redundant kernel removal,” 2018. [8] S. Neupane, “Bearing Fault Detection and Diagnosis Using Case Western Reserve University Dataset With Deep Learning Approaches: A Review,” 2018. [9] T. Priyardishani, “Conditional Deep Learning for Energy-Efficient and Enhanced Pattern Recognition,” 2016. [10] J. Wißing, “HiML edge: An energy-aware optimization framework for hierarchical machine learning in wireless sensor systems at the edge,” 2022. [11] X. Song, J. Yao, Y. Gu, and J. Cao, “A Denoising Autoencoder-Based Bearing Fault Diagnosis System for Time-Domain Vibration Signals,” 2021. [12] R. Piramuthu, V. Jagadeesh, Z. Yan, and H. Zhang, “HDCNN: Hierarchical Deep Convolutional Neural Networks for Large Scale Visual Recognition,” 2015. [13] F. Zhou, “A Novel Multimode Fault Classification Method Based on Deep Learning,” 2022.
15 Discovering and Classifying Digital and Wooden Industries Products’ Defects at the Edge by a Yolo/ResNet-based Approach and Beyond Robin Faro1, Alessandro Strano1, and Francesco Cancelliere2 1Deepsensing SRL, Italy 2University of Catania, Italy Abstract The paper aims to present an AI-based Automated Optical Inspection (AOI) software for both digital and wooden industries developed within the EdgeAI project. Current approaches rely on centralized solutions, where the computation is performed inside the inspection machine itself. Instead, we present algorithms that work at the edge to give rise to competitive solutions to existing ones. In particular, we experiment with two different tasks of defect identification: detecting the defect position within a wooden panel by using YOLO, in which a 96% accuracy is reached. Secondly, concerning the digital industry, we perform a two-step classification between defective and nondefective microchips and then between four possible defect classes in their surface exploiting a ResNet network and obtaining a 97% accuracy. We also exploit explainability tools to understand which parts of the images caused the model’s decision. After developing the AI models we port them to two less power-consuming edge devices, Nvidia Orin Nano, and Nvidia Orin AGX, observing unchanged performance. 289
290 Discovering and Classifying Digital and Wooden Industries Keywords: automated optical inspection, edge computing, key performance indicators, deep learning, computer vision, PCBA defect detection, wooden product defect detection, NVIDIA Jetson ORIN. 15.1 Introduction The paper aims to present an AI-based AOI software developed within the EdgeAI project for defect detection of the PCBAs used in the digital industry to be implemented at the edge with the main expected outcome of being a viable and cost-effective solution for the inspection of many industrial products, not only digital boards. In particular, these algorithms are at the core of the AOI solution shown in Figure 15.1 where visual testing is done at the edge and learning in the cloud. This solution, as illustrated in [1], can reduce purchase and power consumption costs without increasing latency. Indeed, in this architecture, learning is done on the cloud but several tests suitable for highlighting groups of defects can be performed in parallel by competing boards thus decreasing latency. A solution similar to the one adopted in the EdgeAI project has been proposed by Advantech where the NVIDIA board is replaced by the MIC-72 but without discussing the algorithms used in practice for defect detection. On the contrary, several algorithms have been proposed in the literature to manage the abovementioned problems using GPUs. From the literature, we found that these are mainly optimized versions of the DL-powered YOLO algorithm [2]. The mAP (mean Average Precision) of the latest proposed algorithms goes beyond 99%, i.e. 99.17% in [3], 99.5% in [4] , and 99.71% Figure 15.1 An AOI solution consisting of an edge board for testing and a GPU server for learning.
15.4 Spotting Defects in Digital Industry Products 297 Figure 15.4 Example of prediction on validation set. (a) Losses on train and validation sets (b) Precision, Recall and Mean Average Precison Figure 15.5 Training metrics of YOLOv8 model. 15.4 Spotting Defects in Digital Industry Products In this section, we focus on the identification and classification of chip surface defects. Given the challenges associated with acquiring a comprehensive dataset, we employed the one introduced by Wang et al. [25], which, despite its modest size, provides valuable insights into the defect detection process.
298 Discovering and Classifying Digital and Wooden Industries 15.4.1 Defect Detection and Classification Dataset The selected dataset includes a total of 2763 images of chip surfaces, where 2000 images do not contain any defects, and 763 images present one among four possible defects. We implemented a two-step classification approach, aimed to streamline the industrial process and enhance the efficiency of quality control. The initial phase involves a binary classifier designed to quickly and accurately filter out defect-free chips. This step serves as a critical screening process, ensuring that only chips identified as non-defective continue through the subsequent stages of the industrial process. By doing so, we effectively reduce the computational load and focus further inspection efforts only on potentially problematic chips. This approach is particularly valuable in high-throughput manufacturing environments, where minimizing delays and optimizing resource allocation are crucial. For the chips flagged as defective in the first step, a more granular analysis is then conducted in the second step. This phase involves classifying the specific type of defect present on the chip surface. The detailed classification not only aids in determining the exact nature of the defect but also provides essential insights that can be used for root cause analysis. In particular, the dataset contains the following defect classes: • NO_DIE (6b) represents a unique defect scenario where the actual error lies in the absence of the soldered chip: the chip is missing on the substrate. The absence of a chip can be attributed to manufacturing faults, such as misalignment during the chip placement process or incorrect soldering. These errors can occur due to equipment malfunction, human error, or process inconsistencies. • DIE_INK (6c) includes chips with internal ink stains: these defects can arise from ink deposition errors or contamination during the manufacturing process. Detecting and classifying these internal defects is essential for ensuring the integrity and reliability of the chip’s functionality. • DIE_BROKEN (6d) involves chips with visible breaks or fractures along their edges. These breaks can occur during the manufacturing process or as a result of external factors. Detecting and accurately localizing these broken edges is crucial for quality control and identifying potential manufacturing issues.
15.4 Spotting Defects in Digital Industry Products 299 (a) DEFECT_FREE (b) NO_DIE (c) DIE_INK (d) DIE_BROKEN (e) DIE_CRACK Figure 15.6 Defect-free chip and four common chip surface defects Table 15.1 Distribution of the dataset DEFECT_FREE NO_DIE DIE_INK DIE_BROKEN DIE_CRACK 2000 100 135 42 486 • DIE_CRACK (6e) represents chips with cracks occurring internally within the chip structure. These cracks can stem from stress during fabrication, handling, or environmental factors. Detecting and characterizing these cracks aids in identifying structural weaknesses and preventing potential chip failures. The distribution of the different classes in the dataset is the following (1): 15.4.2 Experiments and Results The model we used to perform classification is ResNet [26], in particular the version 18 layers deep (namely ResNet-18). It introduced the innovative concept of “skip connection", which connects the activations of a given layer to further layers by skipping some intermediate layer, thus alleviating the issue of vanishing gradient. The binary classifier was trained for 70 epochs and obtained a 94.5% accuracy and 94% recall, precision, and F1 score. The progress of accuracy and loss function on train and validation is shown in Figure 15.7 a As regards the second phase classifier, it was trained for 80 epochs obtaining 97% accuracy, 87% recall, and 89% F1 score, as shown in Figure 15.8 a
300 Discovering and Classifying Digital and Wooden Industries (a) Loss and accuracy on train and validation sets (b) Confusion Matrix: 0 Defect Free - 1 Defective Figure 15.7 Metrics’ trends during training and test result. (a) Loss and accuracy on train and validation sets (b) Confusion Matrix: 0 NO_DIE - 1 DIE_INK - 2 DIE_BROKEN - 3 DIE_CRACK Figure 15.8 Metrics’ trends during training and test result 15.4.3 XAI Analysis: insigths into ResNet-18 using Grad-CAM Explainable AI methods, often referred as XAI, are essential for enhancing the transparency and interpretability of deep learning models by providing insights into their decision-making processes. Grad-CAM (Gradientweighted Class Activation Mapping) [27] visualizes the regions of an input image that most influence the model’s predictions by highlighting the gradients flowing into the convolutional layers, thus allowing us to identify the specific areas of the chip surface on which the model focuses. Below, we provide examples of the Grad-CAM results for each defect class: • NO_DIE: The activation maps effectively highlight the missing chip region, providing a clear indication of the defect (Figure 15.9 a).
15.5 Porting of the Models on Edge Devices 301 (a) NO_DIE (b) DIE_INK (c) DIE_BROKEN (d) DIE_CRACK Figure 15.9 Grad-CAM activation maps for different chip surface defect classes. • DIE_INK: The maps emphasize the areas surrounding ink stains, accurately identifying the relevant regions for defect classification (Figure 15.9 b). • DIE_BROKEN: It’s the least represented class and in fact the maps show some uncertainty in pinpointing the exact broken regions, reflecting challenges observed in the confusion matrix and indicating potential areas for model improvement (Figure 15.9 c). • DIE_CRACK: The activation maps clearly outline the crack, demonstrating the model’s proficiency in localizing and detecting this defect (Figure 15.9 d). 15.5 Porting of the Models on Edge Devices Porting AI models to edge devices is a critical step in realizing efficient, realtime solutions, especially in industrial settings. The edge devices used in this study include Nvidia Orin Nano and Nvidia Orin AGX, which are specifically chosen for their balance between computational power and energy efficiency. The process of porting models, such as YOLO for defect detection and ResNet for classification, involved converting the models to ONNX (Open Neural Network Exchange) format to ensure they could run efficiently on these resource-constrained devices while maintaining high accuracy. ONNX provides a unified format that allows models to be optimized for different hardware platforms, making them more portable and reducing dependency
302 Discovering and Classifying Digital and Wooden Industries on specific frameworks. By converting to ONNX, we were able to leverage hardware acceleration features available on our tested edge devices, in which we observed improved inference speed and reduced latency. The deployment tests demonstrated that both devices could perform realtime inference, with the Orin AGX showing a higher throughput due to its superior GPU capabilities. However, the Orin Nano, being more costeffective and energy-efficient, also provided satisfactory performance for applications where slightly lower throughput was acceptable. The deployment experiments showed that the models could achieve near-identical accuracy compared to their performance in a centralized server environment, proving the viability of using edge devices for industrial defect detection. In addition, an important aspect of edge deployment is ensuring low latency and robustness under varying conditions. The edge devices managed to maintain high accuracy with inference times well within the acceptable range for real-time operations, demonstrating their suitability for on-site, autonomous quality inspection. By processing data locally, the system also ensures data privacy, which is crucial in industries dealing with sensitive information, such as PCB industries that deal with pre-production boards. Overall, porting the models to edge devices proved successful, providing a viable solution for decentralized, real-time defect detection in industrial applications. The performance metrics, as shown in Tables 15.2 and 15.3, highlight the trade-off between inference speed and power consumption for each device. The Orin AGX, with its higher power consumption, offers better inference times, while the Orin Nano provides a more energy-efficient alternative suitable for less demanding applications. The use of powerful yet efficient edge devices like the Orin Nano and Orin AGX resulted in a system capable of high performance under the constraints typical of edge computing environments. Table 15.2 Performance of Nvidia Orin Nano for model deployment Model Inference Time (ms) Power Consumption (W) Chip Defect/No Defect 25.1 11.2 Chip Defect Classification 64.5 13.4 Wood Defect Classification 72.3 14.2 Table 15.3 Performance of Nvidia Orin AGX for model deployment Model Inference Time (ms) Power Consumption (W) Chip Defect/No Defect 22.0 15.6 Chip Defect Classification 44.4 21.5 Wood Defect Classification 50.4 22.5
15.6 Conclusions and Future Works 303 15.6 Conclusions and Future Works In this paper, we presented an AI-based Automated Optical Inspection solution designed for both the digital and wood industries, focusing on defect detection at the edge. Through a series of experiments, we explored the application of YOLOv8 for wood defect detection and ResNet for chip surface defect classification. The results demonstrated that our proposed approach achieved high accuracy, precision, and recall, proving the effectiveness of deep learning models in industrial defect detection tasks. We also presented an explainability analysis of the ResNet prediction, which can help in understanding the region of the image that led to the model’s decision. Moreover, by porting the models to edge devices such as the Nvidia Orin Nano and Orin AGX, we successfully validated the feasibility of running these complex models in a resource-constrained environment without compromising performance, enabling real-time defect detection. We showed that the edge deployment maintains unvaried the accuracy of the models while reducing latency and power consumption, which are critical aspects in the industrial setup. Future works will focus on several aspects to enhance our solution, for example: • The next step involves integrating our solution into a real-time industrial setup. This will involve testing the models in environments where objects move on conveyor belts, and high-speed cameras capture images of products for defect detection. Such real-world trials will help assess the models’ ability to handle real-time constraints, such as varying lighting conditions, motion blur, and differences in defect types and positions. By doing so, we aim to further optimize the system for seamless operation in an actual production line. • While our models achieved high accuracy, future efforts will focus on enhancing the quality of training data. This involves not only increasing the quantity of data but also ensuring that it represents a wide range of defects across different types of products and environments. High-quality, diverse data is critical for improving the generalization capabilities of the models, reducing false positives and negatives, and adapting to unseen defect types. In particular, augmenting the dataset with hard-to-detect defects and edge cases will be crucial for improving the system’s robustness. • Facing directly the PCB issue by generating a dataset suitable for detecting defects in these boards, either real or synthetic. Successively,
304 Discovering and Classifying Digital and Wooden Industries testing both the one-stage and the two-stage solution presented in the paper. In particular, in the second case, we imagine a first model able to detect each PCB internal component (i.e. resistors, capacitors...) and then several “expert" classifiers that will be able to decide whether a component is correctly mounted or not. • We aim to explore multi-task learning approaches where the models can detect and classify defects from multiple industries simultaneously, improving the efficiency and flexibility of the AOI system. Acknowledgements This work was supported by Chips Joint Undertaking EdgeAI project. The project EdgeAI “Edge AI Technologies for Optimised Performance Embedded Processing” is supported by the Chips Joint Undertaking and its members including top-up funding by Austria, Belgium, France, Greece, Italy, Latvia, Netherlands, and Norway under grant agreement No. 101097300. References [1] AAEON. AI@Edge: AI Vision in Automated Optical Inspection. Available at: http://www.aaeon.com/ru/ai/boxer-6841m-aoi-app-story.2023. http://www.aaeon.com/ru/ai/boxer-6841m-aoi-app-story. [2] Juan Terven, Diana-Margarita Córdova-Esparza and Julio-Alejandro Romero-González. “A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS”. in Machine Learning and Knowledge Extraction: 5.4 (november 2023), 1680–1716. issn: 2504-4990. 10.3390/make5040083. http://dx.doi.org /10.3390/make5040083. [3] JiaYou Lim, JunYi Lim, Vishnu Monn Baskaran and Xin Wang. “A deep context learning based PCB defect detection model with anomalous trend alarming system”. in Results in Engineering: 17 (2023), page 100968. issn: 2590-1230. https://doi.org/10.1016/j.rineng.202 3.100968. https://www.sciencedirect.com/science/article/pii/S2590123 023000956. [4] Yinchao Du, Jiangpeng Chen, Han Zhou, Xiaoling Yang, Zhongqi Wang, Jie Zhang, Yuechun Shi, Xiangfei Chen and Xuezhe Zheng. “An automated optical inspection (AOI) platform for three-dimensional (3D) defects detection on glass micro-optical components (GMOC)”.
References 305 in Optics Communications: 545, 129736 (october 2023), page 129736. 10.1016/j.optcom.2023.129736. [5] Hongjin Zhu, Lina Xing, Honghui Fan and Tao Wu. New PCB Defect Identification and Classification Method Combining MobileNet Algorithm and Improved YOLOv4 Model.april 2022. 10.21203/rs.3.rs-154 4671/v1. [6] Shaojun Song, Junfeng Jing, Yanqing Huang and Mingyang Shi. “EfficientDet for fabric defect detection based on edge computing”. in Journal of Engineered Fibers and Fabrics: 16 (april 2021), page 155892502110083. 10.1177/15589250211008346. [7] P. Cavalin, L. S. Oliveira, A. L. Koerich and A. S. Britto. “Wood Defect Detection using Grayscale Images and an Optimized Feature Set”. in IECON 2006 - 32nd Annual Conference on IEEE Industrial Electronics: 2006, pages 3408–3412. 10.1109/IECON.2006.347618. [8] Selman Jabo. “Machine vision for wood defect detection and classification”. in M.S Thesis, Chalmers University of Technology: (2011). [9] Ting He, Ying Liu, Chengyi Xu, Xiaolin Zhou, Zhongkang Hu and Jianan Fan. “A Fully Convolutional Neural Network for Wood Defect Location and Identification”. in IEEE Access: 7 (2019), pages 123453– 123462. 10.1109/ACCESS.2019.2937461. [10] Wei-Han Lim, Mohammad Babrdel Bonab and Kein Huat Chua. “An Optimized Lightweight Model for Real-Time Wood Defects Detection based on YOLOv4-Tiny”. in 2022 IEEE International Conference on Automatic Control and Intelligent Systems (I2CACIS): 2022, pages 186– 191. 10.1109/I2CACIS54679.2022.9815274. [11] Wenqi Cui, Zhenye Li, Anning Duanmu, Sheng Xue, Yiren Guo, Chao Ni, Tingting Zhu and Yajun Zhang. “CCG-YOLOv7: A Wood Defect Detection Model for Small Targets Using Improved YOLOv7”. in IEEE Access: 12 (2024), pages 10575–10585. 10.1109/ACCESS.2024.3352 445. [12] Rijun Wang, Yesheng Chen, Fulong Liang, Bo Wang, Xiangwei Mou and Guanghao Zhang. “BPNYOLO: A Novel Method for Wood Defect Detection Based on YOLOv7”. in Forests: 15.7 (2024). issn: 1999-4907. 10.3390/f15071096. https://www.mdpi.com/1999-4907/15/7/1096. [13] Zhichao Liu and Baida Qu. “Machine vision based online detection of PCB defect”. in Microprocessors and Microsystems: 82 (2021), page 103807. issn: 0141-9331. https://doi.org/10.1016/j.micpro.2020.1038 07. https://www.sciencedirect.com/science/article/pii/S0141933120309 522.
306 Discovering and Classifying Digital and Wooden Industries [14] Cong Li, Hui-Qing Lan, Ya-Nan Sun and Jun-Qiang Wang. “Detection algorithm of defects on polyethylene gas pipe using image recognition”. in International Journal of Pressure Vessels and Piping: 191 (2021), page 104381. issn: 0308-0161. https://doi.org/10.1016/j.ijpvp.2021.104 381. https://www.sciencedirect.com/science/article/pii/S03080161210 0079X. [15] Yufeng Shu, Bin Li and Hui Lin. “Quality safety monitoring of LED chips using deep learningbased vision inspection methods”. in Measurement: 168 (2021), page 108123. issn: 0263-2241. https://doi.or g/10.1016/j.measurement.2020.108123. https://www.sciencedirect.co m/science/article/pii/S0263224120306618. [16] Nikolaos Dimitriou, Lampros Leontaris, Thanasis Vafeiadis, Dimosthenis Ioannidis, Tracy Wotherspoon, Gregory Tinker and Dimitrios Tzovaras. “A Deep Learning framework for simulation and defect prediction applied in microelectronics”. in Simulation Modelling Practice and Theory: 100 (2020), page 102063. issn: 1569-190X. https://doi.or g/10.1016/j.simpat.2019.102063. https://www.sciencedirect.com/scienc e/article/pii/S1569190X19301947. [17] Young-Jin Cha, Wooram Choi, Gahyun Suh, Sadegh Mahmoudkhani and Oral Büyüköztürk. “Autonomous structural visual inspection using region-based deep learning for detecting multiple damage types”. in Computer-Aided Civil and Infrastructure Engineering: 33.9 (2018), pages 731–747. [18] Sanli Tang, Fan He, Xiaolin Huang and Jie Yang. “Online PCB defect detector on a new PCB defect dataset”. in arXiv preprint arXiv:1902.06197 : (2019). [19] Yiting Li, Haisong Huang, Qingsheng Xie, Liguo Yao and Qipeng Chen. “Research on a surface defect detection algorithm based on MobileNet-SSD”. in Applied Sciences: 8.9 (2018), page 1678. [20] Cheng-Yang Fu, Wei Liu, Ananth Ranga, Ambrish Tyagi and Alexander C Berg. “Dssd: Deconvolutional single shot detector”. in arXiv preprint arXiv:1701.06659 : (2017). [21] Mingjie Liu, Xianhao Wang, Anjian Zhou, Xiuyuan Fu, Yiwei Ma and Changhao Piao. “Uav-yolo: Small object detection on unmanned aerial vehicle perspective”. in Sensors: 20.8 (2020), page 2238. [22] Haixin Huang, Xueduo Tang, Feng Wen and Xin Jin. “Small object detection method with shallow feature fusion network for chip surface defect detection”. in Scientific reports: 12.1 (2022), page 3914.
16.2 Related Research 313 as the rational outcome of these bounded choices. By utilising approximate optimization techniques like reinforcement learning (RL), hypotheses about user goals and processing limits can be generated, allowing for adaptable strategies. The theory also allows the machines to generate explanations of human actions, i.e., answer “why?” and “what if?” questions, which allows them to be more adaptive. The need for adaptiveness is also motivated by the fact that individuals often distribute cognitive processing across their environment to optimise the use of their mental and cognitive resources, this is known as the adaptive distribution of cognition [13]. The works discussed provide a foundation for exploring computational rationality as a framework for agent interaction. Notably, we do not differentiate between human and machine agents in our approach. This leads us to our central question: “If we develop intelligent agents that approximate human cognition, can we integrate humans into the system as equal agents?” Our research objective is to create a multi-agent system (MAS) interaction framework that allows agents to be dynamically added or removed while pursuing both individual and collective system goals. This framework should have adaptable communication strategies to prevent using inefficient methods of problem-solving (e.g., becoming trapped in local optima [13]) when evidence suggests a better approach exists. The central point in the agent architecture we envision is a self-awareness layer (or if we follow terminology from [12], the internal environment). Suppose we aim to avoid hard-coding production recipes or human-machine interaction. In that case, the control program should be able to “reflect” on itself and “decide” how its capabilities and limitations fit into the current set of system goals. For instance, consider an agent which is a program that controls a robotic arm. It can make the arm move and perform grip physically, but if it lacks awareness of these capabilities and is asked to retrieve a book from a shelf, the agent might respond that it cannot perform the task. One might argue that a self-awareness layer can be easily generated since an agent is a control program with the code often known. Surely, if the agent is a white box and the control program is deterministic, one can very well predict its future behaviour knowing, for example, its execution traces. However, this luck is rare, and agents can be represented by a neural network or other machine learning models. The next point is that to achieve the full coverage of the environment relevant to the agent, it needs to learn not only its own capabilities but also how they fit into the whole system and, hence, what other agents do. Thus,
314 Conscious Agents Interaction Framework for Industrial Automation Figure 16.1 Ontology showing the interaction of two agents. Assume-Derive phases are shown for Agent 1. they need to observe the behaviour of other agents, assume what they could do with respect to the current goals, and derive their behaviour based on this knowledge and their goals. The assume-derive model is shown in Figure 16.1. 16.3 Interaction Framework In this section, we outline the framework within which we can achieve the integration of human cognition principles into the MAS system in a way that a human is treated as an equal agent. It consists of four layers: automation, selfawareness, high-level agent communication, and the highest level where the user specifies the system goals. Each layer at any time must ensure safe operation and pass formal checks when necessary. For each of the layers, in addition to their description, we also name the major problems to be solved. 16.3.1 Layer 1: Automation We work with industrial automation systems; thus, our first layer is the layer of pure automation control programs. We can operate out of the idea that our low-level control programs are smart and can rewrite their sequence of actions, and such work is being carried out. However, at this stage, we assume that they operate with a set of atomic operations that can be activated in any order by a scheduler from the higher level (e.g., using OPC UA). An atomic operation has a common definition and different specifics depending on the implementation of the automation system. Commonly, an atomic operation is a single action performed by a machine that starts in the safe state of the system and leaves the system in a safe state ready for next allowed actions.
16.3 Interaction Framework 315 Figure 16.2 Interaction framework. Consider having a drilling machine with a rotating table. Atomic operation “perform drilling” in addition to lowering the drill machine and turning it on for a defined amount of time, will also include checking if there is a workpiece underneath it. The same suffices with any other process constraints; for example, if there is a constraint on the maximum water temperature in the tank, the operation “heat” will not turn on the heater if the temperature has reached its maximum level. An automation system can be represented as PLC control code with a human-machine interface (HMI) and in this case, we must interface the automation program directly, e.g., by sending data and events to function blocks, in case it is implemented in IEC 61499. Another option is to send the sequences of actions to the manufacturing execution system (MES), which has the set of possible actions defined. Thus, this level directly executes the commands from the upper layer and reports on the results of the operations performed back to the top. The PLC code can also be organised in a way that supports the MAS paradigm on this level to ease the debugging process.
316 Conscious Agents Interaction Framework for Industrial Automation 16.3.2 Layer 2: Self-awareness Self-awareness is one of the core layers of the framework, where the concept of an agent is introduced. As highlighted in Figure 16.2, the automation layer only implements control and equipment-specific safety constraints. The selfawareness layer, in turn, defines which automation operations can be grouped into a wholesome entity together with environmental observations specific to the operations. The grouping can be obvious, e.g., a conveyor belt or a robot arm can be self-sufficient agents. On the other hand, if the robot arm is installed on AGV, a composite agent can represent the whole assembly. Another way to define agents is by following the control loops of the system. Thus, having a vertical farming module with lighting and watering systems, the lighting agent would be capable of adjusting the light according to the lighting schedule and, for example, the PAR (Photosynthetically Active Radiation) sensor. The watering agent then is responsible for the regulation of the water flow depending on the watering schedule and a water flow sensor that detects clogs in the pipes. The self-awareness layer can be developed through RL methods. For example, using data from factory processes, an offline RL method can learn the system’s dynamics and generate an optimal control policy based on past experiences. In this approach, agents are trained to become aware of their capacities through iterative learning and feedback [14]. If a digital twin or a simulation model of the process is available, such development can be supported by a closed-loop system, comprising two main components: a modelling toolbox and a controlling toolbox. In [15], the modelling toolbox allows the creation of a plant model that simulates various system components or the entire system, which, in our case, would transform into a simulation of the controller actions and environmental response. Meanwhile, the controlling toolbox, originally, is a collection of autonomous agents that interact with the plant model using reinforcement learning algorithms, which would be replaced by the “consciousness” of the agents being trained. The main function of this layer is to provide specific to the request available abilities and skills of the agent to the upper layer with the data needed for scheduling and planning and to translate these results into a sequence of atomic operations. The definition of appropriate reward functions for training the self-awareness layer to choose suitable capabilities is then necessary to keep the agents from “overthinking”.
16.3 Interaction Framework 317 One interesting challenge in self-awareness is the possibility of discovering skills that were not demonstrated in the self-learning process. Assume, we have a conveyor belt, which is controlled by the motors and set to move workpieces with some fixed speed. Can the speed be changed to achieve production or energy consumption system goals? Whether the agent can discover new skills safely during the runtime is one of the research questions related to this layer. Another challenge, stemming from the fact that the self-awareness of our agents is remote from their physical abilities, is how the self-awareness layer can adjust to the change of equipment executing approximately the same skills and actions. The advantage of our framework is that the upper layers do not have to be aware of such changes, while the self-awareness layer will have to perform self-discovery. 16.3.3 Layer 3: High-level communication and coordination In this layer, agents communicate and coordinate the completion of tasks to achieve the goals passed down from the upper layer, thus we outline the following main functions of the level: Goal selection. On the top level, the goals can be determined for the whole system (e.g., total energy consumption must be X) or for the subsystems (e.g., the human in the packaging unit should be protected from burning out). Agents must be able to select which goals are theirs to complete and which are for the other agents. Goals can also be adopted from other agents when they determine that assistance is needed. Agents’ collaboration. Agents should be able to group (or “team up”) to collaboratively achieve the goals if needed. For example, for a human not to get a burn-out in the packaging unit, probably, a conveyor belt should work slower after lunch and an AGV should pick the packed products faster to keep the workspace of the human clean. For this, agents (1) must assume the goals of other agents, as precisely as they can, (2) assume their common and individual goals, and (3) derive their action based on the requirement inferred using this information. We call these two phases the “Assume” phase and the “Derive” phase (Figure 16.1). Reasoning can help with the Assume phase, where an agent can analyse the trace of decisions of other agents and find their motivation (for example, in [16]). This will allow agents to plan their actions in the context of plans of other agents.
318 Conscious Agents Interaction Framework for Industrial Automation Dynamic optimal policy determination. Some agents might have several goals to satisfy, which, in turn, can be contradictory and belong to different groups of agents. An agent must be able to find an optimal course of action. A crucial part of this decision-making process is evaluating utility [12]. The utility of an action is not limited to the immediate reward it brings but also depends on the future rewards that can be expected when the same policy is followed over time. By considering both short-term and long-term outcomes, agents can make decisions that optimise their overall performance in achieving individual and system-wide objectives. External environment update handling. Agents can be sprouted and destroyed dynamically, depending on the momentary requirements, which they should be able to communicate and, thus, perceive when they are notified about the same events from the other agents. The reconfiguration in this case must happen automatically with as little downtime or human intervention as possible. Communication with humans. Agents that interact with humans should have the mechanisms to assert the same aspects of the human internal environment as they do when communicating with nonhuman agents. In the implementation of the framework for this layer, we also consider a time aspect of the decision-making. If finding an optimal action plan takes a significant number of resources (time, computational power, etc.), the heuristics are to be used. Also, here, we solve the problem of the cost of changing the plan if a new better option is found. For instance, if midway through the plan execution, the resources are already spent, this layer estimates whether replanning is beneficial. 16.3.4 Layer 4: System goals Mainly the goals sprout from the requirements for the system, about which we talked in Section I. Different formulation languages are possible, for example, various temporal logics, diagrams, or natural language, which is becoming more possible with the development of LLMs. This layer is also the place for enlarging the vocabulary of the system, for example, if the goal is “to keep power consumption at 80% from maximum consumption during the nearest month”, the system should at least know what power consumption is. LLMs and users play a key role in vocabulary and semantics enrichment. The requirements the system should satisfy may have a time interval, for example, different production scenarios of the island-based factory might
16.4 Case Study 319 follow different production recipes, however, the general liveness and safety requirements for the production may remain. One important note is that we do not differentiate here goals for machines and goals for humans within the production. Human workers or users participate in goal distribution on the same level as nonhuman agents do. The difference shows only in the way the goals are communicated to a human as special interfaces are required for the purpose. 16.4 Case Study As a leitmotif for the framework trials, we chose the topic of energy consumption, since, in large productions, being able to reduce it even for a small fraction of the total leads to a significant impact on sustainability and production costs. Our case studies explore different parts of the framework implementation starting from communication between the control loops of the assembly, where each of them can straightaway decide on energy consumption to indirect interaction with humans and agents that influence the energy consumption of the main consumers but not decide on it straight away. 16.4.1 Vertical farming module Our first case study is on the optimisation of the energy consumption of a vertical farming module. Figure 16.3 shows the prototype module implemented at the Aalto Factory of the Future and its overall architecture. Our automation code runs on two M251 PLCs to which both digital and analogue sensors and actuators connect. This layer connects to the business logic via OPC UA. The Figure 16.3 Vertical farming module and its controller architecture.
320 Conscious Agents Interaction Framework for Industrial Automation business logic layer is placed on the industrial PC, while it is also possible to develop it in the cloud. Our high-level logic employs the MAS concept. Each agent here is responsible for calculating the consumption of a single control loop, where a control loop is composed of an actuator and the environmental response on the action of the actuator, for example, pump–water level. The whole module itself can also be an agent in case there is more than one. In this case, it becomes a composed agent, incorporating a set of agents—control loops. Agents communicate with each other to converge on the best individual power consumption values, given the total target power consumption of the system. In this example, we explicitly specified the self-conscious layer by inputting the characteristics of the actuators controlled during the experiments, namely LED lights of the vertical farming module. We have also experimented with setting up communication between different agents through distributed optimisation methods (for details, see [16]). 16.4.2 HVAC system of a cruise ship Another case study we have in progress is the energy consumption of a cruise ship. Here, several cabins connect to one Air Handling Unit (AHU) and several AHUs connect to a diesel engine. In HVAC, the parts that consume the most energy are the fans that drive the air inside the main ducts and the chiller that cools it down to a set temperature. The only way we can influence the energy consumption decision is by adjusting the temperature set points in the cabins. Thus, lowering the cabin temperature would mean that the corresponding valve for fresh air intake should open more, which decreases the pressure in the main air duct, making the main fan increase power consumption to stabilise the pressure, which also makes the chillers cool down a larger air volume. Since the solution should be flexible (extendable to various actuator types) and scalable (more cabins can be connected to air ducts/more AHUs, diesel engines in the system), we define each cabin, each AHU, and each diesel as an agent. This relieves us from the need to explicitly encode the physics of the process (unlike in the first case study) and AHU “learns” its dynamics of power consumption based on the traces of the simulation model using RL methods. One of the research questions here is how to dynamically generate not only the agents but also the correct way of communication between them.
16.5 Discussion and Conclusion 321 Another angle of this study is incorporating users’ patience into the system. The energy reserves are restricted on the cruise ship. Suppose that a lot of people set the desired temperature below some minimum. In that case, there is a chance that the diesel engines will not be able (within safety boundaries) to provide enough energy for the ship infrastructure and the HVAC system. This means that the temperature set point should be adjusted automatically; however, this works up to a point where people do not massively complain. 16.5 Discussion and Conclusion While extensive research has been conducted on isolated issues in MAS and HCI, little attention has been paid to the design of intelligent systems as a whole, especially, from the industrial automation perspective. This paper outlines the foundational problems and research questions aimed at shaping machine consciousness, enabling agents to understand one another as well as communicate effectively with humans. We hypothesise that the implementation of such an architecture leads to emergent behaviours akin to needs and motivations for the system itself. As individual agents become more aware and coordinated, the factory’s operations could form a collective consciousness, transforming the production process into something that responds naturally to the environment rather than following pre-set instructions. This represents the basis of a truly reconfigurable and flexible factory—one where manufacturing as a service becomes an inherent response to the factory’s needs, rather than an artificially imposed command structure. Future work will involve a deeper analysis of the solutions proposed for each layer of the system, with special emphasis on the communication aspects, both within the machine agents and between agents and human workers, which involves exploring how intelligent agents can effectively communicate and collaborate with humans. Furthermore, experiments will be conducted in environments with static and dynamic actuators, including mobile workstations, AGVs, and reconfigurable factories. This framework will also have applications in energy communities, where negotiation and flexibility with human input is vital, especially for systems without plug-and-play capabilities, like HVAC.
322 Conscious Agents Interaction Framework for Industrial Automation Acknowledgements This work was supported, in part, by projects I-SWARM X (p/n 411061), MultiEnergy VPP (p/n 13348415), and the project funded by the Academy of Finland, CAFE (p/n 13363691). References [1] G. Zhabelova, V. Vyatkin, and V. N. Dubinin, “Toward Industrially Usable Agent Technology for Smart Grid Automation,” IEEE Transactions on Industrial Electronics, vol. 62, no. 4, pp. 2629–2641, Apr. 2015, [2] A. Kalachev, G. Zhabelova, V. Vyatkin, D. Jarvis and C. Pang, “Intelligent Mechatronic System with Decentralised Control and Multi-Agent Planning,” IECON 2018 - 44th Annual Conference of the IEEE Industrial Electronics Society, Washington, DC, USA, 2018, pp. 3126-3133, [3] J. Kaiser, M. P. Hernández, V. Kaupe, P. Kurrek, and D. McFarlane, “An agent-based approach for energy-efficient sensor networks in logistics,” Engineering Applications of Artificial Intelligence, vol. 127, p. 107198, Jan. 2024, [4] M. J. Wooldridge, “An introduction to multiagent systems”. Chichester: John Wiley, 2012, [5] M. E. Bratman, D. J. Israel, and M. E. Pollack, “Plans and resourcebounded practical reasoning,” Computational Intelligence, vol. 4, no. 3, pp. 349– 355, Sep. 1988, [6] S. Sardiña, L. de Silva, and L. Padgham, “Hierarchical planning in BDI agent programming languages,” 5th international joint conference on Autonomous agents and multiagent systems (AAMAS ’06), Association for Computing Machinery, New York, NY, USA, pp. 1001–1008, May 2006, [7] F. Alzetta, P. Giorgini, M. Marinoni, and D. Calvaresi, “RT-BDI: A RealTime BDI Model,” Advances in Practical Applications of Agents, MultiAgent Systems, and Trustworthiness, The PAAMS Collection, pp. 16– 29, 2020, [8] M. Wooldridge, “Reasoning about Rational Agents.” MIT Press, 2003, [9] G. Boella, “Decision theoretic planning and the bounded rationality of BDI agents,” In Proceedings of GTDT 2002 Workshop, Technical report WS02, vol. 6, pp. 1-10, 2002,
17.3 Free Energy Principle 329 Figure 17.2 depicts the black-box view of the overall system. The IoT network receives inputs from the environment (such as sensor signals like water consumption measurements, weather data, and other internet-sourced data) and executes commands on the environment (e.g., opening/closing pipes, issuing alarms, etc.). Figure 17.2 Black-Box view of the IoT network and the environment. Using hierarchical distributed control and learning with prediction/correction feedback loops, we can minimize communication overhead and energy consumption. Information is primarily processed where it is available; only predictions and corrections are communicated, not commands. A mathematical framework for this approach is provided by Friston’s active inference and free energy principle [13], [14]. 17.3 Free Energy Principle Friston’s free energy principle [13], [14] is a theoretical framework from neuroscience, which posits that the brain minimizes a quantity called free energy to maintain a stable internal state and make sense of the world. This principle is derived from thermodynamics and statistical mechanics, and it explains how biological systems (like the brain) resist disorder (entropy) by maintaining an internal model of the environment. In our case, each level of the IoT hierarchy has its own world model that is updated based on inputs from the next lower level. Active inference has also been applied to IoT in [15]. The free energy principle can also be broken down into terms involving surprise and approximation errors. F=DKL(q(s)∥p(s|o)) −log p(o)(17.1)
[Document text truncated for crawler view.]