Full text
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 Evaluating Knowledge Distillation and Compression Techniques for Edge Class Devices FOTIOS PAPAROUNAS∗,Dept. of Informatics, Democritus University of Thrace, Greece VASILEIOS CHRISTOFAS, Dept. of Informatics, Democritus University of Thrace, Greece PETROS AMANATIDIS,Dept. of Informatics, Democritus University of Thrace, Greece DIMITRIS KARAMPATZAKIS,Dept. of Informatics, Democritus University of Thrace, Greece THOMAS LAGKAS,Dept. of Informatics, Democritus University of Thrace, Greece Recent advances in Machine Learning (ML) technologies have made it possible to execute neural network models in daily devices such as smartphones and IoT. Due to these advancements, intelligent functionalities can now be integrated into devices with limited resources, enabling AI capabilities like speech recognition, and image processing. However as the AI models go deeper, hardware requirements become a substantial issue since the number of computations increases drastically. To mitigate this, researchers discovered ways to compress AI models without sacrificing a lot of performance. Some examples are quantization and knowledge distillation. These techniques aim to reduce the size, and computational complexity, of the AI models, allowing them to be deployed more efficiently on hardware with limited resources. In this research paper, we evaluate these methods and compare them in terms of efficiency and accuracy. The results showed that knowledge distillation can can be applied in real-life scenarios. CCS Concepts: •Knowledge Distillation;•Edge AI; Additional Key Words and Phrases: machine vision, compression, evaluation ACM Reference Format: Fotios Paparounas, Vasileios Christofas, Petros Amanatidis, Dimitris Karampatzakis, and Thomas Lagkas. 2018. Evaluating Knowledge Distillation and Compression Techniques for Edge Class Devices. In Proceedings of Make sure to enter the correct conference title from your rights confirmation emai (Conference acronym ’XX). ACM, New York, NY, USA, 9pages. https://doi.org/XXXXXXX.XXXXXXX 1 Introduction Nowadays, we can [ 13 ]find a wide range of Artificial Intelligence (AI) applications such as real-time object detection and time-series predictions. Such tasks are made possible with Deep Neural Networks (DNNs). Their structure depends on the computation power of the inference device, and the precision requirements of each implementation. Highperformance apps, for example, may require strong hardware, such as GPUs, whereas lighter apps can function on less capable hardware, balancing efficiency and performance. In the end, the selection of hardware and network architecture guarantees that these AI systems can satisfy the requirements for accuracy and resource limitations required for real-world applications. Examples of such networks are Convolutional Neural Networks (CNNs) which consist of Authors’ Contact Information: Fotios Paparounas, Dept. of Informatics, Democritus University of Thrace, Kavala, Greece, [email protected]; Vasileios Christofas, Dept. of Informatics, Democritus University of Thrace, Kavala, Greece, [email protected]; Petros Amanatidis, Dept. of Informatics, Democritus University of Thrace, Kavala, Greece, [email protected]; Dimitris Karampatzakis, Dept. of Informatics, Democritus University of Thrace, Kavala, Greece, [email protected]; Thomas Lagkas, Dept. of Informatics, Democritus University of Thrace, Kavala, Greece, [email protected]. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. ©2018 Copyright held by the owner/author(s). Publication rights licensed to ACM. Manuscript submitted to ACM Manuscript submitted to ACM 1
53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 2 Paparounas et al. multiple convolution layers built on top of a classifier. Convolutions are essentially feature extractors that filter out details that may be not needed by the classifier. By doing this, CNNs are able to focus on the most significant features of images, including edges, textures, or shapes. Multiple pooling layers may be used as well to further improve the extraction by keeping either the most important information or its average. As DNN architectures go deeper, the need for advanced compression methods increases. Researchers have designed algorithms that try to reduce model sizes while keeping the impact on network performance at a minimum [ 15 , 19 ]. Quantization [ 7 , 11 ] is a compression algorithm that converts float weights down to integers. The size reduction is usually big with the accuracy only getting slightly lower, always depending on the amount of bits per weight. It is possible, for example, to convert a float (32-bit) model to an integer (8-bit) model, and the resulting network could be up to four times smaller than the original. Pruning is yet another compression method as explained in [ 9 ] that involves the deactivation of network neurons if their contribution to the results is considered to be small. Both size reduction and network performance depend on the pruning thresholds. Apart from the algorithms mentioned earlier, researchers have found a new approach that results in increased network accuracy and reduced inference times. Knowledge distillation [ 2 , 4 , 10 , 17 , 18 ] works by using a big structured model as a teacher and a smaller architecture model as a student. The student is trained to replicate the results of the teacher by including a distillation loss in its loss function. Despite having a smaller architecture, the student model is trained to mimic the teacher’s behavior, resulting in comparable performance with far less computational complexity. When implementing AI models in situations with limited resources, like embedded systems or mobile devices, this method is quite helpful for maintaining accuracy without sacrificing efficiency. By reducing the execution requirements of a network, inference on edge is made possible. Edge AI reduces the workload on centralized servers by running machine learning models on edge nodes; devices that are closer to the end-user. This way energy costs fall drastically making servers more eco-friendly while the end-user gets his results faster without worrying about data privacy violations. Applying such compression methods in networks that are designed to be run on edge nodes helps reduce the total amount of power that is consumed while improving the execution speed at the same time. However, non-lossless compression causes partial loss of original data and therefore a small reduction in accuracy is to be expected. The contributions of this paper are the comparison between non-compressed models and compressed (distilled) for which we highlight their differences in terms of accuracy, latency, and power consumption, which is very important for edge devices. Furthermore, we illustrate the efficiency of these models and make suggestions based on the cost of each deep learning compression method. 2 Related Work In this Section we present different deep learning compression methods that are implemented on edge class device in recent literature. The authors in paper [ 1 ] addressed the growing need for object detection in edge computing systems, highlighting the need for models based on lightweight convolutional neural networks (CNNs). Current models require demanding deployment on edge devices and provide memory size difficulties, therefore hardware optimization without performance deterioration is essential. In order to achieve this, they investigated a number of model compression methods aimed at reducing the memory footprint and processing load of CNNs without sacrificing accuracy. The results of the testing showed that quantization and binarization were the most successful strategies in terms of achieving significant model compression with the least amount of loss in detection accuracy. Manuscript submitted to ACM
105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 Evaluating Knowledge Distillation and Compression Techniques for Edge Class Devices 3 A summary of effective deep learning techniques designed to reduce the computational complexity of deep neural networks (DNNs) was presented in [ 3 , 12 ], especially for devices with limited resources such as Internet of Things (IoT) devices and mobile phones. Pruning, factorization, and quantization are three techniques the authors examined for model compression. These techniques are essential for lowering the memory footprint and processing requirements of DNNs. They developed the use of AutoML frameworks like neural architecture search (NAS) and automatic pruning, which speed up the design process by minimizing human participation, in addition to manual solutions for model optimization. Another research [ 5 ] highlighted the use of model compression techniques, with an emphasis on the conventional pruning method, to address the problem of deploying deep neural networks on edge devices with restricted computational resources. According to the study, while pruning can successfully reduce the number of model parameters, it frequently causes an enormous decrease in accuracy. Conventional fine-tuning techniques are used to partially restore the lost performance in order to mitigate this. On the other hand, severe pruning reduces the network’s capacity to the point that fine-tuning is insufficient to bring the network back to its initial accuracy. In order to address this weakness, the authors suggested using knowledge distillation, in which the original, unpruned network acts as a teacher, helping the pruned model, or "student," make up for lost performance. A tensor-train video scene segmentation approach is presented in [ 6 ] that is specifically designed for multimedia Internet-of-things (IoT) systems with edge computing devices. Conventional deep learning techniques encounter difficulties with accuracy and memory limitations when implemented on edge devices with limited resources. In order to reduce memory utilization while maintaining acceptable segmentation accuracy, the authors presented a unique technique that makes use of tensor-train decomposition. In order to identify and segment scene boundary boxes, their method compares local background information across adjacent video frames, utilizing an upgraded Faster R-CNN model to provide area recommendations. containing little computational overhead, the model handles image segmentation efficiently by focusing on background boxes for similarity assessment and eliminating foreground boxes containing sparse items. Model compression for the deployment of deep models on resource-constrained edge devices in industrial Internet of Things (IoT) scenarios was intoduced in [ 8 ]. The research focuses on knowledge distillation (KD) in particular as a successful method of knowledge transfer from a big teacher model to a smaller student model. While sample correlations produced from the feature maps of intermediate layers are the basis of traditional relational KD methods, these approaches frequently experience poor results due to overfitting to the feature maps of the teacher model, particularly when important sample regions are missed. The authors suggested an innovative method that uses attention maps to highlight the most informative areas of the data in order to get around this restriction. In article [ 14 ], the authors presented a novel architecture for miniature artificial intelligence that is computationally efficient. A compressing-while-training architecture addresses current issues with deep neural network deployment on low-cost edge devices. The suggested technique, called Random Sketch Learning (Rosler), greatly reduces the memory footprint and processing requirements usually associated with deep neural networks by enabling the direct learning of a compact model. Rosler facilitates the deployment of small AI solutions and improves adaptability in dynamic contexts where data may change over time by enabling on-device learning. Validated on many models and datasets, the architecture achieves impressive memory reductions of about 50-90 times, especially with 16-bit quantization. The difficulties in creating edge intelligence systems are discussed in this work [ 16 ], where a major obstacle is the computational difference between deep learning algorithms that require a lot of processing power and less capable edge systems. Numerous strategies and optimization techniques, including as network compression, lightweight models, and Manuscript submitted to ACM
157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 4 Paparounas et al. effective neural architecture search, have been put forth to close this gap. Though significant progress has been made, there is still a lack of a systematic and thorough examination of the literature that unifies these diverse approaches and highlights their applicability to the deployment of edge intelligence. The authors highlight the necessity for a comprehensive study to assist researchers and engineers working in the edge computing community in understanding the most recent developments in deep learning techniques, which are essential for creating effective edge intelligence systems. Representative and state-of-the-art methods are surveyed in the paper, including adaptive deep learning models, hardware-aware neural architecture search, hand-crafted models, and model compression. 3 Experimental setup Our target compression method is Knowledge Distillation, combined with (and without) Quantization. To measure the difference between the compressed models we used an edge-class device for the inference, the Nvidia Jetson Nano System-On-Module with its Dev kit. 3.1 The inference edge-class board The Nvidia Jetson Nano comes with a Quad-core ARM Cortex-A57 processor, a 128-core Maxwell GPU, and 4GBs of RAM. The latest OS image Nvidia provides includes Ubuntu 18.04 and Jetpack 4.6.1 which is rather outdated but it won’t affect the end results. The idle power consumption was measured to be around 1.07 Watts. The whole package comes with a passive cooler that can maintain low temperatures even when the chip is under load. 3.2 Measuring energy consumption The energy consumption was measured using Monshoon’s High Voltage Power Monitor (HVPM), a lab-grade power monitor. The monitor is factory-calibrated and can sustain loads up to 13.5 Volts and 6 Amps of continuous current; more than enough for our experiment. We power the edge device via its header pins and provide exactly 5.1 Volts which is well within the limits of the chip (4.75-5.25 Volts). The power measurements are gathered with a fixed sample rate of 5000Hz. 3.3 Test dataset For training and testing the ImageNet dataset was preferred, as it includes up to 1000 classes of different objects. The input data were pre-processed to match the model input requirements, resulting in sample images with a center-cropped 224x224 resolution and RGB channels. From the dataset, 500 of them were randomly selected to evaluate the models. There wasn’t any other type of processing afterwards. 3.4 The evaluation models The evaluated networks are as follows: • MobileNet v3, a depth-wise convolution variant of Convolutional Neural Networks (CNNs) designed for low latency inferences on mobile and edge devices. •EfficientNet v2, yet another CNN designed for low latency. •ResNet101, a heavier network that targets accuracy. Manuscript submitted to ACM
209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 Evaluating Knowledge Distillation and Compression Techniques for Edge Class Devices 5 The knowledge distillation part was done by using the ResNet101 as a teacher. MobileNet and EfficientNet were used as students to measure the impact of this method. Of course, this required more time while training since we had to take into account the results of a bigger teacher network alongside the student one. We configured the input size to 3 × 224 × 224 to match the ImageNet image sizes. We exported float, and quantized variants of the networks in TFLite format using Tensorflow. The quantization was done using the 8-bit integer type while the floats use the default 32-bits length. Evaluated models Model Normal Distilled FLOAT32 INT8 FLOAT32 INT8 EfficientNet v2 27.2MB 7.5MB 17.6MB 5.6MB MobileNet v3 Large 21MB 5.8MB 15.2MB 3.9MB ResNet101 169MB 43.3MB - - Table 1. Network specifications of original and compiled models 3.5 Evaluating the models The execution flow consists of three fundamental steps and is as follows: • For the first step, we load the inference model in RAM using the tflite-runtime library and prepare it for inference on the CPU. We do not apply any kind of CUDA acceleration to make the compression differences visible. •Load each sample from the test dataset in RAM separately and execute the inference model. •Upload latency and accuracy results in a local MLflow Tracking server for later analysis. Here we should add that MLflow is a machine learning development platform designed for use cases such as experimenting. We used it to collect and visualize the inference data in an organized manner and export them in CSV files. It’s important to reduce the device factor as much as possible meaning that the inference time we measure includes only the inference itself, and not other tasks such as loading from RAM or uploading the resulting artifacts. The same applies to power related results. The reason we do this is because many tasks can be done in many different (and maybe more optimal) ways so taking them into account could potentially invalidate our measuring points. 4 Evaluation Results The following metrics were taken into account in order to measure the performance of each classification model: •Latency - Measured in milliseconds (ms) •Accuracy - Measured in percents (%) •Average power consumption - Measured in watts (W) •Network size - Measured in megabytes (MB) •Quantization - A model is either FLOAT32 or INT8 Manuscript submitted to ACM
261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 6 Paparounas et al. 4.1 Model performance Starting with accuracy the results showed small differences between the inference models. What’s so impressive is the fact that the distilled variants had similar accuracy while maintaining a much lower latency. Furthermore, it proved that ResNet is capable enough of using its outputs to teach smaller edge-oriented networks. As shown in Figures [1,2], the accuracy difference between the normal and the distilled models barely exceeds 1%. What’s even more impressive is that the distilled EfficientNet v2 model managed to surpass its teacher’s accuracy. Fig. 1. Normal accuracy Fig. 2. Distillation accuracy The same applies to the quantized versions, where the difference from their corresponding float variant is mostly the same, if not better. These versions are also a lot faster with the fastest being the knowledge-distilled MobileNet v3 model with a latency of 44,8 milliseconds per inference and a total inference time of 22,4 seconds. The slowest quantized inference on the other hand was that of EfficientNet with a latency of 232 milliseconds per inference. Overall it’s still adequate for some applications that do not require real-time response. Fig. 3. Normal latency Fig. 4. Distillation latency Manuscript submitted to ACM
313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 Evaluating Knowledge Distillation and Compression Techniques for Edge Class Devices 7 Comparing the results with the respective file sizes confirms the fact that distillation should be thought of when designing networks that will be run in the edge. The lowest file size was 3,92 MB in the quantized MobileNet v3 while the biggest was of course ResNet’s. Distilling the models using ResNet101 resulted in a marginal decrease of almost 30% in all network file sizes. Combining that with quantization reduced them even further. A model small enough can be loaded in RAM and be executed without overheads in resource constrained microprocessors. We should note that there was no acceleration involved and everything was run on Jetson’s CPU only. One could also make a CUDA implementation to help with the runtime of the models and to potentially reduce the power costs. Devices with other types of ML accelerations can also benefit from the combination of the methods that we evaluate since they reduce the calculation workload. 4.2 Energy consumption The average power consumption of each network was kept around 3 watts at all times with some small differences between the models. Figures [5,6] show that the quantized EfficientNet maintained the lowest consumption of 2,87 Watts on average and this is also true for its distilled variant. The results mean that all of the models took advantage of what was available to them, which makes a lot of sense since they all an on the CPU. Fig. 5. Normal consumption Fig. 6. Distillation consumption The distilled variant of MobileNet v3 turned out to be the most efficient with its lower latency, smaller storage requirements, and decent power consumption. EfficientNet v2 is not far behind because it still wins in terms of accuracy. 4.3 Discussion Despite the fact that not every model can perform well in a real-time environment, Knowledge Distillation can have a significant impact in all aspects of efficiency. Other networks that target edge devices could also benefit from such methods, but it’s important to always keep the desired balance between energy efficiency, accuracy, and finally speed. Combining more compression methods in a single model could mean that the network is just not suitable at all for a specific application. An edge network usually consists of multiple edge nodes that get assigned inference tasks dynamically, so latency is not the biggest concern here. This is especially true with the recent advances of technology to the point where accelerators can finish tasks in a matter of a few milliseconds, but they are very limited when it Manuscript submitted to ACM
365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 8 Paparounas et al. comes to memory size, so it has to be taken into account very carefully. The downside of Knowledge Distillation is its training part, where the student model has to be sufficiently trained with a good teacher. This means that transfer learning is essential to avoid training models from the ground up for weeks on high-end hardware. Additionally, while training the smaller model, a forward pass of the teacher model is required in each iteration alongside the student’s one to calculate the KD loss function, therefore increasing training times even more. 5 Conclusions Knowledge Distillation is an effective method for enhancing model efficiency, making it particularly well-suited for edge devices. Compared to conventional methods, it achieves a better balance between energy efficiency, accuracy, and speed by condensing knowledge from a larger model to a smaller one. Transfer learning helps address these problems, even though the memory constraints of edge accelerators and the extra training time brought on by teacher-student interactions can be difficult. All things considered, Knowledge Distillation provides a more workable and effective way to implement models in resource-constrained edge situations. Acknowledgments This project has received funding from the European Union’s Horizon Europe research and innovation programme under grant agreement No 101070181. This work was supported by the MPhil program "Advanced Technologies in Informatics and Computers", which was hosted by the Department of Informatics, Democritus University of Thrace, Kavala, Greece. References [1] Olutosin Ajibola Ademola, Mairo Leier, and Eduard Petlenkov. 2021. Evaluation of deep neural network compression methods for edge devices using weighted score-based ranking scheme. Sensors 21 (11 2021). Issue 22. https://doi.org/10.3390/S21227529 [2] F MohiEldeen Alabbasy, Abdelaziz Said Abohamama, and Mohammed F Alrahmawy. 2023. Compressing medical deep neural network models for edge devices using knowledge distillation. Journal of King Saud University-Computer and Information Sciences 35, 7 (2023), 101616. [3] Han Cai, Ji Lin, Yujun Lin, Zhijian Liu, Haotian Tang, Hanrui Wang, Ligeng Zhu, and Song Han. 2022. Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications. ACM Transactions on Design Automation of Electronic Systems 27 (5 2022). Issue 3. https://doi.org/10.1145/3486618 [4] Haowei Chen, Liekang Zeng, Shuai Yu, and Xu Chen. 2020. Knowledge distillation for mobile edge computation offloading. arXiv preprint arXiv:2004.04366 (2020). [5] Liyang Chen, Yongquan Chen, Juntong Xi, and Xinyi Le. 2022. Knowledge from the original network: restore a better pruned network with knowledge distillation. Complex and Intelligent Systems 8 (4 2022), 709–718. Issue 2. https://doi.org/10.1007/S40747-020-00248-Y [6] Cheng Dai, Xingang Liu, Laurence T. Yang, Minghao Ni, Zhenchao Ma, Qingchen Zhang, and M. Jamal Deen. 2021. Video Scene Segmentation Using Tensor-Train Faster-RCNN for Multimedia IoT Systems. IEEE Internet of Things Journal 8 (6 2021), 9697–9705. Issue 12. https://doi.org/10. 1109/JIOT.2020.3022353 [7] Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer. 2022. A survey of quantization methods for efficient neural network inference. In Low-Power Computer Vision. Chapman and Hall/CRC, 291–326. [8] Jianping Gou, Liyuan Sun, Baosheng Yu, Shaohua Wan, Weihua Ou, and Zhang Yi. 2023. Multilevel Attention-Based Sample Correlations for Knowledge Distillation. IEEE Transactions on Industrial Informatics 19 (5 2023), 7099–7109. Issue 5. https://doi.org/10.1109/TII.2022.3209672 [9] Song Han, Huizi Mao, and William J. Dally. 2016. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding. http://arxiv.org/abs/1510.00149 arXiv:1510.00149 [cs]. [10] Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the Knowledge in a Neural Network. http://arxiv.org/abs/1503.02531 arXiv:1503.02531 [cs, stat]. [11] Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko. 2017. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. http://arxiv.org/abs/1712.05877 arXiv:1712.05877 [cs, stat]. [12] Jinsheng Ji, Zhou Shu, Hongqun Li, Kai Xian Lai, Minshan Lu, Guanlin Jiang, Wensong Wang, Yuanjin Zheng, and Xudong Jiang. 2024. Edgecomputing based knowledge distillation and multi-task learning for partial discharge recognition. IEEE Transactions on Instrumentation and Measurement (2024). Manuscript submitted to ACM
417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 Evaluating Knowledge Distillation and Compression Techniques for Edge Class Devices 9 [13] Farinaz Koushanfar, Vandana Prabhu, Miodrag Potkonjak, and Jan M Rabaey. 2000. Processors for mobile applications. In Proceedings 2000 International Conference on Computer Design. IEEE, 603–608. [14] Bin Li, Peijun Chen, Hongfu Liu, Weisi Guo, Xianbin Cao, Junzhao Du, Chenglin Zhao, and Jun Zhang. 2021. Random sketch learning for deep neural networks in edge computing. Nature Computational Science 1 (3 2021), 221–228. Issue 3. https://doi.org/10.1038/S43588-021-00039-6 [15] Zhuo Li, Hengyi Li, and Lin Meng. 2023. Model Compression for Deep Neural Networks: A Survey. Computers 12, 3 (2023), 60. [16] Di Liu, Hao Kong, Xiangzhong Luo, Weichen Liu, and Ravi Subramaniam. 2022. Bringing AI to edge: From deep learning’s perspective. Neurocomputing 485 (5 2022), 297–320. https://doi.org/10.1016/J.NEUCOM.2021.04.141 [17] Xiaoyang Qu, Jianzong Wang, and Jing Xiao. 2020. Quantization and knowledge distillation for efficient federated learning on edge devices. In 2020 IEEE 22nd International Conference on High Performance Computing and Communications; IEEE 18th International Conference on Smart City; IEEE 6th International Conference on Data Science and Systems (HPCC/SmartCity/DSS). IEEE, 967–972. [18] Stylianos Tsanakas, Aroosa Hameed, John Violos, and Aris Leivadeas. 2024. A light-weight edge-enabled knowledge distillation technique for next location prediction of multitude transportation means. Future Generation Computer Systems 154 (2024), 45–58. [19] Ching-Hao Wang, Kang-Yang Huang, Yi Yao, Jun-Cheng Chen, Hong-Han Shuai, and Wen-Huang Cheng. 2022. Lightweight deep learning: An overview. IEEE Consumer Electronics Magazine (2022). Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009 Manuscript submitted to ACM