scieee AI-readable full text Open interactive document viewer

Multi-Objective Reverse Offloading in Edge Computing for AI Tasks

Amanatidis, Petros; Michailidis, George; Karampatzakis, Dimitris; Kalenteridis, Vasileios; Iosifidis, George; Lagkas, Thomas

Abstract

Offloading tasks between edge nodes is a subject that has drawn a lot of attention since edge computing first emerged. A large number of edge IoT devices utilizing increased computing resources such as autonomous vehicles and UAVs can be used to execute AI tasks close to users. We present a novel approach that deviates from the conventional edge computing offloading concept namely offloading computationally intensive tasks from cloudlets to nearby end nodes. Specifically, we enhance a scenario where end nodes assist more powerful nodes (like cloudlets) in executing AI inference tasks. In edge computing networks, as end nodes grow in number, they build an idle computing capacity which can solve and provide efficient solutions. Our goal is to solve a defined Multi-Objective optimization problem with three objectives namely the overall execution time (slowest substasks), the execution accuracy, and the total energy consumption. We address this challenging optimization problem using a novel method with our released Multi-Objective Edge AI-Adaptive Reverse Offloading, or MOEAI-ARO, algorithm. Using an edge computing testbed and a representative AI service, we demonstrate the effectiveness of our reverse offloading proposal and method. The results indicate that our method further optimizes the system’s performance compared to baseline algorithms.

Full text

Received 27 January 2025; revised 4 March 2025; accepted 25 March 2025. Date of publication 28 March 2025; date of current version 10 April 2025. Digital Object Identifier 10.1109/OJCOMS.2025.3555947 Multi-Objective Reverse Offloading in Edge Computing for AI Tasks PETROS AMANATIDIS 1, GEORGE MICHAILIDIS 1, DIMITRIS KARAMPATZAKIS 1, VASILEIOS KALENTERIDIS 1, GEORGE IOSIFIDIS 2, AND THOMAS LAGKAS 1(Senior Member, IEEE) 1Department of Informatics, Democritus University of Thrace, 65404 Kavala, Greece 2Department of Software Technology, Delft University of Technology, 2600 AA Delft, The Netherlands CORRESPONDING AUTHOR: T. LAGKAS (e-mail: [email protected]r) This work was supported by the European Union’s Horizon Europe Research and Innovation Programme under Grant 101070181. ABSTRACT Offloading tasks between edge nodes is a subject that has drawn a lot of attention since edge computing first emerged. A large number of edge IoT devices utilizing increased computing resources such as autonomous vehicles and UAVs can be used to execute AI tasks close to users. We present a novel approach that deviates from the conventional edge computing offloading concept namely offloading computationally intensive tasks from cloudlets to nearby end nodes. Specifically, we enhance a scenario where end nodes assist more powerful nodes (like cloudlets) in executing AI inference tasks. In edge computing networks, as end nodes grow in number, they build an idle computing capacity which can solve and provide efficient solutions. Our goal is to solve a defined Multi-Objective optimization problem with three objectives namely the overall execution time (slowest substasks), the execution accuracy, and the total energy consumption. We address this challenging optimization problem using a novel method with our released Multi-Objective Edge AI-Adaptive Reverse Offloading, or MOEAI-ARO, algorithm. Using an edge computing testbed and a representative AI service, we demonstrate the effectiveness of our reverse offloading proposal and method. The results indicate that our method further optimizes the system’s performance compared to baseline algorithms. INDEX TERMS AI task offloading, edge computing, multi-objective optimization, resource allocation. I. INTRODUCTION EDGE computing is becoming a necessary component of our lives and is constantly developing. Enabling end nodes (such as mobile phones, Internet of Things devices, etc.) to execute computational tasks closer to the data source is considered a crucial component of future networks. These tasks are often offloaded to more powerful edge devices like cloudlets which enhances the low-latency applications in IoT [1],[2],[3]. Edge computing improves the performance of resource-intensive tasks by bringing computation and storage closer to the user. This reduces the latency caused by sending data to distant cloud servers. On the other hand, communication costs in terms of latency and energy consumption must be taken into account for end devices that are powered by batteries. Interestingly, there is a growing opportunity to investigate the opposite direction: allowing end nodes to perform computational tasks that are assigned to them by nearby edge servers or cloudlets. Through the utilization of numerous end nodes located close to the edge servers, this new architecture has the potential to provide even greater advantages by decreasing response times and improving real-time applications. This potential architecture can be implemented particularly in edge computing applications such as smart cities or in industry, where many end devices either remain idle or execute tasks that require small processing power, thereby enabling them to assist the edge servers in the data processing created within the edge computing network. In the last few years, the research community has thoroughly researched the “traditional” task offloading, from the end devices to the edge servers. Indicating the powerful hardware of the end devices, this new reverse task offloading architecture can prove even more beneficial for processing AI tasks. It is true that as end devices increase in edge c 2025 The Authors. This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License. For more information, see https://creativecommons.org/licenses/by-nc-nd/4.0/ 2474 VOLUME 6, 2025 computing networks, they create an idle computing capacity. This capacity can be useful if it is combined and utilized methodically. The use of Machine Learning (ML) inference from data samples collected at the cloudlets has increased recently. Pre-trained Deep Neural Networks (DNNs) are deployed on the end nodes in large numbers thanks to improvements in end nodes’ hardware and the creation of DNN models that require less processing power and storage at the expense of lower inference accuracy [4],[5],[6]. The deployment of AI models on end nodes is enabled by dedicated AI hardware integrated into current processors. Several AI models of different computational characteristics are being used at the edge. The inference time, accuracy, and energy consumption are directly related to the model’s characteristics. In more detail, larger models achieve more accurate results or predictions but need longer processing time and more energy. Demonstrating their crucial role in reducing task execution delays and energy consumption of end nodes, task offloading has gained significant attention in recent years [7],[8],[9]. Furthermore, the implementation of different kinds of optimization algorithms for task allocation guarantees in most cases the optimal use of the resources in IoT networks which significantly reduces latency and energy usage. The research community is continuously focusing on improving such task allocation methods to handle real-time applications. This enhances the opportunity to investigate the implementation of novel algorithms suited for our proposed task offloading architecture. In this work, we use the reverse offloading approach and solve a multi-objective optimization problem with a newly introduced optimization algorithm. We consider, as a use case, an exemplary AI inference service, namely, an object detection process on images that are located in an edge server. Our goal in our defined multi-objective optimization problem is to find the optimal distribution of the images among the available edge devices (edge server, end devices) and also find the optimal deep learning model selection which takes into account all objectives simultaneously. In our optimization problem, we consider three objectives: the AI inference latency which is determined by the slowest execution of any subtask assigned to some end node, the accuracy of the object detection process, and the total energy consumed for processing the images. Since the above objectives are competitive, there is not a single solution that optimizes the three objectives simultaneously. For this reason, we use the notion of Pareto-optimal solutions [10],[11]. We propose a novel methodology for the automatic selection of a solution that belongs to the Pareto front. To tackle the challenging mixed continuous-discrete optimization problem we introduce a novel adaptive task offloading method that combines a Linear Programming (LP) [12],[13] algorithm with a stochastic Non-dominated Sorting Genetic Algorithm (NSGA) [14]. The results show considerable performance gains and cost savings. A. METHODOLOGY AND CONTRIBUTIONS We explore a scenario of a wireless edge computing network with one access point, an edge server, and end nodes with multiple implemented deep learning models. All end devices are connected via links of different capacities with the edge server. We search for a solution that optimizes simultaneously three objectives, namely the processing speed of the workload, the execution accuracy, and the total energy consumption by splitting the workload (images) among available edge devices via the edge server and additionally selecting the optimal deep learning model for each edge node. The workload consists in processing images (such as videos with a specified duration) that are distributed from the edge server to the edge nodes (including the edge server) where an object detector is used to locate objects of interest. Since each edge device disposes a set of object detectors with a specific neural network size, each combination of device and object detector model essentially has different performance and efficiency. Furthermore, since the devices are connected over wireless links with varying capacities, the transmission time from the edge server to each device differs. The distribution or splitting of the images and the selection of the object detector for the edge nodes are determined by the edge server’s algorithm. For the evaluation of the performance, gains are expressed in terms of all objectives. We provide a complete solution employed as a reverse offloading method. Using a wireless testbed consisting of three Raspberry Pis (RPis) of which one serves as an edge server and two Nvidia Jetsons, this method is evaluated through a series of experiments utilizing the YOLO object detector [15]. When compared to different task offloading baseline methods, the findings demonstrate that our implemented strategy provides an improved overall performance. Furthermore, our offloading technique which we refer to as Multi-Objective Edge AI-Adaptive Reverse Offloading (MOEAI-ARO), tailors its operation by considering resource availability. Consequently, the following contributions are provided by this work: •We enhance the concept that involves edge nodes splitting up their tasks and sending the subtasks to faredge nodes (end devices). This is known as reverse offloading, which is of high importance for edge computing services and IoT networks. •Formulation of a Multi-Objective optimization problem that optimizes simultaneously three crucial objectives, namely the latency, the total energy consumption, and the execution accuracy for edge computing applications. •Consideration of the different object detector models as an optimization variable. •Development of an adaptive reverse offloading algorithm (MOEAI-ARO) that can be tailored to different uniform AI tasks as well as different end devices with VOLUME 6, 2025 2475 AMANATIDIS et al.: MULTI-OBJECTIVE REVERSE OFFLOADING IN EDGE COMPUTING FOR AI TASKS different hardware processing capacities and resource constraints. •Thorough performance evaluation by comparing the proposed MOEAI-ARO algorithm with various baseline algorithms in a wireless edge computing testbed, which consists of end devices with different processing capacities. II. RELATED WORKS Reverse task offloading. The optimization of the system’s latency for vehicular edge computing using a reverse offloading framework was presented in [16]. Resource management and task allocation were optimized via a greedy-based efficient searching (GES) method for binary reverse offloading strategies and the joint alternative optimization-based bi-section searching strategy for partial reverse offloading. Another reverse offloading framework that used a GES algorithm was presented in [17] optimizing the system’s services capacity in a cooperative fashion. Moreover, in [18] the authors developed a reverse offloading strategy for MEC (Multi-access Edge Computing). In more detail, they implemented a heuristic-based and machine learning method, such as Deep Reinforcement Learning, to optimize reverse offloading decisions, effectively reducing latency and energy consumption. In this paper, we focus particularly on handling resource-intensive tasks, such as computer vision tasks, using additional degrees of freedom (DoF) corresponding to different object detector models. AI task offloading. In recent years much work from the scientific community has been devoted to the optimization of AI inference tasks in edge computing networks. Researchers in [19] formulated a Mixed-Integer Nonlinear Programming (MINLP) problem to optimize the latency, accuracy, and energy consumption of AI inference in edge computing networks for videos using the so-called Channel-Aware heuristic algorithm. Offloading of Machine Learning tasks in edge computing networks was examined in [20].An Integer Linear Programming problem was defined, aiming to maximize the processing accuracy of AI inference tasks. In more detail, the authors’ purpose was to find the optimal trade-off between the model size and the processing accuracy using a Dynamic Programming algorithm. In addition, an Automated Machine Learning framework was proposed in [21] that optimizes the inference accuracy of AI tasks while adhering to the minimum frame-rate constraint. The introduced online optimization algorithm improves the performance of the edge computing systems in real time without violating the constraints. All the aforementioned methods only employ continuous or integer optimization variables. Our method shows how to treat discrete and continuous optimization variables efficiently in a model where an edge device (cloudlet) divides and offloads the tasks to multiple smaller devices. Multi-objective optimization. The most reasonable optimization problems in edge computing are those that optimize multiple objectives. In [22] heuristic-based solutions are used to demonstrate significant gains in computational time and energy efficiency. The proposed system, named RAMOS, utilizes different kinds of edge nodes and addresses a multi-objective resource-aware task assignment and scheduling problem with modes for energy efficiency and latency minimization. An interesting work is presented in [23] where the authors tried to identify the optimal trade-off accuracy and latency for deep learning tasks. They formulated an Integer Linear Programming optimization problem to optimize the user’s satisfaction. Additionally, a polynomial constant-time greedy algorithm was presented to get the neat optimal solution. Another work [24] introduced a Constrained Multi-Objective Decomposition Evolutionary Algorithm to optimize energy consumption and latency. This algorithm was implemented in a UAV-assisted MEC network. Another paper [25] attempted to challenge the optimal management of available resources (edge devices) in IoV networks. The NSGA-III algorithm was implemented to optimize objectives such as latency, energy consumption, and load balancing. Authors in [26] used a Genetic Algorithm (GA) [27] to optimize the task processing latency while adhering to several constraints. The presented approach attempts to find the optimal resource allocation in the MEC network. In our proposal, we explore further the implementation of an evolutionary algorithm (NSGA) for our multi-objective optimization problem with the difference that our constraints are handled by an LP algorithm which results in a highly efficient optimization algorithm. This novel strategy efficiently optimizes our three objectives using both discrete and continuous optimization variables. The discrete variables are handled by the NSGA algorithm while the continuous ones by the LP algorithm. III. MODEL AND PROBLEM FORMULATION We present our system model and the corresponding mathematical formulation. A batch of images B,orelsethe Batch step, is distributed by the edge server to a set N of N devices including the edge server for processing. We consider a wireless IoT network represented by the set Nconnected over wireless links with different capacities, all operating with standard IoT technologies such as Wi-Fi. The devices execute an AI task, or more specifically, an object detection task. For object detection, each device in our system implements a deep learning model (neural network). Because the models may vary in size, they may also differ in terms of energy consumption, mean accuracy, and computational performance. A set of Mneural networks is represented by the index yi∈Mfor each device i∈N where M={1,2,...,M}. We set the index i=1for the edge server. We introduce an optimization variable xi∈[0,1] for each edge device i. This variable denotes the part of Bdistributed to the edge device i∈[1,N]. All of these variables are collected into a vector x=[x1,...,xN]∈[0,1]Nand satisfy the constraint N i=1xi=1. We also add a discrete optimization variable y=[y1,...,yN]which denotes the 2476 VOLUME 6, 2025 TABLE 1. Key parameters, functions and variables. neural network selection utilized for object detection at each edge device. Remark: The number of the images distributed from the edge server to the end device i∈[1,N]equals xiB and is rounded to the nearest integer value. As long as a significant number of tasks is distributed at each batch step, this rounding has a negligeable impact on the optimization solution. A. FORMULATION OF THE OPTIMIZATION PROBLEM In this section, we provide the analytical formulation of the optimization problem. All key parameters, functions, and decision variables used in our formulation are summarized in Table 1. Our first objective is to minimize the maximum latency. Here, latency is defined for each device as the total amount of time required for both, the transmission time of images from the server to the end device i, denoted as Ti tx, and the time to execute the object detection process by end device iusing the neural network yi, denoted as Ti dl(yi), i.e., Li(xi,yi)=xiBTi tx +xiBTi dl(yi)=Cixi(1) where: Ci=BTi tx +Ti dl(yi)(2) Ti tx =Datasize DataRate.(3) We denote the maximum latency as: Lmax(x,y)=max iLi(xi,yi)(4) The second objective corresponds to maximizing the mean accuracy of the total object detection process, denoted as mAP(x,y), which is computed as: mAP(x,y)= i ximAP(yi),(5) where mAP(yi)stands for the mean accuracy of the object detection process for one image using the neural network yi. Finally, the third objective is to minimize the total energy consumption (E(x,y)), which corresponds to the energy cost for data transmission (Etr i), data reception (Erec i)and task execution (Eexec i), i.e.: E(x,y)= N  i=1 xiBEexec i(yi)+ N  i=2 xiBErec i+Etr i(6) Remark: In equation (6), we have considered that Erec 1= 0, since the edge server does not consume any energy for data reception, and Etr 1=0, since it does not transmit data to itself. Thus, the objective function of the multi-objective optimization problem can be written in vector form as: min x,y[Lmax(x,y),−mAP(x,y),E(x,y)](7) In real-world applications, it is also necessary to consider the energy availability to ensure that devices are truly capable of completing the assigned tasks. For every device i,we need to include an additional constraint, s.t. Ei(xi,yi)≤Eavailable i,(8) where Ei(xi,yi)corresponds to the energy consumed by the device ifor data transmission, data reception, and task execution, and Eavailable irepresents the energy available in VOLUME 6, 2025 2477 AMANATIDIS et al.: MULTI-OBJECTIVE REVERSE OFFLOADING IN EDGE COMPUTING FOR AI TASKS the end device i. Depending on the device index i,theterm Ei(xi,yi)is further analysed as: E1(x1,y1)= N  i=2 xiBEtr i+x1BEexec 1(y1)(9) Ei(xi,yi)=xiBErec i+xiBEexec i(yi),i=2,...,N(10) Therefore, the complete optimization problem reads: min x,y[Lmax(x,y),−mAP(x,y),E(x,y)](11.1) s.t.Ei(xi,yi)≤Eavailable i,i=1,...,N(11.2) N  i=1 xi=1(11.3) xi∈[0,1],i=1,...,N(11.4) (11) To provide a detailed explanation of the multi-objective optimization problem formulated in (11), the objective in (11.1) is to simultaneously optimize the three objectives, namely to minimize the maximum latency, maximize the mean accuracy, and minimize the total energy consumption. Regarding the constraints of the optimization problem, equation (11.2) ensures that end device ihas the required available energy (Eavailable i)to execute the distributed task. According to equation (11.3), the sum of the portion {xi}i=1,...,Nof Btransmitted to all end devices must be equal to 1. Lastly, the bound constraints for the optimization variables {xi}i=1,...,Nare stated in equation (11.4). As in our previous work [28], to avoid the nondifferentiability of the max operator in Lmax(x,y), we adopt the approach in [29] and introduce an auxiliary optimization variable, denoted x0∈[0,∞)which serves both as a cost function and as a uniform bound for the latency of each device, i.e., we use the equivalence between the following optimization problems: min x,ymax iLi(xi,yi)and min x,yx0 s.t.Li(xi,yi)≤x0,i=1,...,N Finally, the multi-objective optimization problem reads: min x,y[x0,−mAP(x,y),E(x,y)] s.t.Cixi≤x0,i=1,...,N Ei(xi,yi)≤Eavailable i,i=1,...,N N  i=1 xi=1 x0≥0 xi∈[0,1],i=1,...,N(12) The main challenges in efficiently solving the optimization problem (12) and providing a solution that offers a well balanced trade-off in terms of all three objectives are explained in Section III-B and III-C correspondingly. B. SOLUTION OF MULTI-OBJECTIVE OPTIMIZATION PROBLEM The above optimization problem can be written in compact form as: min x,y,x0∈F{f1,f2,f3}(13) where Fdenotes the feasible set defined by the constraints of problem (12), while f1,f2,f3denote the corresponding objectives. Using a classical Genetic Algorithm for this problem would be too costly for a practical implementation of a taskoffloading problem, where decisions need to be taken in very short time. More specifically, GA algorithms are not very efficient in handling several constraints simultaneously, in particular for problems as (12) where continuous variables appear in the constraints formulation. This difficulty becomes even more evident as the number of devices increases, requiring many iterations for the algorithm to converge. For this reason, as in [28], we search to tailor a Genetic Algorithm with an LP algorithm. The former aims to find the optimal choice of neural networks while the latter is occupied with choosing the optimal task distributions and satisfying all the set of constraints for each choice of neural networks. To formulate such an LP problem, we need to construct a single-objective optimization problem. Since all our constraints are linear (for fixed y), the objective function shall also be chosen to be linear. An evident choice is to construct a weighted sum of f1,f2,f3, i.e., min x,y,x0∈Fw1f1+w2f2+(1−w1−w2)f3(14) where w1,w2,w3≥0are positive weight coefficients. However, the choice of the above weights is not straightforward. It seems more natural to produce a front of Pareto-optimal solutions for different values of weights, and then choose accordingly. However, it would be too costly to solve the above optimization problem for a great number of weight combinations. Another idea, adopted in this work, consists in adding w=[w1,w2,w3]as an optimization variable in the Genetic Algorithm, searching at the same time for combinations that create Pareto-optimal solutions. Therefore, the NSGA-II algorithm creates Pareto-optimal solutions for the optimization problem: min x,y,w,x0∈Fw1f1+w2f2+(1−w1−w2)f3(15) by evolving a population of solutions {y,w}while for each member of the population the optimization variables x,x0 are determined via the solution of the LP-problem: min x,x0∈Fw1f1+w2f2+(1−w1−w2)f3(16) To achieve a better scaling for the three objectives in the weighted sum (16), all values are normalized such that fnorm i∈[0,1],i=1,2,3, using the formula: fnorm i=fi−fmin i fmax i−fmin i (17) 2478 VOLUME 6, 2025 where fmax i,fmin iare the maximum and minimum possible values for each objective. Remark: Let us emphasize that several Evolutionary Algorithms could be used instead of the proposed Genetic Algorithm, without expecting any significant influence on the performance of the method. The central idea of the methodology lies in the combination of any Evolutionary Algorithm with an LP algorithm. Our choice of GA lies purely in its simplicity of implementation. C. DEFAULT SELECTION AMONG PARETO-OPTIMAL SOLUTIONS By definition, all Pareto-optimal solutions are equivalent. Without additional specifications, we propose a methodology for default selection among Pareto-optimal solutions, based on a distance function, which ensures a proper equilibrium with respect to all objectives. First, we compute the Center of Gravity (CoG) of the solutions on the Pareto front. Each coordinate of the CoG vector reads: fCoG i=P=100 j=1fnorm i,j P,i=1,2,3(18) where Pis the number of Pareto optimal solutions on the front and fnorm i,jdenotes the ith normalized objective of the jth solution on the front. Then, we define an ellipsoidal distance function which computes a distance between two points xj,xkon the Pareto front as: distancexj,xk=    3  i=1fnorm i,j−fnorm i,k maxlfnorm i,l−minlfnorm i,l2 (19) In the above definition, all three objectives are re-normalized based on their maximum and minimum values among solutions of the front (maxfnorm i,l,minfnorm i,l)which provides a better scaling among objectives of different scales and thus an improved equilibrium. Finally, the default selection is found as: xdef =argmin xj distanceCoG,xj(20) i.e., we choose the solution on the Pareto front that lies closer to the CoG in terms of the distance function (19). IV. PROPOSED ALGORITHM First, the initial population P0is created. It consists of Zy vectors that combine different neural network indices along with different weight vectors w. Every individual in P0has to solve a Linear Programming (LP) problem to evaluate its fitness function. Subsequently, the generation counter tis set to zero. In every iteration, operators for crossover and mutation are applied to Ptin order to produce an offspring population Qt, for which we evaluate its fitness function. Rt=Pt∪Qtis the result of combining the populations of the parents and offspring. Non-dominated sorting is performed on Rt, and crowding distances are computed to measure the Algorithm 1 NSGA-II-LP (MOEAI-ARO) Algorithm 1: Initialize population P0of size Z. 2: Solve the LP problem and evaluate the fitness for each individual in P0. 3: t←0. 4: while termination condition not met do 5: Apply crossover and mutation operators on Ptto generate offspring population Qt. 6: Solve the LP problem and evaluate the fitness for each individual in Qt. 7: Combine parent and offspring populations: Rt=Pt∪ Qt. 8: Perform non-dominated sorting on Rt. 9: Calculate crowding distance for each individual in each front. 10: Form new population Pt+1by selecting the best Z individuals from Rtbased on rank and crowding distance. 11: t←t+1 12: end while 13: Select the default Pareto optimal solution. TABLE 2. NGSA-II algorithm parameters. density of solutions surrounding a specific individual. Based on the rank and crowding distance, the best Zindividuals from Rtare chosen to form the new population Pt+1. Once the termination condition is satisfied (e.g., maximum number of iterations), the algorithm stops and the default Paretooptimal solution is computed. Our selection of parameters for Algorithm 1can be found in Table 2following an extensive amount of numerical testing. A. COMPLEXITY ANALYSIS The computational complexity of the proposed MOEAIARO method can be computed as a combination of the corresponding complexities of the NSGA-II [30] and SIMPLEX algorithms [31]. First, the complexity of the classical NSGA-II algorithm for a population of size Zand Kobjectives is known to be O(ZK2). Then, although the SIMPLEX method has exponential worst-case complexity, its average complexity is polynomial and approximately O(N2) where Nis the number of variables. This holds because the number of variables may grow significantly, but the number of inequalities remains very limited. Therefore, the total average complexity of the MOEAI-ARO algorithm is estimated as O(ZN2K2). VOLUME 6, 2025 2479 AMANATIDIS et al.: MULTI-OBJECTIVE REVERSE OFFLOADING IN EDGE COMPUTING FOR AI TASKS B. SCALABILITY In our previous research [28] we proposed a clustering methodology for handling large-scale testbeds using the EAI-ARO method. The same strategy could be applied to its multi-objective version, i.e., the proposed MOEAI-ARO method, solving one multi-objective optimization problem per cluster of devices. Since the complexity of the MOEAI-ARO method is K2 higher compared to EAI-ARO, one should expect that the size of clusters are smaller in size. However, the exact optimal size depends highly on the corresponding implementation of the MOEAI-ARO method. Using programming languages that are suitable for high performance computing, then exploiting the possibility of massive parallelism, allows to reduce significantly the solution time of the optimization problem and thus the impact on the total latency. Furthermore, in our current implementation the actual dynamic problem is approximated by a series of static problems and the distribution of images occurs after the optimization problem is resolved. This is an unfavourable scenario for the scalability of the method, since the optimization problem could be solved continuously, updating the dynamic parameters at every iteration. Once the tasks have been executed, the current optimal solution could be used for the next distribution. This strategy could improve significantly the scalability of the method. All of the above directions have not been examined in this work and are addressed as future work. V. RESULTS In this section, we demonstrate numerical and experimental findings utilizing the suggested adaptive offloading method. We introduce our testbed and compare the MOEAI-ARO with different baseline methods. For all baseline methods, the same methodology is used to choose a default solution on the Pareto front, illustrated in Section III-C. More specifically, for each baseline method, we compute the first Pareto front of the solutions across all neural network (NN) model combinations (see Figure 5). Then, we select the solution closest to the center of gravity of this front. The evaluation process takes into account all three objectives. The baseline methods used to verify the effectiveness of our proposed method are the following: •Local: Every task is executed locally at the edge server. This can be considered as a greedy method [32] that minimizes the network delay [33]. •First Available (FA): The edge server initially assigns an equal amount of tasks to the end devices and to itself. Once a device completes its assigned tasks and becomes idle, the server reallocates new tasks to it, as described in [22]. More specifically, the total batch of images is divided into equal smaller batches. When a device becomes idle the server immediately sends the next batch of images for processing. It is important to highlight that for this method, the TCP socket connection is maintained open while the devices are processing images. This choice results in lower latency, since we need no additional time to create the connection, but higher power consumption (4.5 Watt for the testbed considered herein). •Random: The edge server simultaneously assigns tasks to all end devices according to a random distribution. •Round Robin: The edge server distributes the tasks uniformly across all edge devices. In more detail, the distribution for this baseline is x= [0.2,0.2,0.2,0.2,0.2] which corresponds to 20% of the batch distributed to the five devices. This method is a commonly used baseline method [22], that ensures that the workload is equally distributed, regardless of the system’s network condition. All the above baseline methods are parameter-free, which permits one to avoid any unintentional bias during the comparison that may occur from parameters’ calibration. A. TESTBED DESCRIPTION Figure 1shows our heterogeneous edge computing setup, containing four end devices, an edge server, and a Wi-Fi access point. The four end devices in the illustration are two Raspberry Pi’s 4 and two Nvidia Jetson’s, each having 4GB of RAM and varying capacities of wireless connectivity to the edge server (see Figure 2). The edge server is also a Raspberry Pi’s 4 with 8 GB RAM. All edge devices including the edge server are equipped with various YOLO object detectors. According to Table 3, each end device may have many object detection models, each implemented with a different input image size, energy consumption value, and mAP value. Regarding the energy consumption, the Nvidia jetson (ENvidia(yi)) values are much lower than those of the Raspberry pi’s (ERasp(yi)) because of the fast processing time of each image (Tdl). Batches of images selected from the COCO dataset [34] for object detection are distributed to the available edge devices via the edge server which is a typical edge device (Raspberry Pi’s 4). The edge server sends the images via sockets. The latency of the device that processes the last image is considered as the latency objective. The edge server sends the images via TCP sockets, each managed by a dedicated thread. The results which are labels for each processed image and its related bounding boxes, are subsequently sent back by the end devices to the edge server. Regarding the energy characteristics of our testbed, all devices are connected to a battery with a capacity of 20000mAh (20A for 1 hour). To provide a clear understanding of our system’s energy consumption for example we assume that the edge server sends 100 images to the Raspberry Pi that consumes 200mA for one minute, or 1/60th of an hour, and the consumption of the device comes to 3.3mAh. This indicates that the battery capacity dropped from 20000mAh to 19996.7mAh, i.e., 0.0165% [35].Itis noteworthy that temperature affects battery life and that supply voltage may decrease with increasing load capacity 2480 VOLUME 6, 2025 FIGURE 1. Testbed topology. FIGURE 2. Example of network status for every batch step. TABLE 3. Characteristics of object detection models. over time, both of which may affect the performance of the device, but we simplify our modelling by neglecting the above factors. In our experimental process, we use a total batch of 20000 images with an average size of 150 KB per image. After 10 batch steps of 2000 images, the images located at the edge server are processed. The current network status is considered to solve the defined multi-objective optimization problem, i.e., a set of static problems is used to approximate the actual dynamic problem. Python is used to code the entire distribution method. FIGURE 3. Pareto front surface. B. TESTBED SIMULATION RESULTS In our previous work [28] we proposed a reverse task offloading algorithm that optimizes the task distribution to minimize the latency of the slowest subtask while adhering to energy and accuracy constraints. In this work, we take our method one step further and solve a multi-objective optimization problem to provide our proposed (default) Pareto optimal solution for the three objectives. We illustrate the Pareto front with 100 optimal solutions in Figure 3for one batch step (2000 images) along with a fitting surface to enhance the visualization of the front. Every solution in the Pareto front corresponds to a vector x=[x1,...,xN]∈[0,1]N, a vector y=[y1,...,yN], and a vector w=[w1,w2,w3]where xrepresents the distribution percentage, ythe selected object detection models for each device and wthe three weight values. As expected, when the accuracy increases more resource-intensive object detection models are selected which results in longer processing time and consequently higher energy consumption. In Figure 4, we illustrate the process of computation described in Section III-C. The Center of Gravity (CoG) of the Pareto front corresponds to the point ‘X’ in green color. Then, using the defined distance metric, the Nearest Point to CoG is shown in red dot and corresponds to the decision variables x=[0.271,0.323,0.052,0.304,0.050], y=[1,4,3,3,4] and w=(0.846,0.095,0.059).Thesame procedure for each baseline method is shown in Figure 5. The decision variables for all methods for this typical batch step are shown in Table 4. 1) COMPARISON WITH BASELINE METHODS To provide an extensive evaluation of our proposed method, we compare it with all previously defined baseline methods. VOLUME 6, 2025 2481 AMANATIDIS et al.: MULTI-OBJECTIVE REVERSE OFFLOADING IN EDGE COMPUTING FOR AI TASKS FIGURE 4. Default solution selection of MOEAI-ARO method (normalized values). We execute all methods 100 times, under dynamic network conditions (10 batch steps), and compute the average values for all objectives. The results are summarized in Table 5and detailed in the sequel: •Local: The MOEAI-ARO method outperforms the Local execution solution in two objectives, namely 95.5% for the latency and 1.5% for the accuracy objective. This significant difference in latency is mainly due to the limited processing power of the edge server used. Regarding the energy objective the Local method outperforms the MOEAI-ARO by a difference of 47.1%, which is expected since the Local baseline method does not consume energy for data transmission and reception. •First Available (FA): Our proposed method outperforms the FA method in two objectives, namely 18.5% for the latency and 39.2% for the energy objective. However, the FA method default solution exhibits better performance in accuracy by 1.2%. •Random: The Random method outperforms our proposed method in terms of accuracy by 2% but is significantly worse in terms of latency 94.2% and energy consumption 45.1%. The difference in latency derives from distributing an important part of images to devices with low processing power and slow network connections. Regarding the accuracy objective, the difference of 2% is due to the fact that images were processed using neural networks with higher accuracy. •Round Robin: MOEAI-ARO outperforms Round Robin in all objectives, namely 88.2% for the latency, 48.2% for the energy, and 2.6% for the accuracy objective. This method is very sensitive to the existence of some FIGURE 5. Selection of the default solution among the Pareto front for all baseline methods (normalized values). TABLE 4. Decision variables for the default solution of each method in Figures 4and 5. slow device and/or network connection and is mainly used when the network conditions are unknown. To enhance the visualization of the results, we provide in Figure 6a bar chart that corresponds to Table 5. The above findings show that our method considerably improves the average system performance compared to all tested baseline methods. Even when some baseline method outperforms the MOEAI-ARO solution for some objective, the average gain for all objectives is significantly higher, providing a better equilibrium for the overall performance of the system. Remark: Note that the gain in accuracy shall not be compared in absolute values to the corresponding gains in latency or energy, since the different neural networks 2482 VOLUME 6, 2025