scieee AI-readable full text Open interactive document viewer

Enhancing the Performance of Edge AI Systems via Energy Equilibrium

Amanatidis, Petros; Vagiotas, Eleftherios; Stamopoulos, Athanasios; Michailidis, George; Karampatzakis, Dimitris; Sarigiannidis, Panagiotis; Lagkas, Thomas

Abstract

Nowadays, an enormous increase in the number of edge IoT devices, such as autonomous vehicles and UAVs, has been noticed. These devices have the available computational power to execute AI tasks at the edge of the network without relying on cloud infrastructures. In this work, we explore a scenario where these end nodes can be used to help larger nodes (like cloudlets) in executing AI tasks. In more detail, we formulate an optimization problem in which every node aims to minimize the total execution time, which is determined by the slowest subtask, while ensuring energy equilibrium in the system in order to expand the operational lifetime of the battery-connected devices. We address this complex MINLP (Mixed Integer Non-Linear Problem) optimization problem using an extended version of the previously introduced EAI-ARO algorithm. The results indicate that, in comparison to the previous version, our extended version of the EAI-ARO algorithm improves the system's performance.

Full text

Enhancing the performance of Edge AI Systems via Energy Equilibrium * 1st Petros Amanatidis Department of Informatics Democritus University of Thrace Kavala, Greece [email protected] 2nd Eleftherios Vagiotas Department of Informatics Democritus University of Thrace Kavala, Greece elv[email protected] 3rd Athanasios Stamopoulos Department of Informatics Democritus University of Thrace Kavala, Greece [email protected] 4th George Michailidis Department of Informatics Democritus University of Thrace Kavala, Greece [email protected] 5th Dimitris Karampatzakis Department of Informatics Democritus University of Thrace Kavala, Greece [email protected] 6th Thomas Lagkas Department of Informatics Democritus University of Thrace Kavala, Greece [email protected] Abstract—Nowadays, an enormous increase in the number of edge IoT devices, such as autonomous vehicles and UAVs, has been noticed, which have the available computational power to execute AI tasks at the edge of the network without relying on cloud infrastructures. In this work, we explore a scenario where these end nodes can be used to help larger nodes (like cloudlets) in executing AI tasks. In more detail, we formulate an optimization problem in which every node aims to minimize the total execution time, which is determined by the slowest subtask, while ensuring energy equilibrium in the system in order to expand the operational lifetime of the battery-connected devices. We address this complex MINLP (Mixed Integer Non-Linear Problem) optimization problem using an extended version of the recently released EAI-ARO algorithm. The results indicate that, in comparison to the previous version, our extended version of the EAI-ARO algorithm improves the system’s performance. Index Terms—EdgeAI, Edge Computing, Optimization, Resource Allocation I. INTRODUCTION Edge computing has become very popular in the last few years due to the rapid development of networks and the processing power of IoT devices. The number of devices connected to the Internet is increasing every day, resulting in a large amount of data overload and slow response times. Traditional cloud computing struggles to support the daily needs of today’s society in terms of data processing. For this reason, the utilization of edge computing becomes very important in providing solutions to handle problems such as latency issues, high network instability and energy costs. In typical edge computing systems, tasks are usually offloaded from end nodes to more powerful resources, such as cloudlets This work received funding from the European Union’s Horizon Europe Research and Innovation Programme under Grant Agreement No. 101070181. This paper only reflects the authors’ views, and the Commission is not responsible for any use that may be made of the information it contains. or cloud data centers, which are capable of executing the attributed tasks, resulting in improving the performance of edge computing applications [1]–[4]. An alternative technique of task offloading, called reverse task offloading, allows end nodes (e.g. mobile phones, IoT devices, sensors, etc.) to assist larger nodes (e.g. cloudlets) in processing resource intensive tasks such as AI processing. In fact, edge servers or cloudlets, which are located between the end nodes and the cloud, send their tasks to nearby end nodes. The main idea behind this strategy of offloading is that the increasing computing power of end nodes in combination with the large number of end nodes that often remain idle, results in the creation of an enormous computing capacity at the edge of the network, which can be proven beneficial when used strategically. Furthermore, with the rapid development of Artificial Intelligence (AI) and optimized Deep Neural Network (DNN) models, the ability to infer DNN models on end nodes has been provided, resulting in low latency and energy costs [5]–[7]. Of course, the performance of inferencing the AI models on the end nodes depends strongly on specific characteristics, namely the image input sizes and architecture of the AI models. In more detail, small models save energy and time, unlike larger ones, which offer greater accuracy but increased latency and energy usage. Demonstrating the critical role in minimizing task execution delays and energy consumption of end nodes, task offloading [8], [9] has led to the development of various types of optimization algorithms for task distribution, aiming to enhance the use of resources in edge computing networks [10]. Researchers continue to investigate new task distribution methods that will help to achieve real-time application management while providing an energy equilibrium among devices in the network. In this study, we examine the reverse offloading approach and address a single-objective optimization problem using our newly developed EAI-ARO algorithm [11]. In our previous work, we focused on minimizing the maximum latency, i.e., the time required to process all the images sent from the server to the end nodes, while respecting constraints related to mean accuracy and energy. The mean accuracy constraint ensured that the system operated above a minimum mean accuracy. The energy constraint simply ensured that all end nodes possessed the necessary energy to execute their current assigned tasks. Since this method did not correlate the energy status of each device with the distribution of the images for each device, it may lead to a rapid reduction in the number of active end devices due to the premature consumption of their available energy, potentially resulting in high latencies in subsequent batches of images. On the contrary, limiting energy consumption at the end nodes may result in higher latency for the some batch steps, but the total time required for the entire process may be reduced, especially for unstable networks, thanks to the higher number of available devices. In this work, to solve this complex Mixed Integer Non-Linear Problem optimization problem, we enhance our previously released adaptive task offloading approach to provide implicitly an energy equilibrium to the system. The results highlight notable improvements in total latency, depending on the network conditions. II. RELATED WORKS Reverse offloading: The reverse offloading strategy is becoming more and more popular, because of the high computational power of the end devices. It is important to note that this approach optimizes both resource usage and energy efficiency. Making dynamic decisions about how and where to execute each task was presented in [12], where the researchers proposed a system based on Deep Q-Networks (DQN). The main advantages of this method included the balance between execution speed and energy consumption. Also important, as mentioned in the same study, is the rapid growth of Mobile Edge Computing (MEC) systems, which enhanced the capabilities of the so-called Collaborative Edge Computing Infrastructure Systems (CMIS). The implementation of reverse offloading has resulted in network and energy cost savings, but also an improvement in latency compared to other methods. Another reverse offloading framework was presented in [13], where the authors proposed an algorithm that efficiently treats discrete and continuous decision variables. In this work, a new algorithm was developed, which optimizes the distribution percentages and the selection of object detection models in order to optimize competitive objectives simultaneously. Energy aware Task Offloading: In the domain of energyaware task offloading in Edge AI systems, many studies focus on methods to improve energy efficiency and latency. Researchers in [14] approached task offloading in MEC systems via the Dual-Decomposition method offering a method to optimize energy and latency. Another study [15] used the auction-based method to offload the tasks, aiming to optimize the energy consumption in Edge Computing environments. In addition, energy-efficient offloading strategies have been proposed in the IoT domain, addressing challenges such as transmission delay, energy consumption, and resource allocation, with the objective of optimizing task execution in MEC [16]. Within this context, a decision making framework for MEC in IoT topologies has been proposed, enabling optimal offloading decisions that enhance the Quality of Experience (QoE) for IoT devices. All these methods optimize the performance of their proposed edge computing systems mainly using the reverse task offloading strategy. However, there has been limited contribution regarding the processing of AI tasks while considering the energy availability of end nodes in the optimization process. More specifically, we enhance the EAI-ARO methodology by promoting energy equilibrium, taking into account the energy status of each end node in the network, to further optimize the system’s total task execution time. III. SYSTEM MODEL AND PROBLEM FORMULATION We present the mathematical formulation of the corresponding problem as well as our system model. We list all of the key variables and decision variables we used in Table I. A batch of images Bor the batch step is sent from the edge server to the available end devices (nodes) for processing. The system consists of a set Nof Ndevices that are used to perform an AI task, or more specifically, an object detection task. For object detection, each device iin our system uses a deep learning model. The models’ accuracy, latency, and total energy consumption may vary based on their input image sizes (see Table II). A set of Mneural networks is represented by M={1,2, ..., M}, and for each end device i∈ N, the object detector used is described by its index yi∈ M. Our goal is to build a system with varying processing power in order to reveal sharp trade-offs. More specifically, the Yolov8 [17] object detector is implemented by the devices of our system. Real-time confidence (accuracy) of the detection performed on the transmitted images by the edge server is provided by end devices with various configurations and characteristics. The part of Bsent to the end device i∈[1, N]is represented by the optimization variable xi∈[0,1], which we introduce for each end device i. A vector x= [x1, ..., xN]∈[0,1]N contains all of these variables. The neural network selection used for object detection at each end device is represented by the added discrete optimization variable y= [y1, ..., yN]. Our objective is to reduce the maximum latency, or the amount of time needed to process every image that is sent from the server to the final devices. Here, the latency for each device is defined as the total amount of time required for both the object detection process (Ti dl) and communication with the server (Ti tx), i.e.: Li(xi, yi) = xiBTi tx +xiBTi dl(yi) = Cixi(1) , where: Ci=B(Ti tx +Ti dl(yi)).(2) TABLE I: Key Parameters and Variables. Description Parameters/Variables xiPortion of the total number of images assigned to device i. yiObject detector selection for every device i. BBatch of images for processing (Batch step). Ti tx Time to transmit one image to device i. Ti dl(yi) Processing time per image for the device iusing the object detector yi. Li(xi, yi)Time for transmitting and processing the images for device i. mAP (x, y)Mean accuracy of the overall object detection process. mAP (yi)Mean accuracy per image of the object detection process using the neural network yi. E(xi, yi)Total energy consumed for the object detection process. Ei(yi)Energy consumed by device ifor the object detection process of one image using neural network yi. Ei available Energy available for device i. x0Auxiliary variable that serves simultaneously as cost function and uniform bound for the latency of each device i. xbound iBound on the maximum allowed distribution percentage according to the energy availability of device i. pExponent pis a tuning parameter used in the formulation of the bounds xbound i. TABLE II: Characteristics of object detection models. yi index Input Image Size mAP (yi)ERasp(yi) (Joule/image) ENvidia(yi) (Joule/image) 1 224×224×3 0.33 0.6 0.15 2 320×320×3 0.438 1.22 0.195 3 480×480×3 0.497 2.16 0.305 4 640×640×3 0.541 4.8 0.55 Furthermore, in real-world applications, it is necessary to consider the energy availability of each device to make sure that it is capable of carrying out the assigned tasks. For every device, this adds an additional constraint, s.t. Ei(xi, yi)≤Ei available (3) , where: Ei(xi, yi) = xiB E(yi)(4) and E(yi)is the energy used by the neural network yito detect objects in a single image. In addition, we include an accuracy constraint that requires the processing procedure’s minimum normalized accuracy to be higher than a user-defined threshold value. We define the normalized mean accuracy of the entire object detection process as follows: A(x, y) = mAP (x, y)−mAPmin mAPmax −mAPmin (5) , where: mAP(x, y) = X i mAP(yi)xi(6) and mAPmin, mAPmax stand for the lowest and highest mean Average Precision that can be achieved, or the average precision that the smallest and largest neural networks can provide, respectively. The optimization problem is then stated as follows: min x,y max iCixi s.t. Ei(xi, yi)≤Ei available, i = 1, ..., N A(x, y)≥Amin, Amin ∈[0,1] PN i=1 xi= 1 xi∈[0,1], i = 1, ..., N x0≥0 (7) To avoid the non-differentiability of the max operator in (7), we reformulate the optimization problem in accordance with [18]. We introduce x0∈[0,∞), a positive auxiliary optimization variable that functions as a uniform bound for the latency of all devices as well as the cost function. The equivalent formulation of the optimization problem reads: min x,y,x0 x0 s.t. Cixi≤x0, i = 1, ..., N Ei(xi, yi)≤Ei available, i = 1, ..., N A(x, y)≥Amin, Amin ∈[0,1] PN i=1 xi= 1 xi∈[0,1], i = 1, ..., N x0≥0 (8) The above optimization problem has been efficiently used in [11] for relatively small tasks. However, when the total number of images is large and several batch steps are required for the task to be completed, the algorithm may consume prematurely the disposed energy of battery-powered devices. In the unfavourable scenario of unstable networks, this can lead to increased latency during the process, due to the reduced number of available devices, which reduces the alternatives for the distribution process. Therefore, to improve the robustness of the distribution process in similar situations, it seems interesting to modify the above formulation and include some information related to a notion of energy equilibrium among battery-powered devices. A simple solution towards this direction consists of modifying the upper bounds of the variables xi, limiting the number of images assigned to devices with low energy availability. We propose the following modification of the bound constraints: xi∈[0, xbound i], i = 1, ..., N , where: xbound i=min 1.0,Ei available Ei limit p.(9) In equation (9), the parameter Ei limit defines the limit energy value for each device i, after which the upper bound is modified, since for Ei available ≥Ei limit one gets xbound i= 1.0. Using a high Ei limit value, bounds are modified early in the task offloading process. Moreover, the exponent parameter pdetermines how fast the bounds are modified. Increasing the value of preduces more drastically the value of xbound i. Higher values are expected to perform more efficiently for increasing network instability. The final formulation of the optimization problem reads: min x,x0,y x0(10.1) s.t. Cixi≤x0, i = 1, ..., N (10.2) Ei(xi, yi)≤Ei available, i = 1, ..., N (10.3) A(x, y)≥Amin, Amin ∈[0,1] (10.4) PN i=1 xi= 1 (10.5) xi∈[0, xbound i]i= 1, ..., N (10.6) x0≥0(10.7) (10) For fixed yvector the defined optimization problem (10) represents a linear programming (LP) problem that can be solved using e.g. the SIMPLEX algorithm [19], [20]. More specifically, only the auxiliary variable x0is present in the objective function (10.1), while (10.2) displays the extra constraints due to the reformulation of the problem. The energy constraints appear in (10.3), guaranteeing that every device disposes the necessary energy to complete the assigned tasks. In (10.4) the accuracy constraint is provided, ensuring that the minimum normalized accuracy is higher than a user-defined threshold value. The sum of the parts {xi}i=1,...,N of Bsent to the end device must equal 1 according to equation (10.5). Lastly, bound constraints for the optimization variables are found in equations (10.6) and (10.7). Remark. There are several different ways to achieve energy equilibrium in the system. For example, another idea is to use the standard deviation of the energy consumption among these devices and set an additional constraint, or even to include it in the objective function and formulate a multi-objective optimization problem. For both options, this choice would turn the optimization problem (10) for fixed yto become non-linear and therefore more challenging to solve in acceptable time, as it is solved for every member of the GA population. On the contrary, the proposed approach allows us to keep the defined optimization problem linear and easily solvable using common linear programming algorithms. IV. SOLUTION OF THE OPTIMIZATION PROBLEM The methodology to solve the defined MINLP optimization problem (10) is algorithmically presented in the following pseudocode. Algorithm 1 MINLP Solving Algorithm 1: g←0, create initial population B(0) p 2: while not converged do 3: for i:= 1 to λdo 4: Random choice of parents Pi 5: Oi←recombination(Pi) 6: ¯ Oi←mutation(Oi) 7: Solve LP problem (10) for ¯ Oi 8: Calculate fitness Fifor ¯ Oi 9: end for 10: B(g) o← { ¯ Oi, Fi}i=1,...,λ 11: B(g) p←best(B(g) p, B(g) o, µ) 12: g←g+ 1 13: end while Algorithm 1 combines a Genetic Algorithm and a Linear Programming algorithm to efficiently solve the defined MINLP optimization problem. In more detail, line 1 generates the population with µ y vectors representing combinations of object detector indices. In the optimization loop (lines 213), the offspring population is generated. Two members of the population are randomly selected and recombined via crossover (line 5) to create two offspring. To obtain the final set of offspring, mutation is applied to a small percentage randomly selected (line 6). The LP problem (10) is solved in line 7 and latency is calculated to evaluate the fitness function (line 8). In line 10, both offspring and parent populations are collected to select the best individuals using the ”µ+λ” selection operator in line 11. Finally, the generation index is updated (line 12). V. RESULTS A. Testbed Our testbed, illustrated in Fig. 1, is similar to the one presented in [13]. It consists of two Nvidia Jetson Nano devices with 4GB RAM and two Raspberry Pi 4 which also have RAM capacity. It also consists of an edge Server and a Wi-Fi access point (AP). All these devices are equipped with a YOLOv8 model, which is mainly used for detecting objects in images. For the edge Server, we used a laptop with 8GB of RAM which we have assigned to distribute batches of images to the user’s end devices via TCP sockets, each of which is managed by a separate thread. Our images come from the COCO dataset and are approximately 150KB in size. The bounding boxes and the prediction labels are sent from the devices back to the edge server for processing. All end devices are powered by a battery, while for simplicity reasons the effect of temperature or the change in operating voltage due to load is not taken into account. The implementation of communication and management of image batches has been TABLE III: Bandwidth intervals which represent stable (low variance) and non-stable (high variance) network conditions. Network Conditions Network Speed Interval (MB/s) Stable 4-7 Non-stable 1.2-7 done entirely in Python. A total of 10000 images are used, which are distributed and processed in 100 batches of 100 images. Fig. 1: Testbed topology. B. Testbed Simulation Results For the numerical results, we compare the previous version of our recently released EAI-ARO algorithm with its enhanced version introduced in this work. More specifically, we compare the gain in total latency both for stable (low variance (MB/s)) and non-stable network conditions (high variance). A typical example of the network’s status is provided in Table III. We arbitrarily fixed the normalized accuracy threshold Amin to 0.47, forcing resource-intensive object detection models to be used, to satisfy the accuracy constraint. Furthermore, the Ei limit value was defined to equal the maximum available energy for each the battery-powered devices in our testbed. We compare the performance for various values of the parameter p, namely for p= 0, ..., 4. Remark. For p= 0 the optimization problem corresponds to the original version of the EAI-ARO algorithm and serves as reference for the comparisons performed in the sequel. In Fig. 2, we illustrate the evolution in latency for 100 batch steps for stable network conditions. During the first steps, the bound constraints on the distribution percentage remain inactive and the performance of the algorithm is similar for all values of p. In fact, differences up to this point are merely due to the randomness of the network. As soon as some bound constraints become active and all devices are active, TABLE IV: Gains in total latency considering stable network conditions. pvalue Gains (%) 1 2.5 2 2.8 3 5.1 4 4.6 TABLE V: Gains in total latency considering non-stable network conditions. pvalue Gains (%) 1 6.8 2 7.8 3 9.6 4 8.5 the performance of the algorithm for p≥1is worse than the reference, as expected, since the algorithm has less freedom to distribute tasks to more powerful nodes with better network conditions. However, in later steps, devices for p= 0 become inactive earlier during the process, resulting in higher jumps in the latency value. In fact, for p= 0 all battery-powered devices have become inactive until batch step 75 and all tasks are processed by the server itself. This situation appears in later steps for p≥1. We provide the total latency for all pvalues in Fig. 3. The performance gains are further provided in Table IV, which shows an improvement in the range of 2.5-5.1%. This difference is a notable gain considering that we compare already with an optimal solution. Fig. 2: Evolution of latency for every batch step under stable network conditions, considering different pvalues. Going our evaluation process a step further, we show the performance gains for non-stable network conditions. We illustrate the corresponding evolution of latency at every batch step in Fig. 4. Additionally, we provide also the results for the total Fig. 3: Performance in total latency for different pvalues considering stable network conditions. latency in Fig. 5 along with the corresponding performance gains in Table V, which shows an improvement in the range of 6.8-9.6%. We observe that, for both network conditions, the optimal total latency is obtained for p= 3, which seems to provide a good compromise between low latency at some specific step and low distribution bounds to ensure a sufficiently long operational lifetime of the device. This observation may change depending principally on the size of the total batch of images and the network conditions. Fig. 4: Evolution of latency for every batch step for different pvalues considering non-stable network conditions. VI. CONCLUSION This paper introduced an enhanced version of our previously released adaptive reverse task offloading algorithm for Fig. 5: Performance in total latency for different pvalues considering non-stable network conditions. executing object detection processes in an edge computing environment. We suggested a novel approach to extend the operational lifetime of the available end devices in order to minimize the total completion time of the object detection process on the transmitted images from the edge server to the end devices. The newly presented enhanced algorithm was extensively evaluated against the original EAI-ARO formulation in a heterogeneous real-world testbed. The derived results show that the introduced modification notably improves the overall performance, reaching up to 10% further latency reduction. The findings provide an opportunity for further research, including a dynamic version of the proposed algorithm, to approximate the defined optimization problem via real-time updates. REFERENCES [1] W. Shi and S. Dustdar, “The promise of edge computing,” Computer, vol. 49, no. 5, pp. 78–81, 2016. [2] G. Premsankar, M. Di Francesco, and T. Taleb, “Edge computing for the internet of things: A case study,” IEEE Internet of Things Journal, vol. 5, no. 2, pp. 1275–1284, 2018. [3] Y. Ai, M. Peng, and K. Zhang, “Edge computing technologies for internet of things: a primer,” Digital Communications and Networks, vol. 4, no. 2, pp. 77–86, 2018. [4] K. Cao, Y. Liu, G. Meng, and Q. Sun, “An overview on edge computing research,” IEEE access, vol. 8, pp. 85 714–85 728, 2020. [5] H. Li, K. Ota, and M. Dong, “Learning iot in edge: Deep learning for the internet of things with edge computing,” IEEE network, vol. 32, no. 1, pp. 96–101, 2018. [6] A. Marchisio, M. A. Hanif, F. Khalid, G. Plastiras, C. Kyrkou, T. Theocharides, and M. Shafique, “Deep learning for edge computing: Current trends, cross-layer optimizations, and open research challenges,” in 2019 IEEE Computer Society Annual Symposium on VLSI (ISVLSI). IEEE, 2019, pp. 553–559. [7] Y. Huang, X. Ma, X. Fan, J. Liu, and W. Gong, “When deep learning meets edge computing,” in 2017 IEEE 25th international conference on network protocols (ICNP). IEEE, 2017, pp. 1–2. [8] L. Lin, X. Liao, H. Jin, and P. Li, “Computation offloading toward edge computing,” Proceedings of the IEEE, vol. 107, no. 8, pp. 1584–1607, 2019. [9] M. Aazam, S. Zeadally, and K. A. Harras, “Offloading in fog computing for iot: Review, enabling technologies, and research opportunities,” Future Generation Computer Systems, vol. 87, pp. 278–289, 2018. [10] A. Islam, A. Debnath, M. Ghose, and S. Chakraborty, “A survey on task offloading in multi-access edge computing,” Journal of Systems Architecture, vol. 118, p. 102225, 2021. [11] P. Amanatidis, D. Karampatzakis, G. Michailidis, T. Lagkas, and G. Iosifidis, “Adaptive reverse task offloading in edge computing for ai processes,” Computer Networks, vol. 255, p. 110844, 2024. [12] M. M. Saeed, R. A. Saeed, R. A. Mokhtar, O. O. Khalifa, Z. E. Ahmed, M. Barakat, and A. A. Elnaim, “Task reverse offloading with deep reinforcement learning in multi-access edge computing,” in 2023 9th International Conference on Computer and Communication Engineering (ICCCE). IEEE, 2023, pp. 322–327. [13] P. Amanatidis, G. Michailidis, D. Karampatzakis, V. Kalenteridis, G. Iosifidis, and T. Lagkas, “Multi-objective reverse offloading in edge computing for ai tasks,” IEEE Open Journal of the Communications Society, 2025. [14] A. Younis, T. X. Tran, and D. Pompili, “Energy-latency-aware task offloading and approximate computing at the mobile edge,” in 2019 IEEE 16th international conference on mobile ad hoc and sensor systems (MASS). IEEE, 2019, pp. 299–307. [15] L. Li, X. Zhang, K. Liu, F. Jiang, and J. Peng, “An energy-aware task offloading mechanism in multiuser mobile-edge cloud computing,” Mobile Information Systems, vol. 2018, no. 1, p. 7646705, 2018. [16] J. Li, M. Dai, and Z. Su, “Energy-aware task offloading in the internet of things,” IEEE Wireless Communications, vol. 27, no. 5, pp. 112–117, 2020. [17] G. Jocher, A. Chaurasia, and J. Qiu, “Yolo by ultralytics. 2023,” URL: https://github. com/ultralytics/ultralytics, 2023. [18] J. Taylor and M. P. Bendsøe, “An interpretation for min-max structural design problems including a method for relaxing constraints,” International Journal of Solids and Structures, vol. 20, no. 4, pp. 301–314, 1984. [19] G. B. Dantzig and M. N. Thapa, Linear programming: Theory and extensions. Springer, 2003, vol. 2. [20] J. A. Nelder and R. Mead, “A simplex method for function minimization,” The computer journal, vol. 7, no. 4, pp. 308–313, 1965.