CloudEdge Fusion: A Cloud-Edge Continuum Computing Platform
Abstract
This abstract discusses an ambitious project that we are pursuing named CloudEdge Fusion, i.e. to enable the development and management of a cloud-edge continuum infrastructure of interconnected resources across the cloud, regional data centers, and multi-access edge computing (MEC) data centers, capable of continuously supporting large numbers of widely distributed edge devices and applications.
Full text
CloudEdge Fusion: A Cloud-Edge Continuum Computing Platform Ryousei Takano∗†, Tomohiro Kudoh†, Hiro Kishimoto†, Takahiro Hirofuchi†, Yutaka Oiwa†, Yusuke Tanimura†, Takeharu Kato†, Koji Yokohata‡, Daiki Orihara‡ †National Institute of Advanced Industrial Science and Technology (AIST), Japan ‡SoftBank Corp., Japan ∗[email protected] Cloud-edge continuum [1] [2] accelerates the emergence of applications where physical space and cyberspace are tightly integrated, such as autonomous driving, automatic manufacturing by robots in a factory, human augmentation for enhancing human abilities, etc. Anticipated post-5G and 6G technologies promise to simultaneously provide low latency and high reliability, massive connections, large capacity, and security. In the post-5G era, all kind of “things” in physical space are connected to the network, and huge amounts of realworld data (RWD) are generated from time to time. Such RWD is collected in cyberspace. In cyberspace, analysis results are fed back to humans and machines in physical space in various forms. Each application has different requirements in terms of processing capacity, latency, security, etc. To realize such RWD processing distributed applications that inherit post-5G capabilities, computing must have the same characteristics as post-5G communications. Edge computing, which uses computing resources near the data source to process data, is one solution, but there are issues such as large-scale optimization, secure data sharing over wide area, and support for mobility management [3] [4]. To achieve this goal, we are pursuing an ambitious project named CloudEdge Fusion, i.e. to enable the development and management of a cloud-edge continuum infrastructure of interconnected resources across the cloud, regional data centers, and multi-access edge computing (MEC) data centers, capable of continuously supporting large numbers of widely distributed edge devices and applications. It enables application developers to create and deploy novel highly responsive, scalable, and reliable services upon a cloud-edge continuum infrastructure. Based on an application requirement, for example, time-sensitive processing should be located close to the user or edge device, while high-performance or cost-effective processing should be located in the cloud. Although platform technologies are required to provide a shared and dynamic service execution environment with high reliability, there is a technological gap with the current cloud and edge computing technologies. Here, “reliability” implies that the system works correctly, that it works within a certain period of time, that it is secure, and that it can be traced if something goes wrong. Ideally, data processing latency and throughput should be fixed, but it is difficult to guarantee it on a shared platform. The cost of guaranteeing 100% reliability is exponentially higher. Therefore, we introduce the concept of probabilistic reliability. The service level is guaranteed probabilistically based on the assumption that there is some fluctuation in processing capacity. In other words, it allows the developers to balance reliability and cost by matching probability with application requirements. The key challenges in bridging the gap between the existing computing platform and our envisioned platform are as follows: •High probabilistic reliability in terms of service latency: Conventional computing technologies prioritize throughput and cost efficiency, and pay little attention to latency, which is important in RWD processing. Tail latency [5] [6] is a well-known problem in distributed systems. Under a high degree of parallelism, the performance impact of tail latency significantly increases. In cloudedge continuum infrastructure, short end-to-end connections enabled in part by geo-positioning resources, so that worst-case delays in the data pipeline do not exceed service request response requirements. To improve temporal determinism, i.e., to reduce fluctuations in latency, several system software techniques are worth considering, including lightweight virtualization [7] [8], hardwareassisted acceleration and offloading [9] [10], and time sensitive networking [11]. •Shared and dynamic service provisioning: Most 5G PoC experiments use dedicated computing infrastructure, which is a major barrier to practical service deployment. Adaptive resource management across the cloud-edge continuum infrastructure is required to meet the demands of multiple services and dynamically ensure service level agreement (SLA) for any user, anytime, anywhere. In addition, a mechanism is needed to predict and adjust the situation in a timely manner to prevent service interruptions or other SLA violations. This requires sophisticated resource allocation and scheduling techniques capable of co-scheduling and co-provisioning the entire system. •Unified security model: A wide range of stakeholders are expected to securely store, share, and use data, including personal and confidential corporate information. There is also a need for real-time access control for large numbers of network edges and device edges, but no such framework is established. Flexible security policies
(1) Service Execution Technology (2) Data Utilization Technology (3) System Integration Technology •Resource management technology to improve time determinacy •Network access control technology to achieve “zero trust” in cloudedge continuum •Data processing technology to efficiently utilize large and diverse data by leveraging cloudedge continuum resources •Pseudo-data technology to facilitate privacy preserving data utilization System integration of the R&D results from (1) and (2) with existing software Resource Provider Data provider Demonstrati on of the proposed platform through industrial application services (4) System Demonstra tion Service Provider (5) Social Implementation and Standardization Service Execution Environment Fig. 1. The overview of CloudEdge Fusion Project. The work package consists of (1) service execution technology, (2) data utilization technology, (3) system integration technology, (4) system demonstration with realistic use cases, and (5) promotion of social implementation and standardization. should be defined and enforced to minimize uncertainty through access control that grants the minimum necessary privileges to users, network devices, and device edges. It is also important to monitor records of these authorizations and access controls to ensure traceability. •Service model and programmability: A novel service model that supports probabilistic reliability is essential. We parameterize the fluctuations in performance provided by the resource and define the quality of service that can be provided. A platform combines sufficient resources that are available at different locations to enable successful application deployment. In addition, the heterogeneity, distribution, and the post-5G requirements should be hidden away and abstracted into an application programming environment that is easy for developers to use. High-level APIs, standard service development tools, and domain specific languages should be considered. The CloudEdge Fusion project [12] is a five-year project through March 2027 with nine organizations led by AIST and SoftBank Corp. It aims to establish cloud-edge continuum computing platform technologies that can scale across multiple data centers and hundreds of millions of devices, and to promote social implementation as a platform service business. ACKNOWLEDGEMENT This paper is based on results obtained from ”Research and Development Project of the Enhanced Infrastructures for Post5G Information and Communication System” (JPNP20017), commissioned by the New Energy and Industrial Technology Development Organization (NEDO). The authors thank Professor Jose Fortes at University of Florida and all the members of the CloudEdge Fusion project for valuable discussions. REFERENCES [1] D. Balouek-Thomert, E. G. Renart, A. R. Zamani, A. Simonet, and M. Parashar, “Towards a computing continuum: Enabling edge-to-cloud integration for data-driven workflows,” The International Journal of High Performance Computing Applications, vol. 33, no. 6, pp. 1159– 1174, 2019. [2] P. Beckman, J. Dongarra, N. Ferrier, G. Fox, T. Moore, D. Reed, and M. Beck, “Harnessing the computing continuum for programming our world,” Fog Computing: Theory and Practice, pp. 215–230, 2020. [3] Y. Mao, C. You, J. Zhang, K. Huang, and K. B. Letaief, “A survey on mobile edge computing: The communication perspective,” IEEE communications surveys & tutorials, vol. 19, no. 4, pp. 2322–2358, 2017. [4] N. Hassan, K.-L. A. Yau, and C. Wu, “Edge computing in 5g: A review,” IEEE Access, vol. 7, pp. 127 276–127 289, 2019. [5] J. Dean and L. A. Barroso, “The tail at scale,” Communications of the ACM, vol. 56, no. 2, pp. 74–80, 2013. [6] J. Li, N. K. Sharma, D. R. Ports, and S. D. Gribble, “Tales of the tail: Hardware, os, and application-level sources of tail latency,” in Proceedings of the ACM Symposium on Cloud Computing, 2014, pp. 1–14. [7] A. Agache, M. Brooker, A. Iordache, A. Liguori, R. Neugebauer, P. Piwonka, and D.-M. Popa, “Firecracker: Lightweight virtualization for serverless applications,” in 17th USENIX symposium on networked systems design and implementation (NSDI 20), 2020, pp. 419–434. [8] H. Tazaki, A. Moroo, Y. Kuga, and R. Nakamura, “How to design a library os for practical containers?” in Proceedings of the 17th ACM SIGPLAN/SIGOPS International Conference on Virtual Execution Environments, 2021, pp. 15–28. [9] M. Tork, L. Maudlej, and M. Silberstein, “Lynx: A smartnic-driven accelerator-centric architecture for network servers,” in Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems, 2020, pp. 117–131. [10] P. Shantharama, A. S. Thyagaturu, and M. Reisslein, “Hardwareaccelerated platforms and infrastructures for network functions: A survey of enabling technologies and research studies,” IEEE Access, vol. 8, pp. 132 021–132 085, 2020. [11] L. L. Bello and W. Steiner, “A perspective on ieee time-sensitive networking for industrial communication and automation systems,” Proceedings of the IEEE, vol. 107, no. 6, pp. 1094–1120, 2019. [12] CloudEdge Fusion Project. (to be available soon). [Online]. Available: https://cloudedge-fusion.org/