Serverless Data Analytics (Finally) Bridging the Gap: Introducing the Ortzi DataFrame
Full text
Accepted for publication in the Proceedings of the IEEE 18th International Conference on Cloud Computing (CLOUD), July 2025 Serverless Data Analytics (Finally) Bridging the Gap: Introducing the ORTZI DATAFRAME Germ´ an T. Eizaguirre∗, Marc Hostau†, Marc S´ anchez-Artigas Department of Computer Engineering and Mathematics Universitat Rovira i Virgili, Tarragona {germantelmo.eizaguirre,marc.hostau,marc.sanchez}@urv.cat Abstract—Serverless technologies have simplified distributed computing by streamlining resource management and offering out-of-the-box usability of cloud resources. However, serverless computing has yet to fully permeate the broader data analytics community. One of the major reasons causing this slow adoption is the lack of a handy serverless interface for seamlessly running recurrent workloads in the cloud. To fill this gap, we introduce in this work the missing piece in serverless analytics: the ORTZI DATAFRAME, a practical and intuitive programming abstraction that mirrors pandas DataFrames, so that users can effortlessly run their local, single-threaded Python code at scale in the cloud. Needless to say, such a powerful abstraction is certainly useless if not backed by a serverless analytics system that can operate over it in parallel. For this reason, another major contribution of this paper is a fully-fledged system that can run jobs in parallel across the cloud continuum using the novel ORTZI DATAFRAMES. The new system leverages the specific capabilities of each serverless backend without user intervention. Our evaluation demonstrates that ORTZI enables exploration of the nuanced trade-offs of heterogeneous backends with minimal programming changes and overhead. By harnessing the seamless nature of ORTZI, we optimize jobs through strategic backend selection, still delivering a user-friendly open source framework for programmers without cloud expertise. Index Terms—Distributed computing, serverless computing, cloud computing, data analytics, programming models I. INTRODUCTION Cloud computing offers virtually unlimited computing capabilities on demand, smoothly scaling resources to meet user requirements. Beyond broad scalability, the cloud provides a rich ecosystem of specialized services tailored to diverse application needs. However, operational tasks and the steep learning curve required to master cloud environments frequently deter general programmers from fully leveraging cloud resources. To address cloud usability challenges, serverless computing simplifies the interaction between programmers and infras- © 2024 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. We appreciate all anonymous reviewers at CLOUD ’25, who provided insightful feedback. This work has been partly funded by the EU Horizon programme under the grant agreements: 101092644 (NearData), 101092646 (CloudSkin), 101093110 (EXTRACT) and 101086248 (CLOUDSTARS). Germ´ an T. Eizaguirre is recipient of a pre-doctoral FPU grant from the Spanish Ministry of Universities (ref. FPU21/00630). Marc S´ anchez-Artigas is a Serra H´ unter Fellow. tructure by abstracting resource management away. Serverless technologies, such as cloud functions and object storage, provide rapid scalability, simplified resource provisioning, and fast deployment times through simple APIs. These features reduce the complexity of cloud development, and enable to request precisely the resources needed, exactly when needed. The long promise of serverless extends to many domains, and particularly to data analytics [1]. Hosting data analytics jobs in the cloud can often be time-consuming, and serverless computing alleviates this hurdle. Ideally, an analyst without cloud expertise should be able to launch data analytics jobs in the cloud straightaway, without worrying about the underlying setup. However, the goal of fully serverless data analytics is in reality far from being achieved, primarily due to three unresolved challenges. These are: 1. No general abstraction for serverless data analytics. The strong synergy between serverless computing and data analytics has led to fruitful development of prototypes. However, to the best of our knowledge, there is no a simple and generalpurpose interface to run serverless data analytics effortlessly in the cloud. In short, the existing open source frameworks can be grouped into three classes: •Prototypes with no general programming abstractions: Jolteon [2], Ditto [3] or Minflow [4]; •Tools with usable data abstractions, but a limited set of operators: PyWren [5], Lithops [6], and Seer [7]; •Tools with usable programming abstractions, but tailored to a specific system/backend requiring manual setup: Pixels-Turbo [8], Wukong [9], Flint [10] or Dexter [11]. 2. The lack of open-source serverless analytics systems. As of today, most of the state-of-the-art serverless data analytics systems have not been open-sourced. Unfortunately, the list is rather long, but some of the most relevant ones are Locus [12], Lambada [13], SONIC [14], Caerus [15], and Smartpick [16]. Code inaccessibility forces other researchers to independently recreate these ideas based on in-paper references. For example, Jolteon (re)implements Caerus’ scheduling 1. Although such reinterpretations are inevitable, they hinder the establishment of general standards for serverless data analytics systems. 3. The cloud continuum is nowadays a de facto standard. Heterogeneous compute backend pools are not an exception 1https://github.com/pkusys/Jolteon/blob/workflow/scheduler.py
but a rule. Modern analytics systems combine on-premise clusters, cloud resources and even high-performance laptops. Yet, most serverless frameworks [2]–[7], [9], [11], [13] overlook this existing computing power, opting instead to offload entire workloads to homogeneous serverless services. We argue that serverless frameworks should not be confined to pure serverless backends. Instead, they should deliver a serverless experience while leveraging all available resources, enabling cost and performance optimizations. For example, frameworks could prioritize already provisioned resources while exploiting cloud functions when necessary. Offloading workloads to cloud functions has been previously proposed, for example, to overlap partial workloads with VM deployment times [8] or to handle sudden load spikes [17]. We believe these approaches should be extended to general data analytics programming, granting users on-demand and managed access to the flexibility of cloud functions. Yet, this holistic vision for leveraging the cloud continuum remains largely unrealized in practical terms. Many contemporary serverless frameworks still narrowly focus on specific cloud services, while even established data processing systems like Spark [18], Ray [19], or Dask [20] typically lack outof-the-box continuum support and demand complex manual configuration even for basic homogeneous cluster setups. This highlights a significant gap in providing truly seamless access to today’s diverse computing environments. Our contribution We introduce ORTZI, a serverless data analytics framework designed for heterogeneous computing environments. With its familiar pandas-like DATAFRAME abstraction, ORTZI enables developers to trivially write custom data analytics jobs, which are automatically parallelized across the available computing backends. Moreover, ORTZI orchestrates end-to-end workload execution, leveraging the unique capabilities of each compute backend to optimize job performance and data exchanges. The key contributions of this work are as follows: •An open-source, general-purpose programming abstraction for serverless data analytics, featuring a comprehensive set of built-in primitives, ranging from simple map to complex sort and groupBy operators. •An adaptive serverless architecture to conduct seamless job parallelization across the cloud continuum, dynamically adjusting to the architectural traits of each backend. •Amanaged data exchange system for complex data analytics, focusing on decoupled and heterogeneous computing environments. ORTZI delivers performance comparable to state-of-the-art serverless frameworks in a pure cloud functions deployment, with only a 6.78% overhead. By transparently leveraging heterogeneous backends, ORTZI delivers a 25.57% performance improvement over Jolteon [2] at comparable cost and reduces latency by 70.59% compared to EMR Serverless [21]. II. ORTZI OVERVIEW ORTZI is a programming platform for data analytics that automatically provisions and leverages computing resources with minimal credential setup. This mirrors the experience of commercial, serverless data analytics products like Coiled [22] or Anyscale [23], but as a fully open-source solution with seamless compatibility across the cloud continuum. The ORTZI DATAFRAME.ORTZI is built on the ORTZI DATAFRAME, a practical abstraction with a Dask [20]- and pandas [24]-like API. This ensures familiarity and ease of adoption, offering an experience similar to local programming. Listing 1shows an example of data analytics programming with the ORTZI DATAFRAME. Internally, a ORTZI DATAFRAME is divided row-wise into partitions. Each partition is a separate instance of a Polars [25] DataFrame. ORTZI automatically distributes the partitions across the underlying compute backend to enable the parallel processing of an ORTZI DATAFRAME. Also, ORTZI manages the persistence of partitions in a transparent way for the user, and stores a partition in memory, on disk, and in an external storage system, depending on the computing needs. In terms of data mutability, ORTZI DATAFRAMES support two main types of transformations: (a) narrow transformations; and (b) wide transformations. Narrow transformations require only one-to-one dependencies between partitions and include primitives to apply a function to the data from all the partitions in parallel, such as map and filter. The parallelization of such operations is straightforward due to the absence of crossdependencies. A far more complex aspect are ORTZI wide transformations, as they require shuffling partitions [7]. Wide transformations, such as sort and groupBy, involve all-to-all communication patterns. ORTZI abstracts the complexities of these patterns from the user by automatically handling the parallel exchange of partitions between the compute nodes. 1from ortzi import DataFrame 2from ortzi.backends import BackendType 3from some_lib import encrypt_func, some_func 4from numpy import mean 5 6sdf = DataFrame.from_csv( 7"s3://my-bucket/data.csv", 8on_premise=True 9) 10 sdf = sdf.sort(by="col1") 11 sdf = sdf.map(encrypt_func) 12 sdf = sdf.set_scaleup(BackendType.CLOUD_FUNCTION) 13 sdf = sdf.map(some_func) 14 result = sdf.reduce(mean) Listing 1: Serverless data analytics in the cloud continuum using ORTZI abstractions. To materialize results and present results to the user, ORTZI implements a set of actions, such as collect and write. Actions define computation boundaries. In other words, when ORTZI finds an action, it runs all the pending transformations
to that point. This is visible in line 15 of Listing 1, where the call to the reduce action triggers the execution of all the transformations from line 6to line 13. ORTZI DATAFRAMES offer specific functionalities tailored for programming in the cloud continuum. For example, the programmer can enforce the execution of certain operations on proprietary devices using the on_premise flag. This feature is especially useful for privacy-sensitive data that cannot be offloaded to the cloud in its raw format. Additionally, users can choose to scale resources (if necessary) using cloud functions, cloud VMs, or both with the set_scaleup function. Concept Definition Task The smallest unit of execution, typically processing a subset of the data. Stage A grouping of functionally homogeneous tasks that share dependencies and can run concurrently. Exchange The data transfer between two or more stages, involving data shuffling or redistribution among tasks. Partition Logical subset of the data being analyzed, processed by a single task. TABLE I: Key concepts in ORTZI’s design. Job compilation. ORTZI internals are designed around widely used patterns in distributed data analytics. Table Iintroduces the foundational concepts integral to ORTZI’s architecture. ORTZI DATAFRAMES adopt a lazy execution model for job computation, inspired by Spark RDDs [26]. Transformations, such as map and sort, are deferred until an action, such as collect, is invoked. As the user adds transformations through the interface, the internal application constructs a logical plan, which is a highlevel representation of the operations required to process and analyze the data. The logical plan consists of stages organized as a directed acyclic graph (DAG). Narrow transformations are pipelined within the same stage, while wide transformations define stage boundaries. At the end of each stage, tasks write their partitions to intermediate storage. These partitions are shuffled, with each task in the subsequent stage pulling its input partitions. We refer to this process of shuffling partitions as an exchange. Figure 1a illustrates the steps in generating the logical plan for a sort operator, as the one in Listing 1. Sorting is implemented in a distributed manner; therefore, it comprises two stages connected by a data exchange. Upon invoking an action, the application automatically compiles the final physical plan based on the logical plan. In ORTZI, the physical plan is formed by breaking down the logical plan into executable tasks, defining partitions, and establishing dependencies. Tasks are versatile, meaning that they can be executed in any underlying backend. In our current implementation, the number of tasks in narrow transformations is determined using a static partition size heuristic. For wide transformations, however, the number of partitions per stage is inferred using the Seer resource provisioning algorithm [7]. We depict a physical plan example for the sort operator in Figure 1a. Adaptive execution. To execute the job, we schedule tasks cross available resources. A worker is defined as a process responsible for receiving and executing these tasks. Workers optimize data exchanges by dynamically adapting to the underlying storage backend, enabling seamless data exchanges across ephemeral, persistent and external storage. Workers utilize shared memory devices for exchanges between collocated tasks, as illustrated in Figure 1b. When shared memory is unavailable—common in most commercial cloud function platforms [1]—workers seamlessly fall back to disk for local exchanges. For exchanges between distributed tasks, we transparently utilize a external storage service, reachable from the entire ORTZI architecture. A partition may reside in the local storage of one worker but be required by a remote worker. In such cases, ORTZI automatically signals the local worker to push the partition to external storage, making it accessible to all. We refer to this type of exchange, which involves multiple storage backends, as a hybrid exchange, as illustrated in Figure 1d. df.sort() 1 2 Stage 0 Exchange Stage 1 T0.0 T0.1 T0.2 T1.0 T1.1 T1.2 3 (a) Worker Worker Worker T0.0 T1.0 T0.1 T1.1 T0.2 T1.2 (b) Worker Worker Worker T0.0 T1.0 T0.1 T1.1 T0.2 T1.2 (c) Worker Worker Worker T0.2 T1.2 T0.0 T1.0 T0.1 T1.1 (d) Fig. 1: Example of an ORTZI execution flow. (a) Translation of a ORTZI DATAFRAME sort operator into a physical plan and Job compilation in ORTZI: 1 Application program, 2 Logical plan generation: DAG construction, 3 Physical plan generation: task and dependency definition. The subsequent images show different data exchange mechanisms: (b) Local exchange (in-memory), (c) External exchange (object storage), and (d) Hybrid exchange. ORTZI fully manages the execution flow, allowing the programmer to remain agnostic beyond their interaction with the DATAFRAME. We choose cloud object storage as the external storage backend of the current version of the system, as it provides a cost-effective, always-on service [12]. However, ORTZI’s modular design facilitates the potential integration of alternative external storage systems, including serverless cache prototypes [27]–[29] that could potentially offer higher performance. III. DESIGN We design ORTZI architecture to ensure broad compatibility in adaptive execution. The system comprises three main com-
ponents: the application program, a centralized driver, and a pool of distributed executors that host the workers. Figure 2 illustrates the architecture. Executor Executor Scheduler Application ServerlessDataframe().run() Driver Partition manager Executor server Task scheduler Disk Partition manager Workers I/O handler Task register Cloud function Executor Executor server Task scheduler Memory Partition manager Workers I/O handler Task register Executor(s) Cloud virtual instance Kubernetes node Executor(s) Task metadata Object storage Client machine Executor Resource manager Fig. 2: ORTZI architecture diagram. A. Application Program The application program runs on the client machine and is abstracted through the ORTZI DATAFRAME, with which the user directly interacts. It is responsible for job compilation and driver configuration, managing the driver’s lifecycle. The application delegates resource provisioning to the driver, by transferring the physical plan along with resource configuration parameters. B. Driver The driver orchestrates job execution and resource management using dedicated modules, outlined as follows. Resource Manager. Launches cloud functions and manages the lifecycle of cloud VMs and/or pods within an existing Kubernetes [30] cluster. It dynamically provisions new cloud resources based on job-specific computational needs, such as the number of vCPUs and required memory. Currently, ORTZI permits the specification of resource allocation at stage-level granularity. By default, it eagerly scales compute resources to ensure the number of allocated vCPUs aligns with the count of executable, independent tasks at any given time. Future iterations of the framework will incorporate more intelligent resource provisioning strategies. Central Task Scheduler. Instantiates executors on provisioned resources and distributes tasks across the executor pool. Central Partition Manager. Maintains a registry of all system partitions, tracking their host executors. It notifies executors to remove already consumed partitions, preventing storage bloat from orphaned data—especially useful in cloud functions where storage resources are limited. C. Executors Executors are transient instances running on provisioned resources, tasked with executing assigned operations. They communicate solely with the driver and the external storage system. Within each executor, task execution is parallelized across a pool of worker processes. All executors share the following components, with variations only in the technical adaptations required for the specific backend. Executor Server. Establishes a persistent connection with the driver and manages the routing of requests to and from the executor modules. Local Task Scheduler. Requests tasks and assigns them to the available workers. Local Partition Manager. Tracks locally stored partitions and removes them from local storage—either memory or disk—when no longer required. It also spills partitions to external storage upon request. In complex jobs with varying resource needs, a worker often executes multiple tasks within the same stage, where the function code remains consistent. To optimize this, each worker initializes an in-memory task register at startup with the function code for each stage. The global task scheduler only has to send metadata and arguments per task at runtime, reducing both network overhead and deserialization costs. IV. IMPLEMENTATION We build ORTZI in Python, with approximately 8k lines of code. The resource manager in ORTZI is based on Lithops [6] and leverages Docker containers to run executors across different backends. Internally, ORTZI uses Polars [25] for data structures and PyArrow [31] for data (de)serialization. We utilize asyncio coroutines to perform I/O requests within workers, which enables concurrency at low context switching overhead. To enhance I/O performance, we apply file consolidation in exchanges, as described in Riffle [32]. Executors and drivers communicate using gRPC [33]. The driver hosts a gRPC server, where executors notify task completion and partition management requests. Executors not deployed in cloud functions also run their own gRPC server to handle notifications from the task scheduler. However, due to the unaddressable nature of cloud functions, deploying a gRPC server on them is not feasible. Instead, cloud function executors establish a persistent streaming connection with the driver. To synchronize and communicate between components of the executor we rely on Python multiprocessing queues. However, cloud functions do not have access to shared memory, and thus cannot instantiate such data structures. We overcome this limitation by implementing a custom abstraction of multiprocessing queues over pipes for cloud function executors. ORTZI is a fully open-source project and is available at https://github.com/GEizaguirre/ortzi-CLOUD2025.
V. EVALUATION Setup. We use an AWS EC2 [34]m4.10xlarge instance and AWS Lambda [35] functions, each allocated 1,769MB of memory, corresponding to one single vCPU allocation 2. To ensure fair comparisons, we limit the memory of the Docker container running the VM executor to match the total aggregate memory of the cloud functions in the equivalent setup, and assign each worker to a dedicated vCPU. We use AWS S3 [36] for the cloud object storage. We use 5 replicas per configuration and deploy all resources in us-east-1. We center our evaluation on three key insights gained from developing ORTZI. R1. The ORTZI DATAFRAME provides seamless backend compatibility with minimal code modifications. R2. Exchange performance can be significantly optimized based on backend selection and job traits. R3. Our architecture is performant on benchmark analytics jobs. Workloads. Our evaluation includes two widely used data analytics benchmarks. •TeraSort is a common job to assess exchange performance in distributed systems [37]. It consists of two stages with an all-to-all communication pattern. It is datapreserving, meaning that the DATAFRAME size remains the same throughout the execution, imposing significant pressure on the intermediate communication. •TPC-DS (Query 95) is a data analytics benchmark comprising eight stages with complex dependencies, including gather communication patterns and join transformations. We use a 10GB input in our experiments. A. Serverless backends: not a rule of thumb To evaluate the impact of backend selection, we run TeraSort on homogeneous deployments using cloud VMs and cloud functions. We set the number of tasks per stage to match the number of workers. Figure 3presents execution times for 1GB, 2GB, and 5GB TeraSort across the tested configurations. In the VM setup, we deploy a single executor with the specified number of workers. We evaluate performance using three storage backends: disk, object storage, and memory. In cloud functions we launch a separate cloud function for each executor, each function running a single worker. We test cloud functions exclusively with object storage, as their shared memory is limited and disk space is restricted. Cloud functions outperform VMs in strong scaling. The assumption that shared memory always provides the best performance should be approached with caution. At low parallelism levels, co-located workers with in-memory communication perform better. However, as parallelism increases, 2https://docs.aws.amazon.com/lambda/latest/dg/configuration-memory.html distributed workers exhibit superior strong scaling, achieving optimal execution times at much higher levels. This advantage stems from two key factors: (1) the overhead of managing multiple workers on the same machine and (2) resource contention across network, disk, and memory when all workers share a device. This insight is particularly relevant in scenarios with memory constraints and intensive data exchanges, necessitating alternative storage backends. Monoliths remain complex. Executing serverless workloads on a single server still provides valuable insights. While memory exchanges offer the best performance, their scaling behavior closely resembles that of persistent storage, likely due to (de)serialization and concurrency overheads. Disk and object storage perform similarly, especially as input sizes increase. Object storage is preferable for deployments with multiple disaggregated executors, as it enables ubiquitous access without requiring the driver to manage partitions. However, local disk exchanges can help reduce network congestion and improve scalability with larger inputs. Our results show that backend and resource provisioning decisions are complex and significantly impact the execution time of data-intensive jobs (R2). The selection of the optimal backend emerges as a multifaceted problem, presenting an opportunity for future research. ORTZI makes backend portability easy. To showcase the versatility and simplicity of ORTZI DATAFRAMES across cloud continuum resources, we use the same application code for all backend evaluations (R1). All experiments in this section employ the sort primitive from Listing 1. Before invoking sort, the set_scale_out function is called to select either cloud functions or VMs. B. A serverless experience at low cost The ORTZI DATAFRAME incorporates multiple management layers to ensure serverless usability. To evaluate its efficiency and discard excessive overhead, in Table II we compare its execution time in the TeraSort benchmark against Seer [7], a serverless sort operator built directly on Lithops [6]—the same resource provisioning engine used by ORTZI. Framework Execution Time (s) Overhead % Seer [7] 18.45 – ORTZI 19.17 6.78 TABLE II: Seer and ORTZI execution time in a 5GB TeraSort. ORTZI introduces minimal management overhead, incurring only a 6.78% latency increase compared to Seer, which operates directly on cloud functions without generalpurpose usability enhancements. This comparable performance is achieved despite ORTZI’s additional abstractions, such as the executor,worker, and its multiple orchestration layers. C. ORTZI within the State-of-the-Art We compare the execution time of ORTZI against Jolteon [2], an automatic resource provisioning framework for serverless data analytics. We configure Jolteon to meet a 30-second
4 8 12 16 20 24 Number of workers 10 20 30 Execution time (s) EC2 - S3 EC2 - EBS EC2 - Memory Lambda - S3 (a) 1GB input 4 8 12 16 20 24 28 Number of workers 20 40 Execution time (s) (b) 2GB input 12 16 20 24 28 32 36 40 Number of workers 25 50 75 Execution time (s) (c) 5GB input Fig. 3: Execution time of Terasort benchmark across varying input sizes and numbers of parallel workers (tasks). We compare four different serverless configurations: co-located workers in an AWS EC2 instance using shared memory, an EBS (HDD) volume, or AWS S3 for the data exchange, and distributed workers in AWS Lambda instances using AWS S3. SLO while minimizing cost and replicate its exact same resource configuration in ORTZI. Additionally, we compare both systems to EMR Serverless [21], the serverless Spark product in AWS. Figure 4presents the latency and cost results. Overall, ORTZI delivers performance on par with or superior to equivalent prototypes (R3), demonstrating that an open-source, general-purpose framework can compete with commercial solutions and specialized research prototypes. We deliver a performance-competitive architecture. ORTZI achieves latency and cost comparable to Jolteon when using identical resource configurations. This result is expected, as both rely on AWS Lambda and S3. Relative to AWS EMR Serverless, ORTZI improves performance by an impressive 70.59%. This gain may be attributed to slow scaling times of EMR Serverless, as previously reported in [38]. Leveraging hybrid backends translates into performance gains. To illustrate the benefits of cloud continuum architectures (R2), we maintain Jolteon’s specifications (i.e., number of vCPUs and memory) but modify the resource allocation strategy: we offload scan transformations to cloud functions while running the remaining computations on a cloud VM. We refer to this configuration as hybrid in Figure 4. The hybrid deployment cuts latency by 25.57% versus Jolteon’s cloud-function approach. This speedup comes mainly from swapping AWS S3 data exchanges with faster in-memory transfers between outpost stages, especially those with low parallelism, thus reducing I/O time. 20 40 60 80 Execution time (s) 0 1 2 Cost ($0.01) Jolteon EMR Serverless Ortzi (cloud functions) Ortzi (hybrid) Fig. 4: TPC-DS latency in Jolteon [2], EMR Serverless [21] and ORTZI (using a cloud function and a hybrid deployment). Further research could explore the advantages of hybrid, continuum-native deployments by incorporating heterogeneous hardware (like GPUs and FPGAs) or lightweight IoT devices. Given this paper’s focus on demonstrating the design of ORTZI, these specific explorations fall outside its immediate purview but offer significant avenues for future investigation. VI. RELATED WORK Early serverless frameworks [39]–[41] introduced abstractions for offloading local computations to cloud functions, paving the way for the first serverless data analytics operators [5], [6]. However, these solutions offer limited built-in operators, especially those with added complexity such as sort. ORTZI DATAFRAMES provide a comprehensive framework for data analytics, including built-in wide transformations. Fully-fledged serverless data analytics services are available today as commercial solutions from cloud providers [21], [42], [43] and specialized companies [22], [44], [45]. However, these offerings incur additional costs for resource management. In research, efforts instead either focus exclusively on cloud functions [9], [10] or require homogeneous resource pools [?]. Building serverless frameworks over hybrid backends has also been explored in prior work [46] and approached through various proposals [8], [16], [47]. Existing projects, however, either do not deliver out-of-the-box usability or are limited to a few backends. The ORTZI DATAFRAME is the first fully open-source, serverless data analytics abstraction to provide seamless resource management in the cloud continuum. Data exchanges in serverless architectures have been revisited multiple times [4], [7], [12], [13], [48], including inmemory exchanges for co-located tasks [4], [38], [49]. Extensive research has also focused on alternative communication methods for cloud functions [14], [50]. However, existing proposals are not still integrated within broader data analytics frameworks, limiting their flexibility and applicability. Similar limitations apply to domain-specific frameworks [51], [52]. VII. CONCLUSION We present ORTZI, a serverless data analytics framework designed to execute programmatic workloads across heterogeneous backends. ORTZI offers an adaptive architecture supporting cloud functions, cloud virtual instances, and onpremise clusters, utilizing shared memory, disk and object storage for exchanges. We demonstrate that ORTZI effectively navigates compute and storage backends with minimal programming effort.
Despite being open-source and general-purpose, ORTZI achieves competitive performance compared to both commercial and research products. ORTZI is continually evolving and is actively used in ongoing research projects on cloud-edge analytics and smart resource allocation. REFERENCES [1] E. Jonas, J. Schleier-Smith, V. Sreekanti, C.-C. Tsai, A. Khandelwal, Q. Pu, V. Shankar, J. Carreira, K. Krauth, N. Yadwadkar, J. E. Gonzalez, R. A. Popa, I. Stoica, and D. A. Patterson, “Cloud Programming Simplified: A Berkeley View on Serverless Computing,” 2019. [Online]. Available: https://arxiv.org/abs/1902.03383 [2] Z. Zhang, C. Jin, and X. Jin, “Jolteon: Unleashing the Promise of Serverless for Serverless Workflows,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24). Santa Clara, CA: USENIX Association, Apr. 2024, pp. 167–183. [Online]. Available: https://dl.acm.org/doi/10.5555/3691825.3691835 [3] C. Jin, Z. Zhang, X. Xiang, S. Zou, G. Huang, X. Liu, and X. Jin, “Ditto: Efficient Serverless Analytics with Elastic Parallelism,” in Proceedings of the ACM SIGCOMM 2023 Conference, ser. ACM SIGCOMM ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 406–419. [Online]. Available: https://doi.org/10.1145/3603269.3604816 [4] T. Li, Y. Li, W. Zhu, Y. Xu, and J. C. S. Lui, “MinFlow: High-performance and Cost-efficient Data Passing for I/O-intensive Stateful Serverless Analytics,” in 22nd USENIX Conference on File and Storage Technologies, ser. FAST ’24. Santa Clara, CA: USENIX Association, Feb. 2024, pp. 311–327. [Online]. Available: https://dl.acm.org/doi/10.5555/3650697.3650716 [5] E. Jonas, Q. Pu, S. Venkataraman, I. Stoica, and B. Recht, “Occupy the cloud: distributed computing for the 99%,” in Proceedings of the 2017 Symposium on Cloud Computing, ser. SoCC ’17. New York, NY, USA: Association for Computing Machinery, 2017, p. 445–451. [Online]. Available: https://doi.org/10.1145/3127479.3128601 [6] J. Samp´ e, G. Vernik, M. S´ anchez-Artigas, and P. Garc´ ıa-L´ opez, “Serverless Data Analytics in the IBM Cloud,” in Proceedings of the 19th International Middleware Conference Industrial Track, ser. Middleware Industrial Track ’18. New York, NY, USA: Association for Computing Machinery, 2018, p. 1–8. [Online]. Available: https://doi.org/10.1145/3284028.328402 [7] M. S´ anchez-Artigas and G. T. Eizaguirre, “A seer knows best: optimized object storage shuffling for serverless analytics,” in Proceedings of the 23rd ACM/IFIP International Middleware Conference, ser. Middleware ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 148–160. [Online]. Available: https://doi.org/10.1145/3528535. 3565241 [8] H. Bian, T. Sha, and A. Ailamaki, “Using Cloud Functions as Accelerator for Elastic Data Analytics,” Proc. ACM Manag. Data, vol. 1, no. 2, Jun. 2023. [Online]. Available: https://doi.org/10.1145/3589306 [9] B. Carver, J. Zhang, A. Wang, A. Anwar, P. Wu, and Y. Cheng, “Wukong: a scalable and locality-enhanced framework for serverless parallel computing,” in Proceedings of the 11th ACM Symposium on Cloud Computing, ser. SoCC ’20. New York, NY, USA: Association for Computing Machinery, 2020, p. 1–15. [Online]. Available: https://doi.org/10.1145/3419111.3421286 [10] Y. Kim and J. Lin, “Serverless Data Analytics with Flint,” in Proceedings of the IEEE 11th International Conference on Cloud Computing, ser. CLOUD ’18, Los Alamitos, CA, USA, Jul. 2018, pp. 451–455. [Online]. Available: https://doi.org/10.1109/CLOUD.2018.00063 [11] A. M. Nestorov, D. Marr´ on, A. Gutierrez-Torre, C. Wang, C. Misale, A. Youssef, D. Carrera, and J. L. Berral, “Dexter: A PerformanceCost Efficient Resource Allocation Manager for Serverless Data Analytics,” in Proceedings of the 25th International Middleware Conference, ser. Middleware ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 117–130. [Online]. Available: https://doi.org/10.1145/3652892.370075 [12] Q. Pu, S. Venkataraman, and I. Stoica, “Shuffling, Fast and Slow: Scalable Analytics on Serverless Infrastructure,” in 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). Boston, MA: USENIX Association, Feb. 2019, pp. 193–206. [Online]. Available: https://dl.acm.org/doi/10.5555/3323234.3323251 [13] I. M¨ uller, R. Marroqu´ ın, and G. Alonso, “Lambada: Interactive Data Analytics on Cold Data Using Serverless Cloud Infrastructure,” in Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data, ser. SIGMOD ’20. New York, NY, USA: Association for Computing Machinery, 2020, p. 115–130. [Online]. Available: https://doi.org/10.1145/3318464.3389758 [14] A. Mahgoub, K. Shankar, S. Mitra, A. Klimovic, S. Chaterji, and S. Bagchi, “SONIC: Application-aware Data Passing for Chained Serverless Applications,” in 2021 USENIX Annual Technical Conference (USENIX ATC 21). USENIX Association, Jul. 2021, pp. 285–301. [Online]. Available: https://www.usenix.org/conference/atc21/ presentation/mahgoub [15] H. Zhang, Y. Tang, A. Khandelwal, J. Chen, and I. Stoica, “Caerus: NIMBLE Task Scheduling for Serverless Analytics,” in 18th USENIX Symposium on Networked Systems Design and Implementation, ser. NSDI ’21. USENIX Association, Apr. 2021, pp. 653–669. [Online]. Available: https://www.usenix.org/conference/ nsdi21/presentation/zhang-hong [16] A. D. Mohapatra and K. Oh, “Smartpick: Workload Prediction for Serverless-enabled Scalable Data Analytics Systems,” in Proceedings of the 24th International Middleware Conference, ser. Middleware ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 29–42. [Online]. Available: https://doi.org/10.1145/3590140.3592850 [17] C. Zhang, M. Yu, W. Wang, and F. Yan, “MArk: Exploiting Cloud Services for Cost-Effective, SLO-Aware Machine Learning Inference Serving,” in 2019 USENIX Annual Technical Conference, ser. USENIX ATC ’19. Renton, WA: USENIX Association, Jul. 2019, pp. 1049–1062. [Online]. Available: https://dl.acm.org/doi/10.5555/3358807.3358897 [18] M. Zaharia, M. Chowdhury, T. Das, A. Dave, J. Ma, M. McCauley, M. J. Franklin, S. Shenker, and I. Stoica, “Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing,” in 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2012. [Online]. Available: https://dl.acm.org/ doi/10.5555/2228298.2228301 [19] P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan, and I. Stoica, “Ray: A Distributed Framework for Emerging AI Applications,” in 13th USENIX Symposium on Operating Systems Design and Implementation, ser. OSDI ’18. Carlsbad, CA: USENIX Association, Oct. 2018, pp. 561–577. [Online]. Available: https://dl.acm.org/doi/10.5555/3291168. 3291210 [20] Dask Development Team, “Dask: Library for Dynamic Task Scheduling,” https://www.dask.org/, 2025, [Online; accessed 2025-05-30]. [21] “EMR Serverless,” https://aws.amazon.com/es/emr/serverless, [Online; accessed 2025-05-30]. [22] “Coiled,” https://www.coiled.io, [Online; accessed 2025-05-30]. [23] Anyscale, “Scaling Ray Workloads with Anyscale,” https://www. anyscale.com, 2025, accessed: 2025-03-05. [24] Pandas Development Team, “pandas: Python Data Analysis Library,” 2025, [Online; accessed 2025-05-30]. [Online]. Available: https: //pandas.pydata.org/ [25] Polars Developers, “Polars: Lightning-fast DataFrame library for Rust and Python,” https://github.com/pola-rs/polars, 2025, [Online; accessed 2025-02-26]. [26] M. Zaharia, M. Chowdhury, T. Das, A. Dave, J. Ma, M. McCauly, M. J. Franklin, S. Shenker, and I. Stoica, “Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing,” in 9th USENIX Symposium on Networked Systems Design and Implementation, ser. NSDI ’12. San Jose, CA: USENIX Association, Apr. 2012, pp. 15–28. [Online]. Available: https://dl.acm.org/doi/10.5555/2228298. 2228301 [27] F. Romero, G. I. Chaudhry, I. n. Goiri, P. Gopa, P. Batum, N. J. Yadwadkar, R. Fonseca, C. Kozyrakis, and R. Bianchini, “Faa$T: A Transparent Auto-Scaling Cache for Serverless Applications,” in Proceedings of the ACM Symposium on Cloud Computing, ser. SoCC ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 122–137. [Online]. Available: https://doi.org/10.1145/3472883. 3486974 [28] D. Mvondo, M. Bacou, K. Nguetchouang, L. Ngale, S. Pouget, J. Kouam, R. Lachaize, J. Hwang, T. Wood, D. Hagimont, N. De Palma, B. Batchakui, and A. Tchana, “OFC: an opportunistic caching system for FaaS platforms,” in Proceedings of the Sixteenth European Conference on Computer Systems, ser. EuroSys ’21. New
York, NY, USA: Association for Computing Machinery, 2021, p. 228–244. [Online]. Available: https://doi.org/10.1145/3447786.3456239 [29] A. Klimovic, Y. Wang, P. Stuedi, A. Trivedi, J. Pfefferle, and C. Kozyrakis, “Pocket: Elastic Ephemeral Storage for Serverless Analytics,” in 13th USENIX Symposium on Operating Systems Design and Implementation, ser. OSDI ’18. Carlsbad, CA: USENIX Association, Oct. 2018, pp. 427–444. [Online]. Available: https: //dl.acm.org/doi/10.5555/3291168.3291200 [30] Kubernetes Authors, “Kubernetes: Production-Grade Container Orchestration,” https://kubernetes.io/, 2025, accessed: 2025-02-26. [31] A. S. Foundation, “PyArrow: Python library for Apache Arrow,” https: //github.com/apache/arrow, 2025, [Online; accessed 2025-02-26]. [32] H. Zhang, B. Cho, E. Seyfe, A. Ching, and M. J. Freedman, “Riffle: optimized shuffle service for large-scale data analytics,” in Proceedings of the Thirteenth EuroSys Conference, ser. EuroSys ’18. New York, NY, USA: Association for Computing Machinery, 2018. [Online]. Available: https://doi.org/10.1145/3190508.3190534 [33] Google, “gRPC: A high-performance, open-source universal RPC framework,” https://grpc.io/, 2025, [Online; accessed 2025-02-26]. [34] “AWS EC2,” 2025, accessed: 2025-02-28. [Online]. Available: https://aws.amazon.com/ec2/ [35] “AWS Lambda,” 2025, accessed: 2025-02-28. [Online]. Available: https://aws.amazon.com/lambda/ [36] “AWS S3,” 2025, accessed: 2025-02-28. [Online]. Available: https: //aws.amazon.com/s3/ [37] O. O’Malley, “TeraByte Sort on Apache Hadoop,” Yahoo!, Tech. Rep., May 2008. [Online]. Available: https://sortbenchmark.org/ YahooHadoop.pdf [38] D. Barcelona-Pons, A. Arjona, P. Garc´ ıa-L´ opez, E. Molina-Gim´ enez, and S. Klymonchuk, “FaaS Is Not Enough: Serverless Handling of Burst-Parallel Jobs,” 2024. [Online]. Available: https://doi.org/10. 48550/arXiv.2407.14331 [39] S. Fouladi, R. S. Wahby, B. Shacklett, K. V. Balasubramaniam, W. Zeng, R. Bhalerao, A. Sivaraman, G. Porter, and K. Winstein, “Encoding, Fast and Slow: Low-Latency Video Processing Using Thousands of Tiny Threads,” in 14th USENIX Symposium on Networked Systems Design and Implementation, ser. NSDI ’17. Boston, MA: USENIX Association, Mar. 2017, pp. 363–376. [Online]. Available: http://dl.acm.org/doi/10.5555/3154630.3154660 [40] S. Fouladi, F. Romero, D. Iter, Q. Li, S. Chatterjee, C. Kozyrakis, M. Zaharia, and K. Winstein, “From Laptop to Lambda: Outsourcing Everyday Jobs to Thousands of Transient Functional Containers,” in 2019 USENIX Annual Technical Conference, ser. USENIX ATC ’19. Renton, WA: USENIX Association, Jul. 2019, pp. 475–488. [Online]. Available: https://dl.acm.org/doi/10.5555/3358807.3358848 [41] A. Arjona, G. Finol, and P. G. L´ opez, “Transparent serverless execution of Python multiprocessing applications,” Future Generation Computer Systems, vol. 140, pp. 436–449, 2023. [Online]. Available: https://doi.org/10.1016/j.future.2022.10.038 [42] “BigQuery,” https://cloud.google.com/bigquery, [Online; accessed 202505-30]. [43] “Dataproc Serverless,” https://cloud.google.com/dataproc-serverless, [Online; accessed 2025-05-30]. [44] “Modal,” https://www.modal.com, [Online; accessed 2025-05-30]. [45] J. Tagliabue, T. Caraza-Harter, and C. Greco, “Bauplan: Zero-copy, Scale-up FaaS for Data Pipelines,” in Proceedings of the 10th International Workshop on Serverless Computing, ser. WoSC 10. New York, NY, USA: Association for Computing Machinery, 2024, p. 31–36. [Online]. Available: https://doi.org/10.1145/3702634.3702955 [46] P. Garc´ ıa-L´ opez, M. S´ anchez-Artigas, S. Shillaker, P. Pietzuch, D. Breitgand, G. Vernik, P. Sutra, T. Tarrant, and A. J. Ferrer, “ServerMix: Tradeoffs and Challenges of Serverless Data Analytics,” 2019. [Online]. Available: https://arxiv.org/abs/1907.11465 [47] G. T. Eizaguirre, D. Barcelona-Pons, A. Arjona, G. Vernik, P. Garc´ ıaL´ opez, and T. Alexandrov, “Serverful Functions: Leveraging Servers in Complex Serverless Workflows (Industry Track),” in Proceedings of the 25th International Middleware Conference Industrial Track, ser. Middleware Industrial Track ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 15–21. [Online]. Available: https://doi.org/10.1145/3700824.3701095 [48] M. S´ anchez-Artigas, G. T. Eizaguirre, G. Vernik, L. Stuart, and P. Garc´ ıa-L´ opez, “Primula: a Practical Shuffle/Sort Operator for Serverless Computing,” in Proceedings of the 21st International Middleware Conference Industrial Track, ser. Middleware Industrial Track ’20. New York, NY, USA: Association for Computing Machinery, 2020, p. 31–37. [Online]. Available: https://doi.org/10.1145/ 3429357.3430522 [49] S. Qi, L. Monis, Z. Zeng, I.-c. Wang, and K. K. Ramakrishnan, “SPRIGHT: extracting the server from serverless computing! highperformance eBPF-based event-driven, shared-memory processing,” in Proceedings of the ACM SIGCOMM 2022 Conference, ser. SIGCOMM ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 780–794. [Online]. Available: https://doi.org/10.1145/3544216. 3544259 [50] M. Wawrzoniak, I. M¨ uller, G. Alonso, and R. Bruno, “Boxer: Data Analytics on Network-enabled Serverless Platforms,” in Proceedings of the 11th Conference on Innovative Data Systems Research. www.cidrdb.org, 2021. [Online]. Available: https://doi.org/10.3929/ ethz-b-000456492 [51] V. Shankar, K. Krauth, K. Vodrahalli, Q. Pu, B. Recht, I. Stoica, J. Ragan-Kelley, E. Jonas, and S. Venkataraman, “Serverless linear algebra,” in Proceedings of the 11th ACM Symposium on Cloud Computing, ser. SoCC ’20. New York, NY, USA: Association for Computing Machinery, 2020, p. 281–295. [Online]. Available: https://doi.org/10.1145/3419111.3421287 [52] L. Toader, A. Uta, A. Musaafir, and A. Iosup, “Graphless: Toward Serverless Graph Processing,” in Proceedings of the 18th International Symposium on Parallel and Distributed Computing, ser. ISPDC ’19, 2019, pp. 66–73. [Online]. Available: https: //doi.org/10.1109/ISPDC.2019.00012