scieee AI-readable full text Open interactive document viewer

Quantifying Serverless Elasticity: The gumeter Benchmark Suite

Eizaguirre, Germán T.; Molina-Giménez, Enrique; Finol, Gerard; Molina, Carlos; García-López, Pedro

Abstract

This preprint has not undergone peer review. This work has been accepted for publication in the 23rd edition of the International Conference on Service-Oriented Computing, 2025. The Version of Record will be published by Springer Nature. [DOI to be added upon publication].

Full text

Quantifying Serverless Elasticity: The gumeter Benchmark Suite Germán T. Eizaguirre1, Enrique Molina-Giménez1, Gerard Finol1, Carlos Molina1, and Pedro García-López1 Universitat Rovira i Virgili, Tarragona, Spain Abstract. Serverless computing has emerged as a powerful paradigm for distributed workflows, offering fine-grained, low-latency resource provisioning to precisely meet job demands. While workflows often utilize hundreds of concurrent CPUs, existing serverless benchmarking suites frequently concentrate on small-scale parallelism and microservice workloads. Furthermore, these benchmarks typically consider only Functionas-a-Service (FaaS ) backends, overlooking their natural cloud successors, Container-as-a-Service (CaaS ). These limitations have created a significant gap in evaluating a crucial feature of serverless platforms: their ability to accurately handle sudden changes in resource allocation, also known as elasticity. To address this, we introduce gumeter, a new benchmarking suite specifically designed to evaluate the elasticity of serverless platforms (both FaaS and CaaS ) for highly parallel distributed workflows. gumeter facilitates a thorough assessment of an underlying platform using a "fire-and-forget" execution model and minimal user intervention. It leverages a set of comprehensive pipelines that sample various typical scaling behaviors, providing an in-depth analysis of elasticity, execution time, cost, and efficiency. We apply gumeter to evaluate the elasticity of three popular serverless platforms: AWS Lambda, Google Cloud Run, and IBM Code Engine. Our results reveal significant differences in elasticity across these platforms, showing that FaaS offerings still outperform CaaS in elasticity by a factor of up to 6.5x. However, despite a lower performance, CaaS can be up to 64.5% less expensive than FaaS in specific scenarios, thereby unveiling an interesting optimization space for cloud workflows. Keywords: Serverless computing ·Function-as-a-Service (FaaS) ·Containeras-a-Service (CaaS) ·Benchmarking 1 Introduction Serverless computing has seen rapid adoption in cloud computing, emerging as a solution to the operational complexities of traditional cloud models, like cloud virtual machines. By offering managed resource provisioning and setup abstractions, serverless platforms eliminate the need for developers to manage infrastructure or runtime environments. Instead, they can focus solely on application logic while the cloud provider handles server allocation, scaling, and execution. 2 Eizaguirre et al. While Function-as-a-Service (FaaS) platforms such as AWS Lambda [3] were the first to popularize serverless computing, the paradigm now encompasses a broad ecosystem. This includes scalable object storage (e.g., AWS S3 [2]), data analytics services (e.g., GCP BigQuery [8]), and other managed solutions. Despite their diversity, all serverless platforms adhere to key defining features [18]: Managed resource allocation. Serverless abstracts infrastructure management, handling not just individual resources, but dynamically orchestrating distributed resource pools behind the scenes. This is especially useful for workloads with high and variable parallelism, where the number of required resources scales in and out frequently. Pay-as-you-go billing. Users are charged only for the actual execution time of their functions, with no cost incurred during idle periods. This model eliminates waste and optimizes costs for applications with fluctuating or unpredictable workloads. Decoupled compute and storage. Serverless separates compute and storage, allowing each to scale independently. This improves efficiency and cost-effectiveness by tailoring resource allocation to an application’s specific needs. Serverless Computing for Distributed Workflows The properties of serverless makes it a particularly well-suited paradigm for workloads with dynamic parallelism demands. Such is the case for distributed workflows, where workloads are often intermittent and bursty and resource requirements change dynamically. This translates into highly parallelism needs both within and between jobs. Workflows specially benefit from services providing computational resources in a serverless manner. We call such services Serverless Computing Platforms, as they enable instant resource deployment precisely when needed and strictly for the required duration. This allows matching the required parallelism almost instantly, in a proactive way (as opposite to reactive scaling of tradicional cluster services, where resources are (de)provisioned in response to real time workload metrics). Notably, FaaS platforms can provision thousands of CPU instances within seconds, making them exceptionally appealing for embarrassingly parallel workloads [17, 25]. However, FaaS are not the only cloud services that provide computational resources in a serverless manner. To this day, alternative service types, such as Container-as-a-Service (CaaS) (like IBM Code Engine [16]), provide highly similar APIs to those of FaaS, but with more control over the provisioned resources and essential, underlying architectural shifts. These services, however, are often overlooked in discussions of serverless computing. What fits in A Serverless Computing Platform? To delineate the scope of this work, we first define what can be considered a serverless computing platform. This definition remains ambiguous in practice, as Quantifying Serverless Elasticity: The gumeter Benchmark Suite 3 modern interpretations of serverless extend beyond original serverless services to include various managed components. In fact, cloud providers have broadened the serverless designation to encompass managed services that maintain core serverless attributes. Notable examples include managed cluster services like AWS EMR Serverless [1] and GCP Dataproc Serverless [9], which abstract infrastructure management while offering turnkey data analytics and AI capabilities. ♦We consider a serverless computing platform that with (1) rapid computational resource provisioning and (2) effectively unbounded horizontal scalability. An analysis of commercial and research serverless workflow processing platforms [17, 23, 6, 29] reveals consistent architectural patterns aligned with serverless principles: (1) a uniform, scalable serverless computing platform for data processing, (2) cloud storage for data exchange and fault tolerance, and (3) auxiliary orchestration services. Figure 1 illustrates this canonical architecture. Programmer Serverless framework Serverless computing platform 1 2 ... 0 <workflow> Cloud resources ∞ Serverless storage backend Auxiliary services (e.g. runtime repositories, logging services) Fig. 1: Segregating Components in Serverless Workflow Processing. A serverless data analytics platform provides a programming framework while abstracting the underlying execution flow. Behind the scenes, it features decoupled storage and compute backends, along with auxiliary services (e.g., for logging or message passing). Serverless compute backends offer unlimited horizontal scaling, adaptive resource allocation, and a pay-as-you-go billing model. Storage backends may include in-memory storage, object storage, or managed database services. Defining a serverless computing platform is one thing, but reaching a consensus on the "best" platform for distributed workflow processing remains an elusive goal. In this work, we’re focusing on precisely that challenge. 4 Eizaguirre et al. Do We Need Yet Another Serverless Benchmark Suite? Properly evaluating serverless computing platforms requires capturing specific performance details while accounting their specific limitations. The literature on serverless benchmarking is broad [5, 22, 24, 30, 19, 12, 27], including function workflows with modest resource requirements [28, 13, 20]. However, to the best of our knowledge there is no specific benchmarking suite specifically designed for highly parallel serverless workflows with cross-platform compatibility. Serverless workflows exhibit distinctive properties that complicate their evaluation. They can experience parallelism fluctuations spanning two to three orders of magnitude within job [31, 26, 21], creating dynamic resource requirements that existing benchmarks cannot adequately assess. However, current solutions remain constrained by their focus on either modest parallelism scenarios (typically limited to tens of functions [28, 20]) or isolated function evaluation. To bridge this gap, we introduce gumeter, a comprehensive benchmarking suite specifically designed to assess serverless computing platforms for distributed workflows. gumeter enables holistic evaluation through representative workloads, incorporates elasticity-aware performance metrics and integrates cloud-specific cost modeling while maintaining cross-platform compatibility. Implemented as an open-source solution, gumeter supports both immediate benchmarking needs and future extensions with both further platforms and user-defined workflows. The key contributions of this paper are: –A comprehensive benchmarking suite for serverless data analytics to evaluate performance, cost, and elasticity of multiple architectures, enabling direct comparison of serverless computing platforms. –A methodology to assess the critical triad of elasticity, performance, and cost in serverless workflow execution, addressing a key gap in current evaluation approaches. –A detailed comparison of relevant computing platforms, including AWS Lambda, GCP Cloud Run and IBM Code Engine, disclosing important and practical trade-offs. 2 Design 2.1 Measuring Elasticity in Workflow Execution Measuring cloud elasticity in data workflows remains a challenge, as no consensus exists on a standard metric despite its critical importance. While resource elasticity is not exclusive to the cloud, it is inherently associated with it: since cloud services provide on-demand resources, it is essential to assess how fast these resources adapt to an application’s needs. However, a universally accepted elasticity metric remains elusive—a challenge recognized for over a decade [14]. Quantifying Serverless Elasticity: The gumeter Benchmark Suite 5 Elasticity is defined by three key features [4]: (1) resource (de)provisioning latency, (2) the maximum number of resources that can be acquired concurrently, and (3) billing granularity. However, our definition of a serverless backend assumes virtually unlimited resources and a pay-per-use billing model, inherently satisfying (2) and (3). Thus, we focus on measuring how quickly a backend adapts computing resources to match workload demands. Isolating elasticity as a function of resource (de)provisioning speed simplifies the problem and enables the use of a more intuitive metric. Building on previous proposals [7], we aim to define an indicator analogous to the physics concept of elasticity. In materials science, Young’s modulus—hereafter referred to as stiffness—quantifies the maximum stretch a material undergoes under a given tension. Denoting ϵas the applied tension, and Las the material’s length, we define stiffness Sas: S=σ △L/L (1) To adapt this concept to cloud systems, we represent the applied tension as the required resources, Cr, and the length as the resources provided by the system, Cp. Thus, the system’s stiffness at a given point in the application’s execution is defined as: S=Cr max(Cp,1) (2) 012345 Time 0 500 1000 Resources Required Provided Stiffness Fig. 2: Measuring stiffness in serverless workflow execution. Assuming resources can scale out indefinitely, we quantify the ratio between resources required by the workflow and the actual resources provided by the system. A static measure of stiffness alone is insufficient to fully assess a data analytics pipeline. As discussed previously, resource demands vary across pipeline stages, each with different requirements. Therefore, in real-world scenarios, evaluating 6 Eizaguirre et al. the stiffness of a serverless backend requires measuring its resource provisioning accuracy throughout the entire execution. We calculate it as the area between the required and provisioned resource curves, illustrated in Figure 2. Let the absolute resource mismatch over the interval [0, T]be: S=ZT 0 (Cr(t)−Cp(t)) dt (3) This represents the cumulative amount of under-provisioned resources. To normalize this area, we define the maximum possible mismatch as: Smax =ZT 0 Cr(t)dt (4) which corresponds to the case where no resources are provisioned at all, i.e., Cp(t)=0for all t. Then, we define the normalized elasticity as: Earea = 1 −S Smax (5) This formulation yields values in the range [0,1], where Earea = 1 indicates a perfect match between provisioned and required resources, and Earea = 0 indicates complete inefficiency (no resources were provisioned at any time). E, denoting elasticity coefficient, represents a useful measure of the elasticity of different platforms hosting the same workflow. However, since Eis highly dependent on the specific resource demands of a given workflow, it is not suitable for cross-workflow comparisons. 2.2 A Representative Yet Practical Workflow Set Table 1 presents a detailed list of all workflows included in gumeter, outlining their functionality, execution plan, and observable insights. These workflows are designed to encompass a variety of typical scaling behaviors, from steep scaleouts (e.g., FLOPS, Monte Carlo Stock Prediction) and drastic scale-ins (e.g., Monte Carlo Pi Estimation, Monte Carlo Stock Prediction) to scenarios with changing parallelism within the workflow itself (e.g., Mandelbrot, TeraSort). For practical application, we’ve aligned the benchmark parallelisms with the standard service quotas of AWS, GCP, and IBM Cloud, assuming one vCPU per serverless instance. Quantifying Serverless Elasticity: The gumeter Benchmark Suite 7 Table 1: Overview of gumeter workflows. Workflow Description Stages Key Insights FLOPS Each function independently executes floatingpoint operations over two N-dimensional matrices. Single stage with 200 functions. (1) Measures the computational power of the platform. (2) Reveals steep scaleout capabilities. Monte Carlo Stock Prediction MapReduce implementation of a Monte Carlo algorithm to predict stock variation trends. Map stage: 10 functions; Reduce stage: 1 function. Reveals moderate scale-in capabilities. Monte Carlo Pi Estimation MapReduce implementation of a Monte Carlo algorithm to approximate the value of π. Map stage: 100 functions; Reduce stage: 1 function. Reveals steep scalein capabilities. TeraSort Distributed sort over a 5GB TeraSort dataset. First stage: 50 functions; Second stage: 100 functions. (1) Evaluates scalability for longlasting functions. (2) Reveals moderate scale-in capabilities. Mandelbrot Runs several stages, each calculating the Mandelbrot set on a region of the linear space. Seven consecutive stages with [4, 16, 36, 64, 100, 144, 196] functions. Reveals iterative, progressive scale-out capabilities. 3 Implementation We use lithops [25], an open source backend listed in the third-party integrations of Code Engine1as the central building block of gumeter.lithops is a serverless framework that allows users to run Python functions on serverless platforms. It provides a unified interface for executing functions across different cloud providers, making it suitable for benchmarking purposes. lithops supports various backends, including AWS Lambda and IBM Code Engine, allowing us to evaluate the elasticity of these platforms using the same codebase. For benchmarking purposes, we extend lithops with an additional GCP Cloud Run backend and prove its extensibility to other serverless platforms. 1https://cloud.ibm.com/docs/codeengine?topic=codeengine-supported-integrations 8 Eizaguirre et al. 4 Evaluation 4.1 Benchmarked Platforms and Setup To demonstrate gumeter’s utility, we compared the performance of three representative serverless platforms from different cloud providers. The setups employed are detailed in Table 2. We selected a pure FaaS platform (AWS Lambda), a pure CaaS platform (IBM Code Engine), and a hybrid alternative (Google Cloud Run). We extracted execution time, cost, and elasticity metrics from each to provide a comparative analysis. Table 2: Overview of evaluated serverless platform setups. Platform Cloud Provider Type Storage Lambda [3] AWS FaaS AWS S3 [2] Cloud Run [11] GCP CaaS /FaaS GCP Cloud Storage [10] Code Engine [16] IBM Cloud CaaS IBM COS [15] All experiments ran in the us-east1 region or its equivalent. We pre-warmed instances for all experiments with three replicas of 200 functions to ensure a sufficiently warm execution environment. Each function was provisioned with 1 vCPU. In AWS Lambda, CPU allocation is directly linked to memory, with 1 vCPU corresponding to 1769MB of RAM. For Code Engine and Cloud Run, which offer more flexible CPU-memory configurations, we allocated 2048MB of memory per function to closely match Lambda’s configuration. 4.2 Elasticity We evaluate the elasticity of each platform when handling different workloads. Figure 3 illustrates the provisioned CPUs (equivalent to the number of functions) over time, offering a visual representation of each workload’s execution timeline. FaaS backends demonstrate faster and more stable scaling compared to CaaS. At low parallelism (Figure 3a), the primary difference appears to be in scaling latency. However, as parallelism increases (Figure 3c), Code Engine exhibits inconsistent scale-out patterns rather than a monotonic increase in provisioned CPUs. This behavior is particularly problematic with frequent changes in parallelism, as shown in Figure 3d. Code Engine’s scaling issues may stem from its architecture. Unlike FaaS platforms, Code Engine instances are provisioned within a Kubernetes cluster and require a membership protocol for provisioning. This could interfere with the scaling latency of Code Engine resources. Lambda and Cloud Run show similar performance, with Lambda demonstrating slightly lower scale-out latency. Interestingly, task execution time in Quantifying Serverless Elasticity: The gumeter Benchmark Suite 9 TeraSort appears significantly slower in Cloud Run. This could be attributed to the I/O performance of GCP Cloud Storage, given that a distributed sort involves a data-intensive all-to-all exchange between its stages. 0 5 10 15 20 25 30 Execution Time (s) 0 10 # CPUs AWS Lambda GCP Cloud Run IBM Code Engine (a) Monte Carlo Stock Prediction. 0 5 10 15 20 25 Execution Time (s) 0 100 # CPUs (b) Montecarlo Pi Estimation. 0 20 40 60 80 100 Execution Time (s) 0 100 # CPUs (c) TeraSort. 0 20 40 60 80 100 Execution Time (s) 0 50 # CPUs (d) Mandelbrot. Fig. 3: CPU provisioning timelines of each workflow, on the evaluated platforms. Figure 3 serves as a validation for our proposed elasticity metric, which we depict in Figure 4. As commented before, the elasticity of CaaS does not match that of FaaS, specially as parallelism increases and becomes more variable.