scieee AI-readable full text Open interactive document viewer

Automated capacity analysis of limitation-aware microservices architectures

Fresno Aranda, Rafael

Abstract

La aparición de las arquitecturas de microservicios (MSAs) ha supuesto un cambio importante en la forma en la que se desarrollan sistemas y aplicaciones. Se trata de una evolución del paradigma de las arquitecturas orientadas a servicios (SOAs), y cuenta con diversas ventajas frente a las arquitecturas monolíticas tradicionales. Por ejemplo, podemos destacar la agilidad y rapidez de despliegue, escalabilidad, rendimiento, mejor mantenimiento o flexibilidad. Esto hace que las MSAs hayan ganado popularidad en los últimos años, habiendo sido adoptadas por muchas empresas como Netflix, Amazon o Spotify. En el contexto de las MSAs, es habitual el uso de interfaces de programación de aplicación (APIs), que sirven como mecanismo de comunicación entre los servicios de la arquitectura. Entre los diversos tipos de APIs, las más populares son las APIs RESTful, que se basan en el uso de peticiones HTTP para controlar el estado de datos o servicios, denominados recursos. Esta aproximación permite una mayor descentralización de los servicios, lo que coincide con la esencia de las MSAs. Ante este escenario, muchas empresas han encontrado en las APIs la posibilidad de vender sus datos y funcionalidades para que puedan ser usados por otras empresas. Este nuevo paradigma se denomina API economy, y se define como el conjunto de modelos y prácticas de negocio que se centran en torno al uso de APIs públicas. Esto abre la puerta a nuevas formas de innovación, colaboración y generación de beneficios. Una de las bases de la API economy es la definición de planes de precios. Las empresas que quieran usar una API deberán elegir entre diversos planes. Estos planes establecen una serie de limitaciones de uso a cambio de un precio específico, habitualmente una suscripción periódica. Las limitaciones más habituales son las restricciones del número de peticiones que se pueden enviar en un determinado periodo de tiempo. No obstante, existen otras muchas limitaciones que dependen del dominio de cada API. Con la aparición de la API economy, se hace necesario analizar cómo influyen las limitaciones y los precios de los planes elegidos en la capacidad de la MSA. Esto es, averiguar la carga de trabajo que la MSA puede soportar sin exceder las limitaciones de las APIs consumidas, así como mantener los costes dentro del presupuesto de las empresas. Además, no es poco frecuente que las empresas ofrezcan sus propias APIs con sus propios planes a sus clientes. En este escenario, las empresas actúan como prosumidores, porque consumen APIs externas a la vez que proveen sus propias APIs. La confrontación de estos dos roles genera un problema de impedancia, porque las empresas necesitan gestionar y equilibrar cuidadosamente los costes de las APIs consumidas, mientras que mantienen condiciones óptimas para sus clientes. Hasta donde nosotros sabemos, este análisis no se ha hecho nunca en la literatura existente, lo que abre una línea de investigación interesante a la vez que útil. Además, el análisis manual de una MSA con APIs externas es tedioso y propenso a errores, por lo que es conveniente tener algún sistema de análisis automático. Por lo tanto, el objetivo principal de esta tesis es el desarrollo de modelos y técnicas para asistir en el análisis de la capacidad de arquitecturas de microservicios que consumen APIs externas con limitaciones. Estas MSAs las hemos denominado arquitecturas de microservicios conscientes de limitaciones (LAMAs). Los resultados principales han sido: (I) Definición de un modelo para la descripción de LAMAs y los planes de precios. Esto comprende: (i) analizar un conjunto representativo de APIs reales para conocer su estructura y elementos habituales; (ii) extensión de un modelo existente con nuevos elementos; (iii) proponer un nuevo modelo para la descripción de la topología de una LAMA; (iv) definir un catálogo de operaciones para resolver cuestiones sobre planes de precios y sobre la capacidad de una LAMA. (II) Implementación de un ecosistema de herramientas para dar soporte al análisis automático de la capacidad de una LAMA. Esto incluye: (i) desarrollar un conjunto de herramientas para el análisis de la validez de un pricing y la capacidad de una LAMA; (ii) implementar un catálogo base de operaciones de análisis; (iii) validar las herramientas con escenarios reales y sintéticos. En esta tesis, presentamos un conjunto de modelos y herramientas que comprende lo siguiente: (i) un modelo para describir planes de precios de APIs RESTful, así como una serialización alineada con el estándar OpenAPI; y (ii) diversas herramientas para validar un pricing, calcular automáticamente la capacidad de una LAMA y dar respuesta a operaciones de análisis. Por lo tanto, los resultados de esta tesis están destinados a ayudar a las empresas que usen LAMAs en la toma de decisiones en función de sus necesidades.

Full text

Automated Capacity Analysis of Limitation-Aware Microservices Architectures Rafael Fresno Aranda Doctoral Thesis Advised by Dr. Pablo Fernández Montes and Dr. Antonio Ruiz Cortés Universidad de Sevilla February 2024 First published in February 2024 by Rafael Fresno Aranda Copyright ©MMXXIV [email protected] This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. Support: Doctoral Thesis supported by the FPU scholarship program, granted by the Spanish Ministry of Science, Innovation and Universities (FPU19/00666). Dedicado a mis padres, Isabel y Rafael, por haber estado ahí siempre que lo he necesitado. Gracias por todo. Abstract i The appearance of microservice architectures (MSAs) has been an important change in the way that systems and applications are developed. They are an evolution of the service-oriented architecture (SOA) paradigm, and have multiple advantages over traditional monolithic architectures; as an example, we can highlight the agility and speed of deployment, scalability, performance, improved maintenance or flexibility amongst others. This made MSAs more and more popular over the last years, and they have been adopted by many companies such as Netflix, Amazon or Spotify. In the context of MSAs, it is common to use application programming interfaces (APIs), that serve as a communication mechanism between the services of the architecture. Among the various types of APIs, the most popular are RESTful APIs, which are based on the usage of HTTP requests to manage the state of data or services, known as resources. This approach contributes to the decentralization of services, and this is aligned with the essence of MSAs. In this scenario, many companies have found in APIs the possibility to sell their data and functionality so that they can be used by other businesses. This new paradigm is known as API economy, and is defined as the set of business models and practices that revolve around the usage of public APIs. This paves the way for new ways of innovation, collaboration and revenue generation. One of the elements of the API economy is the definition of pricings and plans. Businesses who wish to use an API must choose between different plans. These plans define a set of usage limitations for a specific price, which is usually a periodic subscription. The most common limitations are the restrictions to the number of requests that can be sent over a certain period of time. Nonetheless, there are many other limitations that depend on the domain of each API. With the appearance of the API economy, it becomes necessary to analyze how the limitations and the price of the chosen plans have an impact on the capacity of the MSA. That is, calculating the workload that the MSA is able to handle without exceeding the limitations of the consumed APIs, as well as keeping costs within the businesses’ budget. Furthermore, it is not uncommon for businesses to offer their own APIs with their own plans to their customers. In this scenario, businesses act as prosumers, because they consume external APIs but also provide their own APIs. The confrontation of these two roles creates an impedance problem, where businesses need to carefully manage and balance the costs of the consumed APIs, while simultaneously keeping optimal offerings for their customers. To the best of our knowledge, this analysis has never been done in the literature, and this fact opens a line of research that is interesting and also useful. Furthermore, the manual analysis of an MSA with external APIs is tedious and error-prone, so it is convenient for businesses to have some automated analysis systems. Given all of this, the main goal of this thesis is the development of models and techniques to assist in the capacity analysis of microservice architectures that consume external APIs with limitations. We coin this type of MSAs as limitation-aware microservice architectures (or LAMAs). The main results have been the following: (I) Definition of a model for the description of LAMAs and pricings. This comprises: (i) analyzing a representative set of real-world APIs to know their structure and common elements; (ii) extending an existing model with new elements; (iii) proposing a new model for the description of the topology of a LAMA; (iv) defining a catalogue of operations to solve questions about pricings and the capacity of a LAMA. (II) Implementation of an ecosystem of tools to support the automated analysis of the capacity of a LAMA. This includes: (i) developing a set of tools for the automated analysis of the validity of a pricing and the capacity of a LAMA; (ii) implementing a base catalogue of analysis operations; (iii) validating the tools with real-world and synthetic scenarios. In this thesis, we present a set of models and tools that comprises the following: (i) a model to describe RESTful API pricings, and a serialization that is in line with the OpenAPI standard; and (ii) various tools to validate a pricing, automatically calculate the capacity of a LAMA and provide answers to analysis operations. Therefore, the results of this thesis are oriented to help businesses using LAMAs in the process of making decissions according to their needs. ii Resumen iii La aparición de las arquitecturas de microservicios (MSAs) ha supuesto un cambio importante en la forma en la que se desarrollan sistemas y aplicaciones. Se trata de una evolución del paradigma de las arquitecturas orientadas a servicios (SOAs), y cuenta con diversas ventajas frente a las arquitecturas monolíticas tradicionales. Por ejemplo, podemos destacar la agilidad y rapidez de despliegue, escalabilidad, rendimiento, mejor mantenimiento o flexibilidad. Esto hace que las MSAs hayan ganado popularidad en los últimos años, habiendo sido adoptadas por muchas empresas como Netflix, Amazon o Spotify. En el contexto de las MSAs, es habitual el uso de interfaces de programación de aplicación (APIs), que sirven como mecanismo de comunicación entre los servicios de la arquitectura. Entre los diversos tipos de APIs, las más populares son las APIs RESTful, que se basan en el uso de peticiones HTTP para controlar el estado de datos o servicios, denominados recursos. Esta aproximación permite una mayor descentralización de los servicios, lo que coincide con la esencia de las MSAs. Ante este escenario, muchas empresas han encontrado en las APIs la posibilidad de vender sus datos y funcionalidades para que puedan ser usados por otras empresas. Este nuevo paradigma se denomina API economy, y se define como el conjunto de modelos y prácticas de negocio que se centran en torno al uso de APIs públicas. Esto abre la puerta a nuevas formas de innovación, colaboración y generación de beneficios. Una de las bases de la API economy es la definición de planes de precios. Las empresas que quieran usar una API deberán elegir entre diversos planes. Estos planes establecen una serie de limitaciones de uso a cambio de un precio específico, habitualmente una suscripción periódica. Las limitaciones más habituales son las restricciones del número de peticiones que se pueden enviar en un determinado periodo de tiempo. No obstante, existen otras muchas limitaciones que dependen del dominio de cada API. Con la aparición de la API economy, se hace necesario analizar cómo influyen las limitaciones y los precios de los planes elegidos en la capacidad de la MSA. Esto es, averiguar la carga de trabajo que la MSA puede soportar sin exceder las limitaciones de CONTENTS x List of Figures xi 1.1 SendGridAPIplans .............................. 7 1.2 Bluejayoverview ................................ 10 1.3 Example of LAMA with three internal services and two external APIs . . . 11 2.1 Microservice architecture example . . . . . . . . . . . . . . . . . . . . . . . 23 2.2 OpenAPI Initiative members . . . . . . . . . . . . . . . . . . . . . . . . . . 29 2.3 SwaggerUIexample .............................. 30 2.4 Architecture of an AWS EC2 instance deployed on an Amazon Virtual PrivateNetwork................................. 33 2.5 Traditional vs virtualized vs containerized deployments . . . . . . . . . . . 35 2.6 OpenTelemetrydiagram ............................ 36 2.7 Example of Prometheus data visualized in a Grafana dashboard . . . . . . 37 3.1 Overview of the API economy applied to a company that offers an API that can be used in any device . . . . . . . . . . . . . . . . . . . . . . . . . 40 3.2 Sliding (rates) vs fixed (quotas) windows . . . . . . . . . . . . . . . . . . . 43 3.3 RapidAPImainpage.............................. 46 4.1 ITILoverview.................................. 50 5.1 Pricing4APIs model for API pricings . . . . . . . . . . . . . . . . . . . . . 67 5.2 Examples of different consumption scenarios for the same ThresholdedLimitation .................................. 71 LIST OF FIGURES 5.3 Extended RapidAPI SendGrid pricing. Each plan has a price, rate, quota, overage cost and, in some cases, a feature . . . . . . . . . . . . . . . . . . . 77 9.1 Diagram of the database filtering process, starting from the Gamez-Diaz and Neumann datasets and the RapidAPI most popular list . . . . . . . . 115 9.2 Tool running a syntax check . . . . . . . . . . . . . . . . . . . . . . . . . . 119 9.3 Tool running the validity operation with errors . . . . . . . . . . . . . . . . 119 9.4 Simple UI for the sla4oai-analyzer API....................120 10.1 The LAMA of Bluejay, including its internal services and external APIs. Each API is depicted along its plans and limitations . . . . . . . . . . . . . 126 11.1 Simplified LAMA for the monitoring experiment . . . . . . . . . . . . . . . 129 11.2 Inferred topology of the LAMA under analysis . . . . . . . . . . . . . . . . 131 xii INTRODUCTION Part I 1 Introduction 3 This chapter introduces the research context and outlines the goals and contributions of this thesis. Specifically, Section §1.1 describes the concepts of the research context which frame the scope of the work. Next, Section §1.2 exposes the main problem addressed. Section §1.3 describes the main goals of this thesis as well as the research questions that support our research. In Section §1.4 we present a summary of our contributions and a list of publications, research stays and awards. Next, in Section §1.5 we provide details about the research methodology followed in this thesis. Finally, Section §1.6 describes how the contents of this thesis is organized. CHAPTER 1. INTRODUCTION 1.1 Research Context Microservice Architectures Are Blooming As organizations continue to navigate the complexities of digital transformation, the principles of service orientation stand out as key enablers for building future-proof, adaptable and efficient software systems. For more than a decade, this paradigm change driven by the need for more agile, scalable and resilient systems, has profoundly impacted how software is designed, developed and deployed. Service orientation emphasizes the provision of software as a suite of independently deployable services, each performing a distinct function. This approach contrasts with the traditional monolithic architecture where all functionalities are tightly integrated into a single application. Originally, service-oriented architectures (SOAs) laid the foundation for this evolution. SOA’s primary goal was to break down monolithic applications into interoperable services, promoting flexibility and reuse [1,2]. However, during the last years, microservice architectures (MSAs), a finer-grained evolution of SOAs, take these principles to the next level by focusing on small, self-contained services that can be developed, deployed and scaled independently. In fact, the current trend in this direction was facilitated by advancements in cloud computing, containerization, and DevOps practices, which provided the necessary tools and methodologies to manage these distributed services efficiently. There are clear evidences on the benefits of MSAs. In the context of this thesis, we can highlight the following four: (i) agility and speed of deployment, as they allow teams to respond quickly to market changes and customer demands, giving them a competitive advantage [3,4]; (ii) scalability and performance optimization, as each service can be scaled independently based on its specific resource requirements and efficiently utilizing underlying infrastructure resources [5,6] (iii) resilience and fault isolation, as failure of one service does not necessarily bring down the entire system and it is easier to identify, diagnose, and rectify issues [3]; or (iv) improved maintenance and update cycles, as smaller codebases and independent services simplify maintenance and updates and teams can implement changes to a single service without impacting others, reducing the risk associated with deployments [7]. As a consequence of those benefits, various industries, from finance to healthcare, have embraced service-oriented architectures, especially microservices, for their digital transformation initiatives. Companies like Netflix, Amazon, and Spotify are notable examples of successful implementations of microservice architectures. These organizations have 4 1.1. RESEARCH CONTEXT demonstrated how service orientation can support massive scale, continuous innovation and high availability. Moreover, service orientation plays a pivotal role in inter-organization integration, where different organizations need to link their systems and processes. By using serviceoriented architectures, companies can expose certain functionalities as services, which can be consumed by other organizations. This approach simplifies integration and fosters collaboration between businesses. Specifically, in Business to Business (B2B) scenarios, where organizations interact with each other, service orientation allows for the creation of seamless, automated workflows. For instance, a supplier’s inventory system can be integrated directly with a retailer’s ordering system, enabling real-time inventory updates and automated order placement. Also within the context of large organizations, service orientation aids in breaking down silos between different departments or business units. By decomposing complex systems into microservices, different teams can work on individual services with clear interfaces, enhancing collaboration and reducing dependencies [8]. This modularity is particularly beneficial in very large organizations (such as public administrations), where different departments might have varied technology stacks and development practices. Microservices allow these diverse units to integrate their systems more efficiently, leading to more cohesive and unified intra-organizational processes. The Consolidation of an API Economy The concept of APIs (Application Programming Interfaces) is at the heart of MSAs, serving as the fundamental mechanism for interaction and communication between the discrete services that define this architectural style. In a microservice architecture, each service is developed, deployed and operated independently, catering to a specific business function or process. APIs facilitate these services to communicate with each other and with the outside world, acting as well-defined contracts that specify how software components should interact. This design principle allows for a loosely coupled system where services can be updated, replaced or scaled without impacting the overall application. Moreover, APIs enable the MSA to be language-agnostic, allowing different services to be written in the languages best suited for their requirements. Through this central role, APIs not only empower the modularity and flexibility of microservices but also enhance their scalability, resilience and maintainability, which are key advantages that make MSAs highly effective for complex, evolving software applications. Building upon the foundational role of APIs in microservice architectures, REST 5 CHAPTER 1. INTRODUCTION (REpresentational State Transfer) has emerged as the de-facto standard for designing web-based APIs, owing to its simplicity, scalability, and statelessness. RESTful APIs use HTTP requests to manage the state of resources (such as data or services) on the web, making them an ideal fit for the decentralized and distributed nature of microservices. This architectural style’s alignment with the principles of the web has fostered an open ecosystem of supporting tools and technologies, ranging from API gateways for managing API requests and responses, to comprehensive service discovery mechanisms that ensure dynamic routing and load balancing. The ubiquity of REST has also led to the widespread availability of developer resources, frameworks and best practices, significantly lowering the barrier to entry for implementing MSAs. Consequently, RESTful APIs not only facilitate the internal workings of microservices but also enable these architectures to seamlessly integrate with external systems and services, further extending their reach and utility in modern software applications. In this context, the term API economy began to gain traction in the early 2010s as companies like Google, Amazon, and Salesforce started to demonstrate the power of APIs in expanding business operations and creating new revenue streams [9]. These companies leveraged APIs not just as a technical interface, but as a strategic asset, enabling them to tap into networks of developers and other businesses to grow their platforms exponentially. In general, during the last decade, The API economy has transformed how businesses operate, fostering new levels of innovation, collaboration, and revenue generation. Specifically the API economy paradigm describes a set of business models and practices centered around the use of APIs to enable software applications to communicate with each other and leverage third-party services efficiently. Furthermore, the API economy boosted the idea of a marketplace that emerged when companies exposed their internal services and data to external parties through web APIs. Consequently, thanks to the standardization of REST, a consolidated API market has been established composed by a number of services ready to be used with a flexible pricing model. As an example, in Fig. §1.1 [10] we can see the real pricing of the SendGrid API that corresponds to a widely used email service. Specifically, this API offers a suite of features aimed at facilitating robust and scalable email services for applications and businesses; moreover, it is designed to allow businesses to seamlessly integrate email sending capabilities into their applications, enabling automated transactional and marketing email workflows. Concerning its pricing, we can see four distinct tiers adapted to a variety of customer needs: (i) An entry-level tier, labeled Free, that offers a no-cost option with a daily cap of 100 emails that could be an attractive choice for individuals or small busi6 1.1. RESEARCH CONTEXT Figure 1.1: SendGrid API plans nesses taking initial steps towards email integration. (ii) An Essentials tier that extends its offerings with a starting price of $19.95 per month, indicating that additional charges may apply based on usage. This plan is designed for businesses seeking more foundational email features, with the provision for over 100,000 emails per month. (iii) The Pro tier, highlighted as POPULAR, indicates a preferred choice among customers, starting at $89.95 per month with potential taxes and overages. This plan significantly expands the scale of operations, allowing for more than 1.5 million emails per month and broadening collaboration with permissions for up to 1,000 teammates. It also enriches the feature set with five event webhooks and includes dynamic templates alongside testing capabilities, suggesting a comprehensive solution for larger businesses or those with extensive email marketing needs. (iv) The Premier tier adopts a tailored approach with custom pricing, implying a personalized service configuration and pricing structure. The illustrative example of SendGrid showcases how services are available in the global API market as they typically offer multiple plans with different pricing options and limitations; overall, they delineate a clear escalation in service offerings and capabilities, directly correlated with the price, addressing the expected spectrum of customers with their different needs and expectations. 7 CHAPTER 1. INTRODUCTION extended by simply adding new variables and constraints. We implemented a tool that automatically transforms a LAMA into a CSOP. This contribution tackles RQ2 and part of RQ4, and it was presented in [17,18,19,20]. 4. Catalogue of analysis operations for LAMAs. We devised a catalogue of operations that can be solved with the automated tool mentioned before; this catalogue can be easily extended if more variables are added to the CSOP. This contribution completes RQ2, and it was presented in [21] over the scenario detailed in [22,23]. 5. Monitoring framework for LAMAs. We developed a monitoring framework that collects traces and metrics from a LAMA during execution time. With this information, the topology of the LAMA can be automatically inferred. Therefore, it is easier to detect possible issues with unexpected or unwanted requests that could be otherwise difficult and tedious to debug. This contribution addresses RQ3 and part of RQ4, and it was presented in [24]. Publications Accepted Journal Papers: •R. Fresno Aranda, J. S. Ojeda Pérez, P. Fernández, A. Ruiz Cortés. Governify. An agreement-based service governance framework.Software Impacts (Accepted, In press) [22]. JCR IF: 2.1 (Q3 Computer Science). International Conferences: •A. Guerrero, R. Fresno, A. Ju, A. Fox, P. Fernández, C. Müller, A. Ruiz Cortés. Eagle: a team practices audit framework for agile software development.Proceedings of the 27th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE 2019) [23]. GGS class 1 (A+), ICORE rank A*. •R. Fresno Aranda. Automated capacity analysis of limitation-aware microservices architectures.Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE 2022) [20]. GGS class 1 (A+), ICORE rank A*. •R. Fresno Aranda, P. Fernández, A. Ruiz Cortés. SLA4OAI-Analyzer: Automated Validation of RESTful API Pricing Plans.Proceedings of the 23rd International 14 1.4. CONTRIBUTIONS Conference on Web Engineering (ICWE 2023) [16]. GGS class 3 (B-), ICORE rank B. •R. Fresno Aranda, P. Fernández, A. Durán, A. Ruiz Cortés. Semi-automated Capacity Analysis of Limitation-Aware Microservices Architectures.Proceedings of the 19th International Conference on the Economics of Grids, Clouds, Systems and Services (GECON 2022) [19]. GGS class Work in Progress. National Conferences: •R. Fresno Aranda. Easy security management over microservices architectures based on OpenAPI Specification.Actas de las XV Jornadas de Ingeniería de Ciencias e Ingeniería de Servicios (JCIS 2019) [25]. •R. Fresno Aranda, P. Fernández, A. Ruiz Cortés. Towards the automation of design time capacity analysis over microservices architectures.Actas de las XVI Jornadas de Ingeniería de Ciencias e Ingeniería de Servicios (JCIS 2021). [17]. •R. Fresno Aranda, P. Fernández, A. Ruiz Cortés. Smart LAMA API: Automated Capacity Analysis of Limitation-Aware Microservices Architectures.Actas de las XVII Jornadas de Ingeniería de Ciencias e Ingeniería de Servicios (JCIS 2022) [18]. •A. Santisteban, P. Fernández, J. M. García, R. Fresno Aranda, A. Ruiz Cortés. Towards a Telemetry Specification for Capacity Analysis in Limitation-Aware Microservices Architectures.Actas de las XVIII Jornadas de Ingeniería de Ciencias e Ingeniería de Servicios (JCIS 2023) [24]. Not rated. •R. Fresno Aranda, P. Fernández, A. Ruiz Cortés. A Catalogue of Analysis Operations for API Pricing Plans.Actas de las XVIII Jornadas de Ingeniería de Ciencias e Ingeniería de Servicios (JCIS 2023) [21]. Journal Papers Awaiting Response: •R. Fresno Aranda, P. Fernández, A. Gámez Díaz, A. Durán, A. Ruiz Cortés. Pricing4APIs: A Rigorous Model for RESTful API Pricings.Computer Standards & Interfaces [15]. JCR IF: 5 (Q1 Computer Science). 15 CHAPTER 1. INTRODUCTION Research Stays •University of California, Berkeley (United States). From January 15th, 2023 to July 15th, 2023. Supervised by professor Armando Fox. This stay was funded by the Andalusian Government (Junta de Andalucía) through a Fulbright Predoctoral Research grant. Awards •The paper Easy security management over microservices architectures based on OpenAPI Specification, presented at the national conference JCIS 2019, received an award for the Best short paper. 1.5 Research methodology The research methodology employed throughout the PhD project was the Design Science Research Methodology (DSRM) [26]. This methodology is focused on creating and improving artifacts such as systems, methods, procedures and tools, with the purpose of solving specific problems. The artifacts are continuously and iteratively evaluated to improve their performance in solving the defined problem. The DSRM consists of a number of steps that must be followed to successfully apply this methodology. Specifically, the main steps of the design science process are [27]: Problem identification and motivation: Define the specific research problem to justify the value of developing an effective solution. Conceptually atomizing the problem can be useful so that the solution can capture the complexity of the problem. The problem addressed in this thesis is the lack of approaches and to develop an adequate capacity management analysis of LAMAs. This problem was identified through two means: firstly, based on our experience as software developers, where we were required to analyze the capacity manually; and secondly, through a literature review that confirmed the lack of solutions in the research community. The provision of tools and techniques to assist in the capacity analysis in LAMAs would be highly beneficial to the industry, as it would optimize their operating expenses. Definition of the objectives for a solution: Infer the objectives of a solution from the problem definition. Objectives can be quantitative, qualitative, or might be rationally deduced from the problem specification. The resources needed for this include 16 1.6. DOCUMENT STRUCTURE knowledge of the state of the problems and current solutions and their effectiveness. In this thesis, the main objective is to provide techniques and tools for the capacity analysis of LAMAs. Design and development: Create the solution artifacts, these can be broadly defined as constructs, models, methods, or instances. The desired functionality of the artifact and its architecture are determined, then the actual artifact is created. Resources needed to move from the objectives to design and development include knowledge of the theory that can be offered as a solution. In this thesis, the main artifacts are the extended Pricing4APIs and a comprehensive toolset for the capacity analysis of LAMAs that integrates various techniques and was developed during the PhD. Demonstration: Demonstrate the effectiveness of the artifact in solving the problem. This could involve its use in experimentation, simulation, a case study, a test, or other appropriate activity. Resources needed for demonstration include effective knowledge of how to use the artifact to solve the problem. Evaluation: Observe and measure the effectiveness of the artifact in solving the problem. This activity consists of comparing the solution objectives with the actual observed results from using the artifact in the demonstration, using pertinent metrics and analysis techniques. At the end of this activity, researchers may decide to return to design and development to try to improve the artifact’s effectiveness or proceed with communication and leave improvements for later projects. In our case, we focused on developing and applying the model and tool suite in different real and synthetic scenarios. Communication: Communicate the problem and its importance, the artifact, its utility and novelty, the rigor of its design and its effectiveness, through academic or professional publications, which can use the structure of this process to structure their empirical research (problem definition, literature review, hypothesis development, data collection, analysis, results discussion, and conclusion). In this phase, we disseminated our research results to the research community and other relevant audiences through presentations at top international conferences. 1.6 Document Structure This thesis document is organized as follows: 17 CHAPTER 1. INTRODUCTION Part I: INTRODUCTION. First, Chapter §1 includes the research context, the problem statements, the list of research goals and questions, contribution details (including publications) and the present section about the structure of this document. Part II: STATE OF THE ART. It includes the background for the most important concepts of this thesis. In particular, Chapter §2 introduces microservice architectures and RESTful APIs, as well as how to deploy and monitor them. Then, Chapter §3 provides information about the API economy, API pricings and marketplaces. Finally, Chapter §4 defines the concept of capacity and capacity analysis, based on definitions found on ITIL. Part III: PROPOSAL. This is the main part of the document, as it contains the contributions of the thesis. Chapter §5 introduces the exdended pricing model, as well as the validity operation and the catalogue of analysis operations. Next, Chapter §6 provides details about the concept of limitation-aware microservice architecture, including its elements. Chapter §7 describes the concept of capacity of a LAMA and introduces the CSOP transformation and the catalogue of operations. Lastly, Chapter §8 offers information about the automated monitoring framework. Part IV: VALIDATION It contains various examples of synthetic and real-world scenarios where our proposal has been validated. Chapter §9 provides some insights on the expressiveness and automated validation of RESTful API pricings. Then, Chapter §10 introduces an automated tool to transform a LAMA into a CSOP and perform various analysis operations. Finally, Chapter §11 describes an experiment for the monitoring framework. Part V: CONCLUSIONS. It includes the final conclusions and future work in Chapter §12. 18 STATE OF THE ART Part II 2 Microservice Architectures 21 This chapter provides some background for the concepts of microservice architectures and RESTful APIs. Precisely, Section §2.1 contains an introduction to these architectures. In Section §2.2 we present the elements of RESTful APIs that are typically found in real-world APIs, as well as the de facto standard for the description of these APIs. Then, Section §2.3 introduces the basics for the operation of microservice architectures, including the deployment and monitoring of their services. CHAPTER 2. MICROSERVICE ARCHITECTURES 2.1 Introduction Microservice architectures have become more popular over recent years. This is because they offer multiple benefits over other architectures which used to be more common in the past. Some of these advantages are ease of deployment and language independence. Usually, microservices are pieces of software designed to perform a very specific task. For example, common tasks include accessing a database, communicating with an external API, rendering the user interface of a web page. Because of this, they are more maintainable and scalable than a traditional architecture, since only one specific functionality inside a greater architecture is affected. Additionally, making a change in a service means that only that service needs to be deployed again. Big architectures split their functionality over multiple microservices. To communicate between each other, some means of communication are needed. Commonly, microservices make use of RESTful APIs to address this. Traditionally, systems have been developed using monolithic architectures. This means that a single application handles each and every task of the system. In contrast with microservice architectures, this approach increases maintenance difficulty. Because there is just one artifact, making changes becomes much more complex, and, therefore, adding new features or fixing some bugs mean that a complete deployment of the system is needed. Moreover, by constantly working on the same artifact, there is a high risk of causing side effects which have an impact on unintended areas of the system. Fig. §2.1 depicts a fictitious e-commerce application that is divided into multiple microservices [28]. Each one of these services provides a solution for the tasks that the system needs to perform. In this case, there is one specific service which provides the web application which includes the user interface that will be seen by the users of the application through a browser. In addition, there are three more services that take care of the different kinds of data that the shop must store: user accounts, product inventory and shipping. As can be seen in the figure, each service communicates exclusively with its own database, and offers the remaining functionality through RESTful APIs. This figure also introduces the concept of an API gateway. Its main purpose is to proxy any request that is sent to the system, so that a user who needs to access one individual service can do so through the use of the gateway. The gateway then redirects the requests to the corresponding services. API gateways are very common in microservice architectures, because, apart from routing requests, they can be set up to implement other elements, such as security and load balancing. 22 2.1. INTRODUCTION Figure 2.1: Microservice architecture example Features All microservice architectures have a series of features in common. These features are the ones which define this architectural style, and make it different from monolithic applications. The following list of features is extracted from an article by Martin Fowler [4], one of the most famous developers and followers of microservice architectures. •Componentization via services: we try to build software so that pieces can be plugged together, just as we do in the physical world. These pieces have usually been named components, and they are independent units of software that can be interchangeable. In the context of microservice architectures, services behave as components. They offer some mechanisms so that other services can communicate with them. They can also be individually deployed. One drawback of this approach is that remote calls between services are slower and more expensive, in contrast with in-app operations. •Organized around business capabilities: traditionally, development teams have been split into different technology layers: UI, server and database. This leads to costly and slow processes when dealing with changes, because all teams need to be notified and develop their part. In microservice architectures, each team is organized around business capabilities, and as such, they have the skills to take part in any layer of the full stack. 23 CHAPTER 2. MICROSERVICE ARCHITECTURES Figure 2.3: Swagger UI example OpenAPI Specification The OpenAPI Specification defines a schema to be followed when describing a RESTful API. It is language-agnostic, that is, it does not depend on any specific language. The schema consists of certain fields and attributes that specify general information and explain how the API works. OpenAPI descriptions can be written in JSON or YAML, which are similar and easily translatable to each other. For simplicity and readability, we will use YAML for the following examples. All examples are extracted from the official Swagger website [31], with minor modifications to fit into the listings. First, every API description must indicate the specific OpenAPI version that is being used. As an example, an API using version 3.0.0 should contain the line in listing 2.1.   1openapi : 3.0.0   Listing 2.1: OpenAPI version Then, additional information about the API itself should be included, similar to the one in listing 2.2. This information may consist of a title (line 2), a description (line 3), a version number (line 4), developer emails, licenses... This is usually displayed on web portals and is not tied to the actual functionality of the API.   1info: 2title : Sample API 3description : Optional multiline or single -line description in CM or HTML 30 2.2. RESTFUL APIS 4version : 0.1.9   Listing 2.2: API information Now, some servers may be specified. These servers refer to the ones where the API is or will be available. Each server consists of an URL and an optional description. It is important to note that, unlike previous Swagger versions, OpenAPI 3.0 does not support a separate base path (such as /api/v1), and it must be included within the server URL. Listing 2.3 shows an example with two servers.   1servers: 2url : http :// api . example .com /v1 3description : Optional server description , e.g. Main ( production ) server 4url : http :// staging - api . example . com 5description : Optional server description , e.g. Internal staging server   Listing 2.3: Servers information Next, the different API paths must be defined. API paths refer to the endpoints that said API supports. These include the admitted HTTP methods, query and path parameters, requests and response bodies, errors, status codes, etc. In listing 2.4, line 2 refers to the endpoint (to be appended to the server URL), and line 3 indicates the HTTP method. Lines 6 and below show how a response is defined.   1paths : 2/ users : 3get : 4summary : Returns a list of users . 5description : Optional extended description in CommonMark or HTML 6responses : 7’200 ’: 8description : A JSON array of user names 9content: 10 application / json : 11 schema: 12 type : array 13 items : 14 type : string   Listing 2.4: API paths The parameters of an endpoint may be included in the URL path, as query parameters, inside headers or cookies, or as a request body. All of these option can be defined using the OpenAPI Specification. In listing 2.5, line 6 defines the name of a parameter, whose location is specified in line 7. Lines 10 to 13 indicate its type and a range restriction.   1paths : 2/user /{ userId }: 31 CHAPTER 2. MICROSERVICE ARCHITECTURES 3get : 4summary : Returns a user by ID . 5parameters : 6name : userId 7in: path 8required : true 9description : Parameter description in CommonMark or HTML. 10 schema: 11 type : integer 12 format : int64 13 minimum: 1 14 responses : 15 ’200 ’: 16 description : OK   Listing 2.5: Path parameters Serving as a general overview, these are the basic elements to be included in an API description using OpenAPI. Nevertheless, there are many other elements that can be added, such as security measures. Some of these elements can also be parametrized. Additionally, OpenAPI leaves the standard open, so that anyone who is interested is able to contribute. Developers may add their own fields to the specification. Usually, the convention for these additional attributes is to prepend the prefix xbefore them. 2.3 Operation Deployment While microservices can be deployed in any infrastructure as any other monolithic application, they are particularly suitable for cloud deployment options. Many well-known companies offer their own solutions for the deployment of cloud services, such as Amazon (Web Services) or Microsoft (Azure). Typically, these solutions are Infrastructure as a Service (IaaS), meaning that the provider offers and maintains a cloud server and the MSA developers need to provision and manage all required software. Other platforms, such as Heroku, offer a Platform as a Service (PaaS) solution, where developers only need to upload the code of the service and it is automatically deployed. In this section, we will focus on some options offered by Amazon Web Services (AWS), as well as deployment variations using Docker and Kubernetes. AWS Elastic Cloud Computing (EC2)1offers various tiers with different infrastructure 1https://docs.aws.amazon.com/ec2/ 32 2.3. OPERATION features, such as types of CPU and optimization for short bursts of requests. Within each tier, there are multiple options with different types of CPU, storage and RAM, among others. When a developer chooses an option, they get a key to remotely access the server, so that they can install the necessary packages and software to deploy the service. It is the responsibility of the developer to maintain and update the packages, while Amazon takes care of the underlying physical infrastructure. Fig. §2.4 shows the architecture of a sample AWS EC2 instance [32]. Figure 2.4: Architecture of an AWS EC2 instance deployed on an Amazon Virtual Private Network Among the available AWS EC2 tiers, one of the most popular ones is T2. These machines are small and cheap, and are optimized for small services which are expected to receive short bursts of requests. Within the T2 tier, the most used option is the t2.micro instance. It is free to use for up to 750 hours of computing time, meaning that it is used by many developers who want to try AWS or deploy small services which are not expected to handle too much workload. Another option is the A1 tier, which is relatively new (released in 2018). It offers machines with ARM CPUs, which are becoming more popular over the last few years. These instances are more expensive than traditional instances based on x64 CPUs, but offer more performance in some cases. 33 CHAPTER 2. MICROSERVICE ARCHITECTURES AWS offers other alternatives to traditional IaaS solutions. In particular, a popular option is AWS Lambda2, a serverless alternative that is based upon the idea of Function as a Service (FaaS). In this approach, the microservice is not a deployed as a regular server; instead, it is deployed as a single function that receives some parameters, performs some actions and, if necessary, returns some data. Furthermore, in contrast to AWS EC2, the developer does not need to manage any packages within the server. Small services deployed as functions typically run considerably faster than the same service in a standard server, making Lambda an ideal solution for microservices that perform a small and very specific task. A service using AWS Lambda may be used together with other AWS components that act as middlewares between the clients and the function. An example of this is the utilization of an API gateway, which AWS also offers (AWS API Gateway). An alternative to the traditional way of building and sharing applications is Docker3. It is a well-known containerization technology that works by encapsulating a service and all of its dependencies into a single image, which can then be distributed and shared anywhere. This facilitates the replication of deployment infrastructures, as an image includes all requirements for a service to run and will work the same in any computer. Building and downloading images is easily done through a command-line interface, and it only takes a few seconds to share an application with other developers. There are other similar technologies that work very similarly, such as Podman, but Docker remains the main solution for the containerization of services. Fig. §2.5 depicts the architectural differences of containerized deployments versus other types of deployment [33]. A Docker image can be used to deploy a service to a traditional IaaS platform, such as AWS EC2. Instead of installing all packages and dependencies of the service in the EC2 instance, a developer would only need to install Docker and download the image of its service. If it is an MSA with multiple services, databases, etc., Docker can automatically deploy all components thanks to the use of the docker-compose tool. However, a more interesting approach is also available. Kubernetes, or K8s4, is a platform that orchestrates containers to automatically manage their resources, balance their loads, scaling, monitoring and many other tasks. It was originally developed by Google, but is now open-source. Kubernetes works as a mix of IaaS and PaaS, giving developers the ability to control some parts of the infrastructure, while also automating many of the most common tasks. 2https://docs.aws.amazon.com/lambda/ 3https://www.docker.com/ 4https://kubernetes.io/ 34 2.3. OPERATION Figure 2.5: Traditional vs virtualized vs containerized deployments Monitoring Once the microservices have been deployed, it is necessary to ensure that each service operates and communicates correctly with other services and APIs, as each service functions independently within the MSA [34]. While research on telemetry for the Internet of Things (IoT) field has been widely explored, its application for RESTful architectures is still under-researched. Even in the IoT field, most of the telemetry data extracted from IoT devices comes from the network layer [35], while REST APIs operate and limit requests at the application layer. This is why there is a need to collect data coming from the application layer, which is referred to as application level telemetry. Fortunately, there exist specifications that define application level telemetry in a standardized format, such as OpenTelemetry [36], which provides a vendor-agnostic way to collect telemetry traces from external sources and send them to different targets through collectors for analysis and visualization. A high-level overview of OpenTelemetry is shown in Fig. §2.6 [37], and the collected traces follow the format shown in listing 2.6.   1{ 2"trace_id": "7bba9f33312b3dbb8b2c2c62bb7abe2d", 3" span_id ": "086 e83747d0e381e ", 4" name ": "/ v1 / sys / health ", 5" start_time ": "2021 -10 -22 16:04:01.209458162 +0000 UTC ", 6" end_time ": "2021 -10 -22 16:04:01.209514132 +0000 UTC ", 7" attributes ": { 8" net . transport ": "IP. TCP ", 9" net . peer. ip ": "172.17.0.1" , 10 "net. peer . port ": "51820" , 11 " net . host. ip ": "10.177.2.152" , 12 "net. host . port ": "26040" , 35 CHAPTER 2. MICROSERVICE ARCHITECTURES Figure 2.6: OpenTelemetry diagram 13 " http . method ": " GET ", 14 " http . target ": "/ v1 /sys / health " , 15 " http . server_name ": " mortar - gateway ", 16 " http . user_agent ": " Consul Health Check " , 17 " http .host ": "10.177.2.152:26040" 18 } 19 }   Listing 2.6: Fragment of an OpenTelemetry trace Prometheus [38] is another popular specification used for collecting metrics from remote sources and exposing them in a consistent manner, as shown in listing 2.7. It has gained widespread adoption in the industry due to its ease of use and flexibility in handling a variety of metric types. Fig. §2.7 shows an example of a dashboard with data collected from Prometheus [39]. However, as with OpenTelemetry, it is a general-purpose specification that may be used for different tasks apart from RESTful architectures, such as monitoring storage or detecting system anomalies. As a result, the transmitted packages become quite large due to the increase of attributes, which may introduce unnecessary overhead in the network, ultimately affecting the performance of microservices. 36 2.3. OPERATION Figure 2.7: Example of Prometheus data visualized in a Grafana dashboard   1# TYPE assets_parse_errors_total counter 2join_parse_errors_total 0 3# TYPE join_http_response_time_seconds summary 4join_http_response_time_secon ds_sum { status ="200"} 0.6420005 5join_http_response_time_secon ds_sum { status ="404"} 0.59900004 6join_http_response_time_secon ds_sum { status ="301"} 0.1440001   Listing 2.7: Fragment of Prometheus metrics In the cloud native landscape5, tools that leverage OpenTelemetry and Prometheus data, such as NewRelic6or Honeycomb7, provide useful information and visual representations of telemetry data. 5https://landscape.cncf.io/ 6https://newrelic.com/ 7https://www.honeycomb.io/ 37 CHAPTER 2. MICROSERVICE ARCHITECTURES The increasing adoption of MSAs in modern software development has brought about the necessity of analyzing their capacity. Further information and definitions about the capacity of an MSA is available in Chapter §4. While researchers have tackled this issue in the past, their focus has often been on performance analysis rather than capacity, which is a more complex metric that encompasses factors beyond other metrics used in performance analysis. Consequently, the literature on the subject has offered a range of frameworks and methods for measuring Quality of Service (QoS), with most of them relying primarily on performance metrics such as response times, network latency or resource consumption. However, none of these approaches have adequately addressed the problem of capacity. Telemetry data can provide valuable insights into capacity analysis of RESTful architectures by allowing the inference of their topology through service endpoint call graphs [40]. This can be achieved by analyzing traces containing information about the source and target, as well as the operation being performed. Additionally, this approach can provide the actual number of requests made by each service to an API per operation, which is a crucial aspect of the capacity problem. Efforts have been made to develop frameworks for extracting telemetry data [41,42]. However, neither of these frameworks was designed to use the resulting metrics for a specific purpose, nor do they provide a concrete model or specification for application level telemetry. This creates a gap in the monitoring of MSAs, as it limits the ability to measure more complex metrics like capacity. This same problem arises with specifications like OpenTelemetry and Prometheus. They have not been designed for a specific purpose other than collect telemetry data, why is why they are intended to be used as frameworks that leverage tools implemented over them, such as NewRelic or Honeycomb. However, currently these tools are not able to provide effective means to answer capacity questions, and the underlying technology is prone to input undesired overhead into the network. In response, researchers have proposed models for measurable QoS goals and frameworks for extracting metrics to evaluate the monitored system against these goals [43]. Tundo et al. [44] proposed a model that contains measurable QoS goals based on the ISO/IEC25011 standard and implemented a framework that extracts key performance indicators (KPIs) based on user needs. Nevertheless, both of these approaches focus on performance analysis since their purpose is to reduce the impact of monitoring overhead on performance and meet monitoring goals based on performance indicators, respectively. 38 3 API Economy 39 This chapter introduces the API economy, a term that has gained traction over the last years. First, Section §3.1 provides an introduction to the concept of API economy. Next, Section §3.2 contains information about the concept of API pricings and their elements and features. Then, Section §3.3 presents the idea of API marketplaces and their benefits for consumers and providers. CHAPTER 3. API ECONOMY vastly evolved over the last years, as it has absorbed and joined multiple marketplaces and API mashup websites. Fig. §3.3 shows the main page of RapidAPI [50]. Figure 3.3: RapidAPI main page Besides RapidAPI, there are other marketplaces such as APILayer3. Some major companies also offer their own versions of marketplaces, such as Amazon Web Services. 3https://apilayer.com/ 46 4 Capacity Analysis 47 This chapter introduces the concept of capacity as it is known in the context of IT service management and engineering.Section §4.1 provides an introduction to the capacity of a system. Then, Section §4.2 describes the concept itself and contextualizes it within ITIL 4 (a widely used management framework), including some insights on the capacity management practice. Next, Section §4.3 introduces the capacity analysis of a microservice architecture, detailing current approaches for this analysis and indicating why they lack support for the specific features of MSAs that consume external APIs with limitations. CHAPTER 4. CAPACITY ANALYSIS 4.1 Introduction The capacity of a system is a relatively wide term. In general, the capacity refers to the maximum workload that the system is able to handle, under certain circumstances. Furthermore, the workload does not have a standard metric, although most analyses use the number of users or requests. The circumstances depend on the system under analysis. For example, a business might be interested in analyzing the capacity of one of its systems considering the limitations of the deployment infrastructure. If this infrastructure is robust and scalable, the system should be able to handle bigger workloads. Therefore, if the system supports a bigger workload, it means that it can support more users and in turn generate more revenue for the business. This scenario showcases the importance of the analysis of the capacity. Additionally, costs also play an important role in the capacity analysis of a system. Following the example in the previous paragraph, if the system is not expected to have a high number of users, it is not advisable to invest a disproportionate amount of money in a very expensive deployment infrastructure. This is because the revenue would not be high enough to sustain the infrastructure costs, and the business would lose money. Again, this demonstrates the necessity of analyzing the capacity of a system to ensure that it works under optimal operating conditions, while keeping costs under control. In the following sections, we provide a comprehensive analysis of the concept of capacity within ITIL 4, a widely used IT management framework where the capacity management plays a very important role. After the importance of this management is stated, we then introduce the concept of capacity analysis to the scope of microservice architectures. 4.2 Capacity Analysis in ITIL The Information Technology Infrastructure Library (ITIL) emerged in the late 1980s, as a response to the growing dependence on Information Technology (IT) in the business world. Developed by the Central Computer and Telecommunications Agency (CCTA), a government agency in the United Kingdom, ITIL was initially focused on standardizing the processes involved in managing IT services. This standardization was aimed at ensuring efficiency and predictability in IT service management (ITSM), which was becoming increasingly complex. The first version of ITIL, known as ITIL v1, was a collection of more than 30 volumes, 48 4.2. CAPACITY ANALYSIS IN ITIL each detailing various aspects of IT service management. As ITIL gained popularity, it underwent several revisions to keep up with the evolving landscape of IT. The most notable of these revisions were ITIL v2, introduced in the early 2000s; and ITIL v3, which was released in 2007 and later updated in 2011. ITIL v3 expanded the framework to an approach based on the life-cycle of the service management, covering service strategy, service design, service transition, service operation and continual service improvement. The latest iteration, ITIL 4, was launched in 2019. This version focuses on integrating ITIL guidance with best practices from other methodologies like Agile, DevOps, and Lean, acknowledging the need for a more flexible, collaborative, and customer-centric approach in the digital era. ITIL has become a global standard in IT service management for several reasons. Firstly, it provides a comprehensive, consistent, and coherent set of best practices, promoting efficiency and effectiveness in IT service management. By adopting ITIL, organizations can improve service delivery, increase customer satisfaction and optimize resource utilization, leading to cost savings. In such a context, ITIL helps in aligning IT services with the overall business strategy. This alignment is crucial as IT plays a pivotal role in the modern business environment. It ensures that IT services support business goals, enhancing the overall value creation of the organization. Additionally, ITIL offers a common language for IT professionals, enabling better communication and understanding within and between organizations. This aspect is particularly important in a globalized business world where cross-functional and cross-organizational collaboration is the norm. From a high level perspective, ITIL is structured around a service life-cycle composed by five stages (see Fig. §4.1): •Service Strategy: This stage involves the development of policies and strategies to serve the overarching business goals. It includes aspects like service management as a strategic asset, defining the market and financial management. •Service Design: In this stage, ITIL focuses on designing IT services, including architectural, process, policy and documentation design, to meet current and future business requirements. •Service Transition: This involves the transformation of services into a live operational environment, focusing on change management, release and deployment, service validation and testing. •Service Operation: At this stage, ITIL addresses the efficient and effective deliv49 CHAPTER 4. CAPACITY ANALYSIS ery and support of services, ensuring that the value is realized in the operational phase. Key processes include incident management, event management and request fulfillment. •Continual Service Improvement: This stage is about aligning and realigning IT services to the changing business needs by identifying and implementing improvements to IT services and processes. Figure 4.1: ITIL overview Each stage of the life-cycle includes processes and functions that are guided by ITIL’s principles. These principles are designed to be adaptable and flexible, allowing organizations to apply them according to their specific context, needs and maturity level. In the ITIL framework, the capacity appears as a cornerstone of the Service Design and Service Operation stages and including an specific capacity and performance management practice, which is classified under the service management category, defined as: 50 4.2. CAPACITY ANALYSIS IN ITIL This practice helps organisations ensure that their products and services meet expected performance levels. It also addresses current and future demands, helping organisations identify any changes that could affect their capacity1. In a complementary manner, the capacity and performance management is defined as: The purpose of the capacity and performance management practice is to ensure that services achieve agreed and expected performance, satisfying current and future demand in a cost-effective way2or The objective of ITIL capacity and performance management is to ensure that your IT capacity meets your business needs. Satisfying current and future demand in a timely and cost-effective way is key to this ITIL practice3. As observed in these definitions, the capacity of a service is related to the demand and the business needs of an organization. It involves the analysis of the current demand, so that the organization can act accordingly to meet customer needs and also take the corresponding actions to satisfy the expected future demand. The fact that the term capacity is always paired with performance leads to the conclusion that both concepts are closely related and, therefore, successfully managing the demand of a service also means ensuring that the performance of the service is optimal. Furthermore, two of the three definitions above talk about cost-effectiveness. This means that capacity management directly affects the cost of a service, and thus it is vital for organizations which do not wish to waste any of their capital. In fact, the proper management of the capacity affects both income and expenses. By ensuring that the service meets the customer requirements, organizations can earn their trust and thus continue using the service, while avoiding unwanted service disruptions derived from underprovisioned infrastructure (e.g. cheap servers that allow a lower load than needed). On the other hand, organizations also need to make sure that the capacity of the service is not exceedingly high, so that a considerable part is wasted (e.g. very expensive servers that allow a very high load that is never achieved). Trying to balance costs is a fundamental part of capacity management. The third definition also mentions the need to satisfy the demand in a timely way. This means that the service must meet the capacity requirements when its customers actually need it, and predict potential usage spikes. By correctly analyzing the evolution of demand over time, organizations can adapt their capacity accordingly and, therefore, successfully meet the customer demand. This analysis can also help in the optimization of 1https://www.itgovernance.co.uk/blog/what-are-the-itil-4-management-practices 2https://wiki.process-symphony.com.au/framework/lifecycle/process/ capacity-and-performance-management-itil-4 3https://www.smartsheet.com/content/itil-capacity-management 51 CHAPTER 4. CAPACITY ANALYSIS costs (e.g. choosing and changing to a cheaper or more expensive server based on demand predictions). In the following subsections we explore how the management of Service Level and Capacity in ITIL contextualize the need of an operational analysis that is specially useful in the context of IT infrastructure management based on MSAs. Service Level One of the key processes related to capacity within the ITIL framework is the Service Level Management (SLM) process. In this context, the SLM process is critical to the success of any IT service organization as it is focused on ensuring that the agreed-upon service levels are met and that customers are satisfied with the services they receive. This process works closely with the Business Relationship Management (BRM) process, which is responsible for managing the relationship between the IT organization and its customers. The SLM process is responsible for understanding the Service Level Requirements (SLRs) of customers and translating them into Service Level Agreements (SLAs). These agreements serve as the rulebook for measuring and reporting the level of service delivered. The SLRs determine the contingencies to be provided in the architecture, such as high-availability architecture, and the resources required for delivering the service. This, in turn, has a direct bearing on the cost of the services. Moreover, with the rise of paradigms like DevOps, there is a growing need to factor in development measurements during customer discussions and negotiations. DevOps is a methodology that emphasizes collaboration between development and operations teams to deliver service and services more quickly and efficiently. The SLM process needs to be able to work with DevOps teams to ensure that development measurements are taken into account when negotiating SLAs. In addition to measuring the operational aspects of a service, the SLM process also needs to be able to measure the customer experience. This includes factors such as response times, resolution times and overall satisfaction with the service. By measuring the customer experience, the SLM process can identify areas for improvement and work with other processes within the ITIL framework to address these issues. Overall, the SLM process is a critical function within the ITIL framework. It is responsible for ensuring that services are delivered to customers at the agreed-upon service levels and that customers are satisfied with the services they receive. With the rise of DevOps and the need to factor in development measurements, the SLM process needs to 52 4.2. CAPACITY ANALYSIS IN ITIL be able to adapt to changing circumstances and work closely with other processes within the ITIL framework to deliver high-quality services. Moreover, the SLM process should be seen as an integral part of the ITIL framework that ensures that services are delivered at the agreed-upon level. This process is responsible for monitoring service performance against SLAs and taking corrective action when necessary. The SLM process also involves regular service reviews with customers to ensure that their changing requirements are met, and service levels are adjusted accordingly. As a consequence, the SLM process is critical for organizations that want to deliver high-quality services to their customers. It helps to establish a clear understanding of the expectations and requirements of the customers and ensures that the services provided meet those expectations. The process also helps to identify any areas where the service may be falling short and take corrective action to improve it. In this context, the SLM process needs to factor in development measurements during customer discussions and negotiations since current service engineering practices focus on delivering high-quality services at a faster pace. To achieve this goal, ITIL promotes the idea that it is essential to have a clear understanding of the key performance indicators (KPIs) that will enable teams to measure their progress and success. In fact, the SLM process plays a critical role in ensuring that IT service engineers can deliver high-quality service that meets the changing demands of their customers. By incorporating development metrics into the SLM process, organizations can improve their service development and deployment processes and stay competitive in the market. Consequently, the SLM process is a crucial component of the ITIL framework that helps organizations deliver services at the agreed-upon level and ensures that customers’ changing requirements are met. Moreover, the SLM process ideally must adapting to incorporate development metrics to enable teams to measure their progress and success. By doing so, organizations can deliver high-quality services that meet the changing demands of their customers and stay competitive in the market. In the industry, ensuring high-quality performance is essential. However, many times, expected measurements are done at a contractual level without a process defining the ins and outs of the data sources and the measurement of performances. This can lead to a disconnect between the goals of the service provider organization and the actual results achieved. This is where service level management comes in as a crucial element for the service development and operation. Specifically, service level management involves identifying KPIs and keeping track of how things are moving along. By doing so, it helps IT service engineers stay on track and deliver high-quality service at a faster pace. This 53 CHAPTER 4. CAPACITY ANALYSIS is especially important in the current market where speed and efficiency are critical to staying ahead of the competition. Agreement of service levels happens at multiple levels, including SLAs with customers, operational level agreements (OLAs) within the same organization and SLAs with suppliers. These agreements ensure that everyone involved in the service development and deployment process is on the same page and understands their roles and responsibilities. This not only helps to avoid potential conflicts but also ensures a smooth and efficient workflow. To enable speed, multiple toolsets or frameworks are employed across various environments, and they are usually managed by separate teams. Therefore, it is imperative that the agreements (either OLA or SLA) exist to ensure guarantees for speedy deliveries. With the right agreements in place, IT teams can work more efficiently and deliver high-quality services faster than ever before. Another crucial aspect in the context IT services deployed in cloud environments, is the availability. Customers expect seamless access to services they have subscribed to or use, and any downtime can have a significant impact on their satisfaction levels. This is where service level management and availability management come into play. Service level management involves agreeing on the expected levels of service with the customers. This includes factors such as response times, uptime and other performance metrics. Once these levels are established, availability management takes over to ensure that the service is built to meet and even surpass these expectations. The level of availability required can vary significantly depending on the nature of the service. For instance, a service that requires 99.95% availability may have a very different architecture compared to one that needs 99.99% availability. Even though the difference is just a few decimals, the impact on the underlying infrastructure can be massive. To achieve a higher availability, the service architects have to factor in multiple layers of contingencies to ensure that even multiple failures do not take down the service. This could involve redundant systems, backups and failover mechanisms to ensure that the service remains available even in the event of a catastrophic failure. The cost of providing services with such high levels of availability can be multiples of those offered at lower levels. Managing availability in ITIL is highly aligned with the DevOps philosophy. As an addition, DevOps introduce a more concrete operational framework that promotes the usage of automation and continuous monitoring to detect and fix issues before they 54 4.2. CAPACITY ANALYSIS IN ITIL impact the service. They may also represent the usage of tools such as load balancers and auto-scaling to ensure that the service can handle spikes in traffic without affecting availability. In today’s fast-paced and highly competitive business environment, organizations are constantly striving to deliver services that meet and exceed customer expectations. This is where perspectives such as DevOps comes into the picture, as they provide a comprehensive framework for managing the entire service development lifecycle, from planning and development to deployment and maintenance in an operational way. However, its implementation comes with its own set of challenges, one of which is setting availability targets. In order to achieve this, organizations need to have a clear understanding of the different environments that are required for building and testing service. This includes environments that are not customer-facing but are still critical to the overall development process. Moreover, multiple infrastructures and platforms are used in the service development, deployment and operation, and these elements need to be available as per the team’s needs. Any lapse in the availability management of test environments, infrastructures or platforms can result in delayed delivery of services, which can have a negative impact on the overall delivery levels set forth by the service level management process. Therefore, it is important to carefully analyze the availability requirements to ensure that they are in perfect alignment with the delivery rate and the speed at which the team can deliver. This requires a rigorous approach to availability management, which involves monitoring and tracking the availability of various environments, and taking proactive measures to address any issues that arise. Another aspect to consider is over-delivering on availability. While this may seem like a good thing, it can actually be detrimental to the business angle of services. Every additional decimal of availability adds exponentially to the cost of services, which can have a significant impact on the bottom line. Therefore, organizations need to strike a balance between availability and cost-effectiveness to ensure that they are delivering services that meet customer expectations while also being financially sustainable. In conclusion, service level management is an essential aspect from the ITIL Perspective. It helps to ensure that the goals of the project are aligned with the actual results achieved. By identifying KPIs and tracking progress, service level management helps teams stay on track and deliver high-quality services at a faster pace. Agreements at multiple levels ensure that everyone is on the same page and understands their roles and 55 CHAPTER 4. CAPACITY ANALYSIS was not the case in the previous approach, where a provider only had one plan for a task. Additionally, an MSA might send multiple requests to the same provider. •There might be cases where there is no solution to the problem, e.g. the MSA is not able to serve enough requests to meet user requirements given the subscribed pricing plans. In previous approaches, this aspect was not taken into account. We are not aware of any existing proposal which analyses the capacity of an MSA with external APIs regulated by pricing plans. The most similar proposal, which has also been a major inspiration for this thesis, is ELeCTRA by Gamez–Diaz et al. [54]. Based on the limitations of an external API (specified in its pricing) and the topology of an MSA with a single entrypoint, ELeCTRA computes the maximum values of the limitations that the entrypoint of the MSA will be able to offer to its customers. Assuming that the topology of the MSA does not vary, these maximum values are determined solely by the values of the external API limitations, i.e., they are induced by them. This analysis of the limitations induced in an MSA is performed by ELeCTRA by interpreting the problem as a CSOP and using MiniZinc [55] as a solver. Unfortunately, ELeCTRA’s capabilities are insufficient to automatically analyze the capacity of a MSA. Its main limitation is to consider that a pricing consists of a single constraint, or a single quota or a single limit. Thus, it is not possible to model prices, overage cost, or specify several limits (quotas and rates) in the same pricing. 62 PROPOSAL Part III 5 Extended Pricing Model 65 This chapter presents our proposal for a pricing model that includes the required elements for the capacity analysis of MSAs. In particular, Section §5.1 contains a short introduction. Next, Section §5.2 introduces the pricing model itself, which extends an existing model to add some elements that were originally missing. Then, Section §5.3 includes the serialization of the model, aligned with the OpenAPI Specification. Finally, Section §5.4 contains a catalogue of analysis operations for API pricings, classified by different scenarios and with some examples. CHAPTER 5. EXTENDED PRICING MODEL 5.1 Introduction As discussed, the non-functional description of an API (e.g. its pricing) has not been addressed in the most used standard in the industry, the Open API Specification (OAS); consequently, the actual pricing definition depends on ad-hoc approaches by the API provider [56]. This makes it harder to leverage the information included in a pricing, because each organization may describe it in a completely different way. These differences include different formats in the definition of the billing conditions or the usage limitations, mixing and confusing terms such as rates and quotas that do not have specific definitions. Furthermore, it is difficult to check the validity and correctness of a pricing when its description is not standardized, as it is not possible to have an exact definition of the validity of a pricing. In previous works, Gamez-Diaz et al. presented Governify4APIs [49], a standardized model for RESTful API pricings that was devised after an exhaustive analysis of realworld APIs in the industry. Based on this model, the same authors created SLA4OAI [57], a proposed extension of OAS that includes non-functional information and aims to overcome the aforementioned issues. It is a work-in-progress specification, written in JSON or YAML, that is under discussion by organizations and practitioners [58]. The formal description of an API pricing is useful, for example, to check for its correctness or automatically choose the best plans based on user needs [19]. Consequently, in this chapter we present the Pricing4APIs model, an extension of the Governify4APIs model that includes some missing elements, along with its serialization in SLA4OAI. Furthermore, we present a first iteration of a catalogue of analysis operations for API pricings, including various examples to demonstrate how to solve them. 5.2 Pricing4APIs In order to address a more capable capacity analysis in LAMAs, the proposed Pricing4APIs model takes the model presented in [49] as starting point and extends it. As a high-level overview, the Pricing4APIs model allows to define a set of plans with its associated cost; for each plan, a set of limitations (i.e. quotas and rates) over the potential API operations can be defined. In the context of the RESTful paradigm, those operations are bounded to an HTTP path and method. Fig. §5.1 illustrates the complete Pricing4APIs model. To enhance clarity, we have divided it into three distinct areas: (i) the yellow elements, which are related to pricing, 66 5.2. PRICING4APIS plans and cost; (ii) the pink elements area, which addresses limitations and limits; and (iii) the blue area, which concerns capacity. In the subsequent sections, we will elaborate on each aspect of the model, utilizing examples from the external API E1 in Fig. §1.3. To do so, we will consider each part individually: the plan area, the limitations area, and the capacity area. Figure 5.1: Pricing4APIs model for API pricings Pricing, Plans, and Cost As illustrated with yellow elements in Fig. §5.1), it is shown that Pricing is structured around various Plans, each labeled with a unique name and associated with a specific Cost that defines the charge for accessing the service. For instance, API E1 offers two types of plans: a Basic plan and a Premium plan. The Cost associated with a Plan might be straightforward, such as a set fee of $5 or $8 in the given scenario. However, the cost can also vary based on additional parameters. In scenarios where the cost is contingent upon a Limitation, it bifurcates into two distinct categories: OperationCost, applied each instance an Operation is executed, 67 CHAPTER 5. EXTENDED PRICING MODEL and OverageCost, applied after surpassing a predetermined threshold of the Limitation, with fees calculated based on the excess usage. These Cost categories may adopt a periodic structure, articulated through a Period that specifies the duration and the TimeUnit associated with it. For example, the charge for the Basic plan is set at $5 monthly, indicating a Period of 1 TimeUnit (month). According to the model, the fee tied to performing a service operation hinges on the relevant Limitation. This could lead to differing plans, such as Plan A, which levies a $0.10 charge per request, and Plan B, which offers a package of 1,000 requests at $75. This arrangement allows clients to select the option that aligns with their usage patterns and budgetary constraints. Should a client’s usage exceed the predefined quota, an OverageCost may be applied, ensuring continued service access rather than termination. For example, surpassing the limit of 1,000 requests per day in our scenario results in a $0.01 fee for each additional request. Limitations and Limits As illustrated (highlighted by pink elements in Fig. §5.1), for the effective regulation of API usage, each Operation within a Plan is subject to Limitations based on a specified Metric. The prevalent form of Limitation is the ThresholdedLimitation, which sets one or several ThresholdLimits for the quantity of Metric units within a Period, typically adhering to a max ThresholdType. This signifies the maximum allowable Metric units. Through the imposition of Limitations, providers can tailor API consumption to align with the overarching Capacity of the platform. An Operation is characterized by a combination of an HTTP method and path, essentially defining an endpoint. Metrics such as the number of requests are commonly used, but others like storage,bandwidth, or CPU usage are also viable. Metrics can interrelate, for instance, where each API request consumes 2KB of bandwidth, thereby establishing a MetricRelationship of 2between the number of requests and bandwidth per request metrics. The methodology for updating Limitations Metric distinguishes two categories of ThresholdedLimit: a Quota, calculated over a static time frame, and a Rate, which employs a sliding time window relative to the initiation of the metric count. For instance, aone-week sliding window starts and ends based on the timing of the initial metric unit, unlike a one-week static window which operates on a fixed schedule from Monday to 68 5.2. PRICING4APIS Sunday. The distinction between sliding and static windows is visually represented in the bottom right of Fig. §5.2, particularly for the number of requests metric. This visual aids in understanding how occurrences differ based on the type of window applied: a 1-second sliding window shows four occurrences, whereas a 1-second static window records only two. This demonstrates the choice between sliding (rate) and static (quota) windows affects the tally of observed events. To prevent more than 4 requests per second, two strategies are viable: a 1-second sliding window or a 1-second static window, each with its limits to control request frequency. This dual approach showcases the flexibility in managing API usage limits. Industry practices often adhere to defined patterns [56] for setting ThresholdedLimits, with Quotas spanning longer periods (daily to yearly) and Rates focusing on the number of requests over shorter intervals. The challenge lies in the ambiguity of ThresholdLimit types in documentation, necessitating empirical testing by API users for clarification. In this scenario, the Basic plan incorporates both a Rate and a Quota, demonstrating a common strategy for defining usage limitations. The model differentiates between ThresholdedLimitation and ThresholdedLimit, with the former setting a portion of the service’s Capacity for a specific metric and operation. A ThresholdedLimitation might be detailed as a series of ThresholdedLimits, like 30 requests every 1 week and 1 request every 1 second, offering a structured limit on capacity usage. Alternative representations for Limitations, such as frequency distributions [59], provide further flexibility, although this discussion primarily focuses on ThresholdedLimits due to their prevalence in the industry. This example illustrates the typical application of rates for immediate capacity limits and quotas for long-term business considerations (as identified in [56]). Capacity A vital detail often omitted from public documentation regarding pricing or plans is the concept of Capacity. This element is intrinsic to the service provider’s infrastructure and encapsulates the various limitations imposed by the platform or system’s capabilities. These limitations are inherently tied to both technical specifications and financial considerations, such as processing power, memory allocation, and the architecture’s scale. The process of determining a service’s Capacity is essential for establishing Pricing strategies and evaluating Limitations. Specifically, the service’s Limitations should 69 CHAPTER 5. EXTENDED PRICING MODEL align with its Capacity to ensure they are not exceeded, maintaining operational integrity. Illustrated in the framework (marked by blue elements in Fig. §5.1), identifying the Capacity leads to its categorization akin to a Limitation, specifying the allowable quantity of a particular Metric within a predetermined Period. Consequently, similar to Limitation,ThresholdedCapacity is defined by a threshold value and a ThresholdType (typically max refering to the maximum function) for a specified Period and TimeUnit. Expressing Capacity in terms of the requests per second (RPS) metric for each operation and plan is common. For instance, a 10,000 RPS limit for the GET /contacts operation under the free plan implies that all users under this plan can collectively issue up to 10,000 requests per second. Notably, Capacity may vary across different plans to reflect the diverse levels of service provided through varied infrastructure setups. For instance, based on rigorous performance and stress testing, an organization may determine that its system can support up to 10 000 RPS. If there were a restriction of 10 requests per second per client, the maximum number of clients that could be served concurrently would be calculated as 10,000/10 =1,000. When evaluating Limitations, the percentage of capacity utilization or the percentage of utilization (PU) becomes an invaluable metric. This figure crucially influences the feasibility of imposing a Limitation, as setting one becomes untenable if the PU exceeds 100%. The PU will depend on how a consumer consumes the API. There are two interpretations given a Limitation: uniform and burst. Therefore, the PU can be calculated in two different ways. To illustrate this idea, let us consider a ThresholdedLimitation with a single ThresholdedLimit of 43,200 requests every 1 day: An API user might infer that, because a day encompasses 86,400 seconds, the rate per second is computed as 43,200/86,400 =0.5 requests. This estimation is predicated on a uniform distribution, gradually enabling the user to exhaust the daily quota of 43,200 requests. This illustration typifies the minimum PU. On the contrary, the ThresholdedLimitation clarifies that up to 43,200 requests are permissible within a single day, without any restrictions on executing all these requests instantaneously at the start of the day. Consequently, a user is theoretically capable of deploying all 43,200 requests in a singular second. This situation denotes a burst distribution and is indicative of the maximum PU. Consequently, the PU must take both these models into account, so that we define the bounded PU (BPU) as this range: 70 5.2. PRICING4APIS 1. The lower bound is the minimum PU, in which a uniform distribution of utilization over the period is assumed. 2. The upper bound is the maximum PU, which assumes the utilization of the maximum allowed in a single burst. Fig. §5.2 illustrates different consumption scenarios for the same ThresholdedLimitation of 60 requests every 60 seconds. 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 1 22 23 24 25 26 27 28 29 30 21 57 58 59 60 1 req/sec 60 req/min 2 req/sec 60 req/min 3 req/sec 60 req/min 6 req/sec 60 req/min t Rate 4 req/sec t+1 t+2 Quota 2 req/sec Figure 5.2: Examples of different consumption scenarios for the same ThresholdedLimitation If we consider an uniform rate of consumption, where 60 requests are made in 60 seconds, this would be the same as making 1 request per second. However, in a burst scenario, the number of requests made in 1 second can range from 2 to a maximum of 60. So, when calculating the BPU (Burst Per Unit) within the limitation of 60 requests every 60 seconds, we need to consider the minimum value of 1 request per second for a uniform distribution, and the maximum value of 60 requests in 1 second for a burst within a 1-minute timeframe. It is crucial to emphasize that an accurate evaluation of capacity plays a vital role in ensuring reliable service and efficient resource utilization in the ever-changing API economy. When dealing with a diverse range of clients, each with their own unique plans and service requirements, the challenges are multiplied. For example, consider multiple consumers, each with different limits on the number of requests they can make per second. The provider must not only guarantee that each consumer receives their entitled 71 CHAPTER 5. EXTENDED PRICING MODEL If the answer is negative, a useful auxiliary operation is Why not?. This would help the consumer pinpoint the specific rate or quota that is acting as a bottleneck. Example: we have one subscription to plan Pro and want to send 10 RPS and 2,000 RPD. This would not be possible, as 2,000 RPD is equivalent to 60,000 requests per month (RPM), which exceeds the quota of Pro. We would either need to upgrade to Ultra or use overage costs. O3.What is the maximum speed at which I can send requests, and for how long? The rates of the subscribed plans will define the maximum speed at which the consumer may send short bursts of requests, while the quotas will determine how many bursts can be sent. Example: we have one subscription to plan Pro. Therefore, the maximum speed at which we can send requests is 10 RPS, which is the rate of Pro. At this speed, we would consume the entire quota in 4,000 seconds. We would then need to wait until next month for the quota to be reset. If we are willing to use overage costs, we could extend this time. O4.What is the maximum time that I will need to wait after consuming all quotas? If the consumer uses all of their available quotas and there are no overage costs (or the consumer does not want to use them), they will need to wait until the quotas are reset. Some APIs actually return the waiting time when a consumer runs out of requests, but only for a single subscription. If the consumer has multiple subscriptions, this information might not be as useful. Example: let us take the example in O3as a starting point. At full speed, we would consume the quota of Pro in 4,000 seconds. After that, we would need to wait until next month for the quota to be reset. If we consumed the quota during the first 4,000 seconds (which is a bit more than 1 hour) of a 30-day month, we would need to wait 29 days and a bit less than 23 hours for the quota to be reset. Given a Maximum Budget. For this category, all operations O1through O4still apply. However the first step before solving these operations is determining which subscriptions to get for the given budget. This leads to an additional operation: O5.What is the optimal set of subscriptions to get in a specific scenario? Depending on the circumstances and the consumer needs (e.g. the scenarios in the previous 4 operations), the optimal set will vary. Nonetheless, this set must always take the maximum budget into account. Example: let us assume that we want to send 50,000 RPM and have a maximum 78 5.4. ANALYSIS OPERATIONS budget of $100. At a first glance, we could simply get one subscription to Ultra, which is $79.95 and has a quota of 100,000 RPM. Nonetheless, we could potentially get many Basic subscriptions for $0, enough to allow our desired 50,000 RPM. However, let us assume that SendGrid limits Basic subscriptions to 1 per client. In that case, we could get 2 subscriptions to Pro, which would be $19.8 and have a total quota of 80,000 RPM. Given a Desired Capacity. Under this category, the analysis operations are the following: O6.What is the cheapest set of subscriptions that meets the desired capacity requirements? This operation would return the set of subscriptions with the lowest total cost that allows the consumer to fulfill their needs. Example: assuming that we want a capacity of 50,000 RPM, the example in O5would still apply. The difference is that there is no maximum budget in this operation. O7.What percentage of the desired capacity is met with a specific set of subscriptions? If the consumer already has a set of subscriptions or a budget constraint, they may want to know how much of the desired capacity is actually able to be fulfilled. Example: we want a capacity of 50,000 RPM, and, for any specific reason, we want to use one subscription to Pro, which has a quota of 40,000 RPM. In this scenario, we are meeting 80% of the desired capacity. Given a Desired Set of Features. This category contains the following operations: O8.What is the cheapest set of subscriptions that meets the desired feature requirements? Similar to O6, this operation would return the optimal set of subscriptions that meets the consumer needs, but focusing on features instead of capacity. Example: let us assume that we want to have phone support from SendGrid. Even if we just wanted to send 1,000 RPM, neither the Basic or Pro plans include phone support. Therefore, the cheapest option in this case is one subscription to the Ultra plan. O9.What amount of desired features is met with a specific set of subscriptions? Similar to O7, the consumer may want to know how many of the desired features can be fulfilled with a given set of subscriptions or a budget constraint. Example: if we want to have phone support and we are subscribed to the Pro plan, then we are meeting 0 of the desired features. Given a Desired Pricing. In a similar way to the previous two categories, this category 79 CHAPTER 5. EXTENDED PRICING MODEL includes two analysis operations: O10.What is the cheapest set of subscriptions that meets the desired pricing requirements? Similar to O6and O8, the consumer wants to know the cheapest subscriptions to get in order to meet the needs of their own customers. Example: let us assume that we want to offer a pricing with 10 RPS and 50,000 RPM. In this scenario, the cheapest option would be to get two subscriptions to Pro. However, if we want to offer this pricing to multiple simultaneous customers, we would need to multiply the limitations by the number of customers, e.g. for 10 customers we would need 100 RPS and 500,000 RPM. This would increase the required cost to subscribe to the SendGrid API. O11.What percentage of the desired pricing is met with a specific set of subscriptions? Again, similar to O7and O9, the consumer already has a set of subscriptions or a budget constraints and may want to know how much of their needs can be fulfilled. Example: if we want a pricing with 10 RPS and 50,000 RPM and we only get one subscription to Pro, then we are meeting 100% of the desired rate but only 80% of the quota. These percentages could be aggregated into one, depending on the consumer needs and preferences. Validity of a Pricing. The validity of a pricing is defined as the absence of conflicts between the different elements of the pricing [16]. Let us recall the basics of a pricing. A pricing consists of a series of plans, each of them having a series of limitations, and each limitation having a series of limits. A pricing is therefore valid when all of its elements are valid. For example, a pricing is not valid if a single limitation is not valid and has a conflict. An API provider would benefit from the automated validation of pricings, in order to ensure the coherence of its elements and, potentially, prevent users from exploiting the API. Another element of an API worth mentioning is its capacity. To sum up the definition in the previous section, the capacity of an API is the maximum workload that it can handle over a specific period of time. It usually depends on the internal architecture and the deployment setup of the API. For example, a capacity of 1000 requests per second (RPS) means that the API can serve up to 1000 requests each second to its users. In the following paragraphs, we present the validity criteria in a hierarchy. We start from fine-grained elements (limits and limitations) to coarse-grained ones (plans and pricing). Each criterion has multiple subcriteria, and all of them must be valid for the criterion 80 5.4. ANALYSIS OPERATIONS to be valid. VC1 - Valid limit Alimit is valid if its threshold is a natural number (VC1.1). VC2 - Valid limitation Alimitation is valid if: all its limits are valid (VC2.1); there are no limit consistency conflicts between any pair of its limits, i.e. a limit over a longer period of time has a lower threshold than a limit over a shorter period (VC2.2); there are no ambiguity conflicts between any pair of its limits, i.e. multiple limits with different thresholds over the same period of time (VC2.3); and there is no capacity conflict, i.e. a limit is less restrictive than the capacity of the API (VC2.4). VC3 - Valid plan Aplan is valid if: all its limitations are valid (VC3.1); and there are no limitation consistency conflicts between any pair of its limitations, i.e. a limitation over a metric allows another limitation over a related metric (by a certain factor) to be exceeded (VC3.2). VC4 - Valid pricing Apricing is valid if: all its plans are valid (VC4.1); and there are no cost consistency conflicts between any pair of its plans, i.e. a limitation in a plan with a lower cost is less restrictive than the equivalent limitation in a plan with a higher cost (VC4.2). Listing 5.5 shows a simplified example of a pricing with a limit consistency conflict (VC2.2). In this example, there is a conflict because there are two quotas defined over the same metric (requests) and the quota with the longest time unit (week) is more restrictive than the quota with the shortest time unit (day).   1Limitations : 2Quota : 100 requests / 1 day 3Quota : 10 requests / 1 week   Listing 5.5: Example of validity criterion VC2.2 (limit consistency conflict). Thus, the main operation under this category would be O12,Is my pricing valid?. 81 CHAPTER 5. EXTENDED PRICING MODEL 82 6 Limitation-Aware Microservice Architectures 83 This chapter introduces the concept of limitation-aware microservice architectures or LAMAs. First, Section §6.1 presents some background and the elements that differentiate a LAMA from a regular MSA. Next, Section §6.2 provides further details about the impedance mismatch problem that was originally introduced in Chapter §1. Then, Section §6.3 contains a description of the elements of a LAMA, as well as a domain-specific language for their formal definition. CHAPTER 6. LIMITATION-AWARE MICROSERVICE ARCHITECTURES 6.1 Introduction As discussed in Chapter §3, it is common for MSAs to consume external APIs offered by third party providers. In this scenario, where the developers of the MSA are aware of the limitations of these external APIs, we coin the term Limitation-Aware Microservice Architectures, or LAMA. In short, a LAMA is an MSA that consumes external APIs with limitations. The LAMA might offer its own plans to its customers in a SaaS scenario. Let us recall the example in Fig. §1.3. The scenario depicted in the figure is a LAMA, and, from now on, we will refer to it as such. The LAMA consumes two external APIs, each of them with two different plans and various limitations. Furthermore, the LAMA offers two plans to its customers, from which they are able to choose. To simplify calculations in the following chapters and sections, we use the same time units for all rates, quotas and billing periods. However, in real-world LAMAs that consume multiple external APIs, this is not always true, thus complicating the analysis of the aggregated impact of all limitations. Additionally, the internal services of the LAMA might also have their own limitations. They are usually derived from the deployment infrastructure of the services, and they may have a considerable impact on the overall ability of the LAMA to accept requests from its customers. For example, a service that is deployed in a small server (either locally or in the cloud) might have a relatively high response time. If many requests are sent in a short period of time, they may take too long to be fulfilled and therefore create a queue. Depending on the available memory, CPU or storage, these queued requests could potentially result in a disruption of the service, making the LAMA unavailable. Analyzing the internal limitations of a LAMA is an interesting topic, but it adds a whole new level of complexity to the capacity analysis of the LAMA, so we let it out of the scope of this thesis. 6.2 Impedance Mismatch As introduced in Section §1.2, potential impedance mismatch is derived from the confronted roles of providers and consumers in a prosumer scenario. On the one hand, as consumers, a business relies on third party providers with their pricing and limitations. On the other hand, as providers, they deliver services to their customers, for a price and with certain limits. However, the dynamics on those different realities are intrinsically different. As service consumers, they have to select amongst the potential providers that 84 6.2. IMPEDANCE MISMATCH evolve their offering (i.e. pricing and limitations) which one (or ones) is adequate from the perspective of their business model. Conversely, as service providers, they have to design and change their offering to be competitive and grow while satisfying their customers and augment its capacity if necessary. Specifically, in the middle of these two realities, the prosumers need to solve the impedance of pricing and limitations by taking decisions on which is the best offering that they will put on the market to search for customers, while selecting the best plan to use of each provider. In many scenarios, organizations delivering services are not actually doing it in the context of market but there will always be a set of operational requirements that will evolve. Those requirements specify the acceptable conditions of load in which the service is required to operate (e.g. how many requests per second, for a certain endpoint, should be accommodated). As a consequence, in this case, the same impedance problem is also apparent. The following subsections provide a deeper insight on the impedance mismatch in LAMAs. The capacity analysis problem will be tackled later in Chapter §7. Dynamic Impedance Mismatch The first problem that we identify is the dynamic impedance mismatch problem. In some cases, the LAMA developers notice that their demand within the period of a quota (e.g. a month) is either too high or too low, meaning that they are close to exceeding the quota limitation or wasting too many requests. However, they cannot do anything to solve this issue as the quota resets at a specific time (e.g. at the end of the month), so they are unable to change the subscribed plan until the next period. When a business subscribes to an API, they commonly choose a plan based on their quotas. In most cases, quotas reset at a certain time independently of when clients send their requests. For example, a quota may reset at 12 AM on the first day of each month, even if the LAMA sent all its requests the last day of the previous month. This means that the requests sent to the external APIs (the demand) do not always align with the periods of the quotas. This may create a misalignment between the billing lifecycle and the fluctuations in the demand. As an example, we will use Bluejay, a real-world IS that works as an auditor framework for agile software development and was presented in Section §1.2. It collects information from various external sources (GitHub, Pivotal Tracker, etc.) and draws multiple graphs in a dashboard that show the evolution of the developers’ adherence to a series of best 85 CHAPTER 6. LIMITATION-AWARE MICROSERVICE ARCHITECTURES practices, known as team practices. Bluejay has been used to audit multiple software engineering courses at the University of Seville and the University of California, Berkeley. In the context of academic experiments, there are two types of users: students, which work on their projects and trigger new updates; and teachers, which check the dashboards and do not typically do any changes. When a student updates their repository, either by pushing some code or updating an issue, this change triggers a new calculation in Bluejay. To check the adherence to the team practices, the system needs to check the status of the project in each one of the different external services used in the project. With that information, Bluejay then computes a series of metrics for each team practice and stores them in a database. These metrics are then used to create graphs that students can use to evaluate their progress, and are useful for teachers too. In terms of a LAMA, this means that each new change implies a request to the entrypoint of the LAMA, that then sends the appropriate requests to the external APIs. Therefore, the demand to these external APIs varies depending on the amount of changes that students do. Usually, these changes are mostly concentrated towards the end of a sprint or a deliverable, and are more spread out at other times. Consequently, there are periodic activity spikes at these times (e.g. weekly, every two weeks, monthly...). Furthermore, there may be an even higher spike when the course is near its end and students are finishing their projects. However, the quotas of the external APIs used to compute the metrics are reset monthly, independently of the students’ activity. For this reason, we need to monitor the demand to know when these spikes happen, and adapt the subscriptions accordingly. This task is not always easy, as we need to wait until the next quota period to be able to increase or decrease said quota. Thus, we need to anticipate to the demand spikes. Static Impedance Mismatch A second problem is the static impedance mismatch problem, which occurs when the initially expected demand to the external APIs does not match the actual demand. When designing a LAMA, the system architects or developers may obtain an initial calculation of the number of requests sent to each external API, that is, the demand to these APIs. Taking the API pricings into account, they may adapt the demand to make the most out of each subscribed API plan. Usually, this initial analysis assumes a worstcase scenario, that is, all requests are always sent and all users send their requests at the same time. Thus, the business may get the appropriate number of API keys to support 86 6.2. IMPEDANCE MISMATCH this maximum demand. In other cases, the analysis assumes an expected demand based on initial calculations of expected users and workload. This is a static demand analysis. Nonetheless, in the real world, the system load commonly fluctuates and is not always the same. The LAMA customers send their requests when they need to do so, each one of them at a different time, following a chaotic load. Furthermore, the specific functionality used by each customer may be different, thus following different paths within the internal services of the LAMA and, in turn, having a different demand to the external APIs. Additionally, it is possible to accidentally alter the demand by adding a new feature to the system or by releasing a new version which contains a bug that sends more requests than expected. These situations may lead to one of two different scenarios: (i) the system sends less requests than initially expected, thus wasting available requests and losing money by assuming a worst-case scenario; (ii) the system sends more requests than expected, so the external API limitations are exceeded and a service disruption occurs. Therefore, the initially calculated demand does not match the real demand. This is what we call the static impedance mismatch problem. Demand Analysis In both previous problems, it is necessary to obtain the actual demand of the LAMA at any given point in time so that the business can adapt accordingly. Consequently, a third problem is the demand analysis, that is, analyzing the current demand of a LAMA. The simplest way to perform this analysis is to monitor the internal services of the LAMA and check the requests that they are sending to each other and to the external APIs. To do this, the services need to be instrumentalized in order to collect information about their usage. By collecting traces of the various services, it is possible to infer the demand and even the internal topology of the LAMA, which may differ from the expected topology if an implementation bug exists (thus creating a static impedance mismatch). Besides using this information to adapt the subscribed plans as required, the business is also able to leverage the collected traces to modify or correct the internal topology if needed. Section §8 provides further information about the requirements of a monitoring framework for LAMAs, and shows a first implementation that can be used to automatically draw the topology of the LAMA and also answer multiple capacity analysis operations described in Section §7.4. 87 CHAPTER 7. CAPACITY ANALYSIS 7.1 Introduction In Chapter §4, we introduced the concept of capacity and the capacity analysis of a regular MSA, showcasing its necessity. In the following sections, we will ground these concepts within the scope of LAMAs, considering their particular features. Besides the definition of the capacity analysis problem, we provide further details about the elements that are involved in it. Furthermore, we now introduce a first approach to automate the capacity analysis of a LAMA. This approach is based on a transformational model, which converts the topology of a LAMA and the external API pricings into a constraint satisfaction and optimization problem (CSOP). Each element of the LAMA is transformed into a parameter within the CSOP, while the topology is represented as a set of constraints. Then, an analysis operation would be converted into additional constraints so that the solution of the CSOP provides the solution to the operation. As far as we know, this transformation has never been explored in the literature for the particular features of LAMAs. Furthermore, the declarative nature of a CSOP makes it easy to further expand the model if new elements are added to the LAMA, or if new operations are introduced. We also introduce a catalogue of three basic analysis operations, from which a wide variety of other operations can be defined. 7.2 Capacity of a LAMA The capacity of a LAMA refers to the maximum workload that it can handle over a given period of time and at a maximum cost, without exceeding any of the external limitations derived from subscribed pricing plans. This definition is in line with the capacity and performance management practice in ITIL 4 [64]. The capacity analysis of a LAMA should provide answers to the software architects and DevOps to make decisions over the subscribed external APIs and the potential operating conditions for the LAMA users. In particular, this analysis should take into account three dimensions that are intertwined: •Metrics. This dimension addresses the metrics (bounded to a scale) that have an impact on the capacity or are constrained by external APIs. In this thesis we focus on a single metric, number of requests, that is the most widely used metric in the industry [56] and is constrained and limited in most commercial API pricing plans. 94 7.2. CAPACITY OF A LAMA It is important to note that the metric should always be bounded to a particular scale. In the case of number of requests, we could have different time scales such as Requests Per Second (RPS) or Requests Per Hour (RPH). •Temporality. This dimension represents the temporal boundaries for the capacity to be analyzed. In this context, the same LAMA could have different capacities depending on the time period when it is calculated. These boundaries are typically linked with the desired operating conditions or, in case of a SaaS, the defined pricing plans. In a realistic setting, there could be scenarios where different external APIs have different plan periods and consequently, the capacity analysis should combine multiple temporal perspectives involved. •Cost. This dimension takes into account the derived costs from the infrastructure operation and the cost derived from the contracted plans with the different external APIs. In the example of Fig. §1.3, multiple options are possible, depending on the number of plans contracted; we assume that it is possible to contract multiple times a particular plan as this is the norm in the real API market. For example, given the LAMA in Fig. §1.3, the capacity can be analyzed by manual calculations. In 1 second, using the cheapest plans (Basic and Silver) and no overage cost, the capacity of the LAMA is 1 RPS, because 1 request to S1 results in 8 requests to E1, and one more request to S1 would result in 16 requests to E1, thus exceeding the limitation of the Basic plan. The cost is a fixed value, $9 in this case. In 2 seconds, the maximum number of requests allowed to E1 using the Basic plan is 30; therefore, the capacity is 3 RPS, resulting in 24 requests to E1 and 12 to E2. The cost, however, remains the same. When dealing with real-world architectures, the number of internal services and external APIs is considerably high, and thus there is a great number of plans and possible combinations. Additionally, when defining the pricing plans to be offered to the LAMA customers, it is fundamental to know the limitations derived from the usage by the external APIs together with its associated cost. In fact, these costs will be part of the operational costs of the LAMA, and are essential when analyzing the OpEx (Operational Expenditures) [65] for the desired operating conditions in general, and to have profitable pricing plans in the case of a SaaS LAMA. 95 CHAPTER 7. CAPACITY ANALYSIS 7.3 Automated Capacity Analysis Automated LAMA capacity analysis deals with extracting information from the model of a LAMA using automated mechanisms. Analyzing LAMA models is an error-prone and tedious task, and it is infeasible to do it manually with large-scale and complex LAMA models. In this thesis we propose a similar approach to that followed in other fields, i.e., to support the analysis process from a catalogue of analysis operations (analysis of feature models [66,67], service level agreements [68,69,70] and Business Process [71]). In this sense, all the analysis operations we have faced so far can be interpreted in terms of optimal search problems. Therefore, they can be solved through Search Based Software Engineering (SBSE) techniques, similarly to other cloud engineering problems [72]. We tackle this problem as a Constraint Satisfaction and Optimization Problem (CSOP), where, grosso modo, the search space corresponds to the set of tuples (Requests,T ime,Cost)that conform valid operating conditions of the LAMA. The objective function is defined on the variable that needs to be optimised in each case: requests, time or cost. Formal Description of LAMAs The primary objective of formalizing a LAMA is to establish a sound basis for the automated support. Following the formalization principles defined by Hofstede et al. [73], we follow a transformational style by translating the LAMA specification to a target domain suitable for the automated analysis (Primary Goal Principle). Specifically, we propose translating the specification to a CSOP that can be then analyzed using stateof-the-art constraint programming tools. A CSOP is defined as a 3-tuple (V,D,C)composed of a set of variables V, their domains Dand a number of constraints C. A solution for a CSOP is an assignment of values to the variables in Vfrom their domains in Dso that all the constraints in Care satisfied. Table §7.1 describes the mapping of a LAMA to a CSOP. Because of its complexity, we describe each transformation from the table in the following paragraphs. We recommend that readers get familiarized with this mapping. Table §7.2 summarizes the meaning of each of the abbreviations used in the mapping. •Positive number of requests. All internal services and external APIs must serve a positive number of requests. Therefore, all variables reqSiand reqEithat denote the 96 7.3. AUTOMATED CAPACITY ANALYSIS request served by internal services and external APIs respectively must be greater than or equal to 0. •Requests served by internal services. Each internal service in the LAMA Simust serve all requests sent to it by every other service Sj, denoted as reqSjSi. Thus, for each internal service there is a constraint reqSi=Pn j=1reqSjSi·reqSj. •Requests served by external APIs. Each external API Eimust serve all requests sent to it by the internal services Sj, denoted as reqSjEi. External APIs do not send requests between them. Thus, for each external API there is a constraint reqEi=Pn j=1reqSjEi·reqSj. Additionally, the total number of served requests is the sum of the requests sent to each plan below its limitations, limReqij, and the requests sent over the limitations, ovgReqij. This differentiation in two variables helps us obtain the number of overage requests more easily. Thus, for each plan Pij of external API Eithere is a constraint reqEi=Pn j=1limReqij +ovgReqij. Furthermore, no requests can be sent using a plan with no keys, so for each plan there is a constraint limReqij >0→keysij >0. Also, no overage requests can be sent if there are no requests below limitations, so for each plan there is another constraint ovgReqij >0→limReqij >0. •Quota of each pricing plan. The number of requests served by each external API Ei must not exceed any quota qij, defined over a time unit quij. Multiple keys keysij for each plan may be obtained. For each external API Eiand each of its respective plans Pij, there is a constraint limReqij <=keysij ·qij · ⌈time/quij⌉. •Rate of each pricing plan. The number of requests served by each external API Ei must not exceed any rate rij, defined over a time unit ruij. Note that rates need to account for the time unit of the quota, as the rate is reset at the beginning of each unit. Therefore, for each external API Eiand each of its respective plans Pij, there is a constraint limReqij −quij · ⌊time/quij⌋<=keysij ·rij · ⌈time mod quij /ruij⌉. If a plan has no no quota, the constraint is simplified to limReqi<=keysij ·rij · ⌈time/ruij⌉. •OpEx of each external API. The cost of each external API Eiis the sum of the subscriptions to each plan Pij plus overage costs. For each external API Ei, there is a constraint OpExi=Pn j=1keysij ·costij +ovgij ·ovgReqij. •Total OpEx. The total cost of the LAMA is the sum of the cost of each external API. There is a constraint OpEx =Pn i=1OpExi. 97 CHAPTER 7. CAPACITY ANALYSIS Table 7.1: LAMA to CSOP mapping. LAMA Services and APIs CSOP Mapping Services S∀Siin S,        V←V∪reqSi D←D∪domain(reqSi) C←C∪reqSi>=0 External E∀Eiin E,        V←V∪reqEi D←D∪domain(reqEi) C←C∪reqEi>=0 Pricings Ei→Pi Pricing Pi:Pi1,...,Pin ∀Eiin E ∀Pij in Pi,              V←V∪limReqij ∪ovgReqij ∪keysij D←D∪domain(limReqij )∪domain(ovgReqij ) C←C∪limReqij >0→keysij >0∪ ovgReqij >0→limReqij >0 ∀Eiin E,C←C∪reqEi=Pn j=1limReqij +ovgReqij LAMA Elements CSOP Mapping Entry Entry SiV←V∪reqL D←D∪domain(reqL) C←C∪reqSi>=reqL Consumption [API] Rltshp S1< reqS1Ei> Ei,..., Sn< reqSnEi> Ei ∀Eiin E,C←C∪reqEi=Pn j=1reqSjEi·reqSj [Service] Rltshp S1< reqS1Si> Si,..., Sn< reqSnSi> Si ∀Siin S,C←C∪reqSi=Pn j=1reqSjSi·reqSj Limitations Plan Pij rate: rij /ruij quota: qij /quij V←V∪time D←D∪domain(time) ∀Eiin E ∀Pij in Pi,C←C∪limReqij <=keysij ·qij · ⌈time/quij⌉ ∪ limReqij −quij · ⌊time/quij⌋<=keysij ·rij · ⌈time mod quij /ruij⌉ Individual cost Plan Pijcostij ovg: ovgij ∀Eiin E ∀Pij in Pi,C←C∪ovgij =0→ovgReqij =0 ∀Eiin E,        V←V∪OpExi D←D∪domain(OpExi) C←C∪OpExi=Pn j=1keysij ·costij +ovgReqij ·ovgij Total cost Plan Pijcostij ovg: ovgij V←V∪OpEx D←D∪domain(OpEx) OpEx =Pn i=1OpExi 98 7.3. AUTOMATED CAPACITY ANALYSIS Table 7.2: Parameters and variables glossary. Parameters Definition SiAn internal service of the LAMA. EiAn external API consumed by the LAMA. PiPricing offered by Ei Pij j-esim pricing plan of Pi. reqSiSjNumber of requests served by service Sjeach time Sjis invoked. reqSiEjNumber of requests served by external API Ejeach time Siis invoked. rij Rate of plan Pij. ruij Time unit of the rate of plan Pij . qij Quota of plan Pij. quij Time unit of the quota of plan Pij. costij Subscription cost of plan Pij. ovgij Overage cost of plan Pij. Variables Definition reqL Number of requests served by the LAMA, equivalent to the number of requests of the entrypoint. reqSiNumber of requests served by service Si. reqEiNumber of requests served by external API Ei. limReqij Number of requests served using plan Pij within its limitations. ovgReqij Number of requests served using plan Pij beyond is limitations. keysij Number of keys subscribed for plan Pij. time Time period. OpExiTotal cost of external API Ei. OpEx Total cost of the entire LAMA. 99 CHAPTER 7. CAPACITY ANALYSIS 7.4 Analysis Operations We propose a catalogue of three analysis operations that leverage the formal description of LAMAs as a CSOP to automatically extract helpful information. Analogous analysis operations have been defined in the context of the automated analysis of feature models [66], service level agreements [68,69,70] and in the area of MSAs [54]. We may remark that it is not our intention to propose an exhaustive set of analysis operations, as an unbounded number of operations could be potentially defined by adding or removing constraints and parameters. For the description of the operations as CSOPs, we will refer to the input specification of a LAMA Land a variable v. Additionally, we will use the following auxiliary operations: •map(L). This operation translates a LAMA specification Lto a CSP following the mapping described in the previous section. •minimize(CSP, v). This standard CSOP-based operation returns a solution for the input CSP (if any) with the minimum value of variable v. •maximize(CSP, v). Same that prior operation but with the maximum value of variable v. In what follows, we present three basic analysis operations, and, for the first operations identified in Section §4.3 (now replacing the term MSA with LAMA), and using the LAMA in Fig. §1.3, we provide an explanation of how it is mapped to a CSOP from the corresponding basic operation. Maximum number of requests. This operation returns the maximum number of requests that a LAMA Lis able to serve, over a specific time window tand for a maximum total cost c. This operation can be translated to a CSOP as follows: maxRequests(L,t,c)⇐⇒ maximize(map(L)∧time =t∧OpEx <=c,reqL) With this operation we can answer question Q2 (Assuming we have a Basic plan and a Gold plan already contracted what is the maximal RPM I can guarantee to all my 20 customers?) in Section §4.3 resulting in 5.6 RPS to each customer: Q2 ⇐⇒ maxRequests(L,60s,15)/20 =5.6 req 100 7.4. ANALYSIS OPERATIONS Similarly, question Q3 (Assuming we have a monthly budget limit of $120 in my LAMA, which is the maximum RPS to each of 20 customers?) is translated into the expression maxRequests(L,1s,120)/20, resulting in 1.35 (that is, 1) RPS to each customer: Q3 ⇐⇒ maxRequests(L,1s,120)/20 =1.35 req Minimum cost. This operation returns the minimum cost of the LAMA L, so that it can serve a minimum of RL requests over a time window t. From the result of this operation we can obtain the optimum (cheapest) plan combination (including the number of keys to be subscribed for each plan and possible overage requests). The translation of this operation to a CSOP is as follows: minCost(L,RL,t)⇐⇒ minimize(map(L)∧reqL =RL ∧time =t,OpEx) The question Q1 (What is the cheapest operational cost for my LAMA in order to offer 2 RPS to 20 customers?) is translated into minCost(L,2 ·20,1s), resulting in a total cost of $174: Q1 ⇐⇒ minCost(L,2 ·20,1s) =$174 Minimum time. This operation returns the minimum time that a LAMA Lneeds to serve at least RL requests, given a maximum total cost c. This operation can be translated to a CSOP as follows: minTime(L,RL,c)⇐⇒ minimize(map(L)∧reqL =RL ∧OpEx <=c,time) 101 CHAPTER 7. CAPACITY ANALYSIS 102 8 Monitoring Model 103 This chapter presents a monitoring framework to automatically collect traces and metrics from a LAMA and infer its topology. First, Section §8.1 gives an overview of the framework and its usefulness for the capacity analysis of a LAMA. Next, Section §8.2 introduces a list of requirements that need to be met by a framework to be useful for the capacity analysis of a LAMA. Then, Section §8.3 details the framework itself and its components. 9 Pricing Model 111 This chapter describes the validation of our extended pricing model to determine its expressiveness and usefulness. Section §9.1 provides an overview of the validation process. Section §9.2 presents a detailed analysis of the model in an extensive set of real-world APIs, including various metrics and statistics. Then, Section §9.3 introduces an automated tool for the validation of a pricing according to a set of validity criteria. CHAPTER 9. PRICING MODEL 9.1 Introduction In order to validate the usefulness of our Pricing4APIs model, and the SLA4OAI serialization in particular, we carried out two different experiments. First, we analyzed a representative set of real-world and publicly available APIs found in the industry to check whether their elements can be modeled using Pricing4APIs. We provide a detailed analysis including various metrics and statistics to showcase the most common elements in these pricings. Then, we manually modeled a subset of these realworld API pricings to validate if our serialization is actually able to cover all of their aspects. We concluded the experiment with a series of remarks and a clear description of the elements that, as of today, are not supported by Pricing4APIs or SLA4OAI. Next, we developed a tool to automatically validate an API pricing, that is, automatically answering the analysis operation O12 in Section §5.4. This tool is publicly available as a command-line tool and as a RESTful API, and can be used to validate any pricing modeled using SLA4OAI. We validated the tool with the subset of modeled APIs from the previous experiments. Initially, thanks to this tool, we found out some modeling mistakes that we made ourselves. When corrected, we observed that none of the validated pricings had any conflict, which is to be expected because real-world API providers try to carefully craft their pricings so that their customers do not have any issues or doubts when choosing between plans. 9.2 Expressiveness Analyzing API Limitations and Pricing For this analysis, we considered three different sources: (i) the previous work in GamezDiaz et al. [56] in which the authors analyzed a set of 69 APIs from two of the largest API directories; (ii) the work of Neumann et al. [75] in which the authors analysed a set of 500 APIs from the top most popular 4,000 websites in the Alexa ranking [76]; (iii) the 27 most popular APIs from RapidAPI1. We adapted and applied the process described in contribution [56] (i), screening the API repositories and applying the inclusion criterion described by the authors (which 1https://rapidapi.com/collection/popular-apis. Accessed on January 2023. Note that this list is regularly updated, and some APIs may be added, deprecated or removed. Some of the APIs included in our analysis are no longer available. 112 9.2. EXPRESSIVENESS includes more than 5,000 APIs with a last update in 2020). The result was the selection of one source: ProgrammableWeb. We extracted the most popular API categories (97th percentile, i.e., 14 categories selected, with more than 16,500 APIs). We filtered this dataset by removing duplicates (only one API per company was chosen at random). As a result, we had 2,966 potential APIs to study. Out of them, 30 APIs were selected. In contribution [75] (ii), the authors analyzed a set of attributes of 500 APIs by focusing on their general features such as their fit to REST best practices and design decisions rather than on their specific pricing aspects. Nevertheless, this dataset is interesting as a starting point for our analysis since it includes a variety of APIs and provides a comprehensive analysis of certain attributes. From this dataset, we filtered out any rows which were not RESTful APIs, leaving a subset of 499 unique APIs. First, we selected those APIs with a Payment Plan, as specified in a column in the dataset, obtaining 55 APIs, which represents the 11.02% of the total 499 APIs. We noted some errata in the classification of some APIs that we are very familiar with (e.g., GitHub was wrongly classified in the not having plans section). The reason behind these errors might be that the APIs were analysed some time ago, when they did not have plans at the time; alternatively, some APIs might have their pricing plans hidden within their documentation (e.g., we found that Yelp has an implicit VIP plan). This led us to analyze the rest of the dataset (444) manually to check whether the API still existed and whether it had API limitations. This analysis resulted in 162 APIs to be included (67 with pricing plans and 95 without but with API limitations, which represent 15.09% and 21.4% respectively of these 444 APIs originally classified as "not having plans"). Adding these to the first set of 55 APIs, led to 217 APIs to be analyzed. The list of most popular APIs from RapidAPI (iii) includes 27 different RESTful APIs. RapidAPI acts as a gateway to these APIs and provides simple pricing options. Some of these pricings differ from the pricings offered by the API providers in their official websites, which are often more complex. After a manual analysis of all of them, we found that 22 APIs had pricing plans and limitations, while 5 had no plans or limitations. Combining the 30 APIs extracted according to [56], the 217 from the dataset in [75] and the 22 from RapidAPI and removing duplicates left a dataset of 268 APIs – the Pricing4APIs dataset. Table §9.1 presents the overall picture of the analysis that was carried out. Out of more than 17 027 APIs, we manually modeled 54 of them, having a 90% confidence level and an 11% margin of error [77]. The full Pricing4APIs dataset, including details about their attributes, is available at [78] as part of Dataset D01. We analyzed 268 APIs in regard to two main types of attributes: limitations and 113 CHAPTER 9. PRICING MODEL Table 9.1: Main numbers of our Pricing4APIs dataset APIs from Gamez-Diaz (N=2966) 30 APIs from Neumann (N=217) 217 APIs from RapidAPI (N=27) 22 Total APIs after removing dups. 268 APIs having pricing 176 out of 268 Manually modeled APIs (N=266) 54 pricing. Both types include a wide range of other attributes, some supported by our model but others not. For example, our model does support overage costs (e.g., $0.1 per exceeded request), but it does not support complex metrics based upon HTTP protocolrelated aspects (i.e., headers, parameters, etc.). Although the Pricing4APIs dataset comprises 268 APIs, only 176 of them, the 65.7%, present a pricing or a plan. Consequently, the analysis of pricing is limited to this reduced dataset. Table §9.2 presents some results of the analysis. Table 9.2: Results of the analysis in real-world APIs N=268 N=176 Has limitations 95.9% 94.9% Has quotas 59.7% 72.2% Has rates 78.7% 69.9% Has quotas and rates 42.5% 46.6% Simple cost (e.g., monthly price) 60.8% 92.6% Has a pay-as-you-go cost model 9.3% 14.2% Includes overage cost 11.9% 18.2% Limitations analysis: Most APIs (95.9%) have limitations in terms of quotas (59.7%) or rates (78.7%). Almost half use a combination of the two (42.5%). These limitations are usually rather simple (such as monthly requests for quotas and secondly requests for rates). However, a minority tend to have a higher level of expressiveness. For example, they use the information from the HTTP request – from query parameters (2.2%) to other low-level aspects of the HTTP message such as headers, body, etc. (9.3%). A marginal number of APIs allow consumers to exceed the limitation value once or many times per month (1.1%). 114 9.2. EXPRESSIVENESS Pricing analysis: the vast majority of the APIs (92.6%) include a simple (e.g., monthly) cost. Nonetheless, they may have operation costs (14.2%) or include overage costs (18.2%). Finally, a minority have purchasable add-ons or extras (5.7%) or their pricing is calculated based on the number of users (10.2%). ProgrammableWeb (>16500) Neumann dataset (500) Without duplicates (2966) RESTful APIs (499) With pricing plans and limitations (55) Without pricing plans (444) With pricing plans and limitations (67) With limitations but no pricings (95) Pricing4APIs dataset (268) With pricing plans and limitations (176) Modeled APIs (54) Manually selected APIs (30) With pricing plans and limitations (22) RapidAPI Popular (27) Figure 9.1: Diagram of the database filtering process, starting from the Gamez-Diaz and Neumann datasets and the RapidAPI most popular list Given the 54 modeled APIs, we analyzed the different metrics included in the documentation of each of them. We found a total of 145 metrics, although different providers may name a same metric with different names (e.g. requests and transactions). 61.38% of these metrics are domain independent (such as requests, storage or users), while 38,62% are dependent (such as emails, documents or invoices). We grouped the 145 metrics in different categories based on their similarity, resulting in 14 categories. The most populated one is requests, including 58 metrics. The second one is AI (Artificial Intelligence), with 19 metrics, as some of the analyzed APIs are related to artificial intelligence and include a considerable amount of metrics. Many categories only include a few metrics 115 CHAPTER 9. PRICING MODEL because they are difficult to group together. This analysis is available at [78] as part of Dataset D02. Modeling API Pricings This subsection describes a validation of our Pricing4APIs model by modeling a number of real-world APIs, first describing the modeling process and then the issues that arise during this process. This process included the construction of a curated list of 54 API pricings with the model in Section §5.3, which represents the variability found in the industry. As noted above, we analyzed different attributes of the Pricing4APIs dataset of 268 APIs. The next step would be to write the SLA4OAI specification of every API so that it can be passed to the automated analyzer. However, since this is a time-consuming task, we decided to follow a hybrid sampling approach, using purposive and probabilistic sampling [79], to obtain a subset of APIs. With the former, we wanted to ensure that our model covers the most representative elements of API pricings, so we modeled all 22 APIs from RapidAPI’s most popular list. With the latter, we aimed to reduce the threats of purposive sampling by modeling 32 additional pricings. According to [79], purposive sampling is the most common approach in software engineering research, used in 76% of studies. The resulting subset includes 54 APIs. Note that the process of modeling a single API pricing consists of (i) reading and understanding the entire API documentation, (ii) extracting the API endpoints and methods (skipped if OAS documentation is available), (iii) reading and understanding every limitation in every plan of the API pricing, and (iv) specifying the metrics and API limitations in accordance with our proposed model for each API path and method. The process of modeling the API itself is tedious, which is why APIs having a public OAS documentation greatly facilitate the subsequent modeling task. With the introduction of globbing, step (ii) is simplified when the same limitation applies to multiple endpoints, and even completely unnecessary when the limitation applies to all endpoints. The 22 APIs from RapidAPI were modeled using globbing, which resulted in much simpler files. In the following sections, we determine the issues found during this OAS modeling process. In the process of modeling the pricings of this subset of 54 APIs, we encountered several issues. We have classified them into two categories: modeling issues and open issues, depending on whether they are issues that can be partially modeled with SLA4OAI or issues that need changes that will be taken into consideration when establishing future 116 9.2. EXPRESSIVENESS work. Modeling issues: MI-01 In pay-as-you-go plans, users are only charged with the requests that they actually consume (e.g., FacePlusPlus). This situation was modeled as a quota, with no max field (or max: unlimited) with its corresponding OverageCost. As an alternative, we could also have modeled this as an OperationCost. MI-02 In some APIs (e.g., FacePlusPlus), the operation cost depends on the HTTP status code that is returned to the consumer. Hence, the same request to the same endpoint might well be billed differently with regard to the status code (e.g., $0.01 if 200 OKs and $0.005 if 400 Bad Requests). We modeled this situation as a new metric for each status code. For example, in FacePlusPlus, the QPS metric has been split into QPS_OK, QPS_timeout and QPS_invalidParam. MI-03 If a certain plan explicitly denies access to certain API operations (e.g., Azure Search), those operations are not included in the model. MI-04 If the actual value for a quota or rate is unknown (e.g., Accuweather), we omit this rate/quota. For example, a number of APIs explicitly mention that they apply some rate-limiting value, but they do not mention what the actual value is. MI-05 Some metrics are dependent on some aspect of the HTTP request (body, parameters, etc.) and do not have any associated period (e.g., FacePlusPlus). In this case, the period property is removed. MI-06 There are also pricings with unknown cost (such as educational plans, non-profit organizations, enterprise, etc.). These are modeled with custom: true (e.g., GeoRanker). Additionally, if a limitation has a custom value to be negotiated with its provider, it is also modeled with custom: true (e.g., Yelp). MI-07 In plans whose billing depends on the number of users (e.g., Box) or on other variables affecting the cost (number of organizations, consumers, accounts, etc.), we considered the minimum number allowed. For example, plan Business Starter of Box requires a minimum of 3 users, and it becomes more expensive if more users join the plan. Therefore, for the sake of simplicity, we consider the cost of Business Starter to be the cost for 3 users. MI-08 Finally, in APIs whose documentation does not specify whether the time window in which limits are calculated is fixed or sliding, we assumed that limits with longer 117 CHAPTER 9. PRICING MODEL periods (e.g. years or months) use fixed windows, and limits with shorter periods (e.g. seconds or minutes) use sliding windows. This decision is based on the research in [56]. Open issues: OI-01 Some HTTP query parameters are limited to a certain range of allowed values instead of a maximum value (e.g., Scopus). Despite the fact that we modeled some parameters as a metric (e.g., number of results), parameters within a range were not modeled. In the Scopus case, Scopus Search limits the number of results to 25 in the nonsubscriber plan, whereas this number rises to 200 in the subscriber plan. Nevertheless, it also limits the parameter view to STANDARD in the former case and allows COMPLETE only in the latter. OI-02 Another open issue arises when the overage cost is also limited (e.g., Georanker). Some providers force one to move to another plan if one surpasses a certain value of the overage cost. This situation has not been modeled. For example, the small plan includes 300,000 requests, with an overage cost of $0.001 per request. However, this overage cost goes up to 750,000 requests. Once this amount is reached, one has to move to the medium plan. 9.3 Automation In this section, we focus on automating the validation of a pricing, that is operation O12 in Section §5.4. This operation, in order to be useful for practitioners, needs to be automated by means of a specific tool. To this end, we have developed SLA4OAIAnalyzer2, a publicly available command-line tool prototype [80]. Once installed, given a SLA4OAI file, the command sla4oai-analyzer -o <operation> -f <myFile.yaml> will initiate the validity analysis for this file. For example, for the validity operation, sla4oai-analyzer first checks the syntax validity according to the JSON Schema defined in the repository, and then checks each validity criterion in each part (pricing, plan, limitation, and limit). Fig. §9.3 depicts a consistency conflict detected by this tool, caused by a modeling mistake. As illustrations of some outputs of the tool, Fig. §9.2 shows a pricing with syntax errors and Fig. §9.3 a consistency conflict. 2https://github.com/isa-group/sla4oai-analyzer 118 9.3. AUTOMATION > sla4oai-analyzer -o validity -f ‘.\yelp.yaml’ ------ BEGIN CHECKING FILE: .\yelp.yaml ------ CHECKING SYNTAX... SYNTAX ERRORS in yelp.yaml SYNTAX ERROR: in path "#/": Missing required property: metrics ------ END CHECKING FILE: .\yelp.yaml ------ Figure 9.2: Tool running a syntax check > sla4oai-analyzer -o validity -f ‘.\inconsistent-ex.yaml’ ------ BEGIN CHECKING FILE: .\inconsistent-ex.yaml ------ CHECKING SYNTAX... SYNTAX OK CHECKING VALIDITY... USING DEFAULT CAPACITY LIMIT CONSISTENCY CONFLICT: in Plan1>/method1>get>requests ('60 per 60/second' and '1 per 1/second') VALIDITY ERROR ------ END CHECKING FILE: .\inconsistent-ex.yaml ------ Figure 9.3: Tool running the validity operation with errors Pricing Validation API The SLA4OAI-Analyzer is also available as an API at [81]. Its code is publicly available at [82]. The API takes a URL that points to an SLA4OAI specification file as an input and checks for all conflicts that were presented in the previous subsection. The API returns a Boolean value that indicates whether the file is valid, and a detailed log showing the validity of each individual element of the pricing. To validate the API, we provide an online Jupyter notebook that is available at [83]. Note that the notebook is shared with execute access, meaning that the cells can be executed but not modified. Nonetheless, they may be modified if the notebook is duplicated into your own Deepnote account. The notebook includes two examples for each validity subcriterion: one example has a conflict and the other does not. Because the full SLA4OAI specification files are complex and contain additional information that is not relevant for the detection of conflicts, each example includes a short pseudocode fragment, that shows the pricing elements that are actually relevant for the corresponding example. 119 CHAPTER 10. CAPACITY ANALYSIS OF A LAMA Reporter Registry Collector nPeriods nPages · nMembers GitHub GitLab Heroku Pivotal Tracker nScopes · nMetrics · nGuarantees Free (Auth) Enterprise $0/month $21/month 5000 req/h 15000 req/h Free $0/month 1000 req/min Free $0/month 4600 req/day Free $0/month No limit Figure 10.1: The LAMA of Bluejay, including its internal services and external APIs. Each API is depicted along its plans and limitations periods in 1 day (288): maxRequests(L,$0,1 day)/288 =0.73 projects What is the cheapest combination of plans to analyzed the projects in S2 with a resolution of 30 minutes? This operation is solved by using minCost for the number of projects in S1 (23) multiplied by the number of 30 minute periods in 1 day (48). Once the minimum cost is obtained, we can induce how many keys are needed for each API. This information can be directly included in the output of MiniZinc together with the total cost: minCost(L,23 ·48,1 day) = Unsatisfiable There is no combination of plans (keeping in mind the limitation of 10 keys per API) that is able to analyze all 23 projects each 30 minutes. 126 11 Monitoring Model 127 This chapter showcases the applicability and utility of the proposed monitoring framework. Section §11.1 provides an overview of the validation process. Section §11.2 explains the design of the validation experiment, including the description of a synthetic LAMA. Next, Section §11.3 introduces the experiment itself and the results of the process to automatically infer the topology of the LAMA, and also performs various analysis operations using Smart LAMA. CHAPTER 11. MONITORING MODEL 11.1 Introduction After stating the importance of the impedance mismatch problem and introducing our first implementation of a monitoring framework in Section §8.3, we now validate said framework with a sample LAMA. As our current implementation of the framework requires all services to be implemented using Node.js and also to be instrumented with the agent, we decided to use a synthetic LAMA for this experiment, based on Fig. §1.3. Nonetheless, any real-world LAMA using Node.js can be monitored using our framework. The first step towards performing this experiment is creating the synthetic LAMA and simulating the requests between the internal services and the external APIs. To make this task easier, we created a simple, generic API whose only purpose is sending requests to other APIs. Then, the LAMA is deployed using Docker and the traces are stored through a locally deployed instance of the collector. Next, we take the collected traces and clean their data to facilitate their analysis. Using these traces, we can automatically draw various diagrams to visually show the chain of requests and, more importantly, the topology of the LAMA. With this inferred topology, we are able to automatically obtain the LDL description of the LAMA and use it with the tool presented in the previous chapter. This demonstrates that the monitoring framework is useful for the capacity analysis of a LAMA and can be easily used together with other tools. 11.2 Experiment Design For the validation of the monitoring framework, we created the synthetic LAMA shown in Fig. §11.1, that is a simplified version of the one in Figure §1.3. The services are configured and deployed using a docker-compose file. The following fragment of code shows the configuration of one of the services, in particular, the Microservice 2. The Docker image named isagroup/api-requester is a custom-made RESTful API whose only purpose is sending requests to other services. The remaining lines show how to set up the name of the service within the LAMA, as well as the agent to collect traces and metrics.  1microservice -two: 2container -name: microservice - two 3image : isagroup / api - requester 4environment : 5RESTSENSE_SERVICE_NAME : microservice - two 128 11.2. EXPERIMENT DESIGN Figure 11.1: Simplified LAMA for the monitoring experiment 6RESTSENSE_COLLECTOR_URL : host . docker . internal :4317 7command: 8- /bin/sh 9- -c 10 - | 11 npm i @restsense / agent 12 node -r @restsense / agent / tracing index . js   Listing 11.1: Configuration for the deployment of Microservice 2 using docker-compose Microservices 2 and 3 need to send requests to the GitHub API. Note that, for the design of this experiment, the limitations are irrelevant as our only goal is monitoring the LAMA to obtain its topology. As long as Microservices 2 and 3 consume the GitHub APIs, it is enough. External APIs do not neet do be instrumented. To specify the number of requests between each service, the api-requester service has a RESTful API with a single endpoint. This endpoint receives information about how many requests will be sent and where. To start the chain of requests, we must send one request to the entrypoint (in our case, Microservice 1), describing the entire chain of requests between the other services and the external APIs. Then, Microservice 1 will send further requests to Microservices 2 and S with information about their corresponding requests. Finally, Microservices 2 and 3 will send their requests to the GitHub API. Fragment 11.2 shows the structure of a sample request for a service using api-requester.   1[{ 2" url ": " http :// microservice - two / api /v1/ request " , 3" method ": " POST ", 4" body ": { 5" url ": " ANOTHER_URL ", 6"body ": { "..." }, 7" headers ": { "..." } 8} 129 CHAPTER 11. MONITORING MODEL 9}]   Listing 11.2: Request example for a service using api-requester This way, any arbitrary LAMA can be easily described and simulated. Note that api-requester was designed for the only purpose of simulating requests between services and, as such, it is necessary to know the number of requests beforehand in order to define the chain of requests. This means that we already know the topology of the LAMA, so we do not really need to infer it. Nonetheless, this is only for experimentation and in the case of a real-world LAMA we would not now its topology. 11.3 Experiment Execution Having already defined the LAMA using docker-compose and api-requester, we can now obtain information about its topology using traces and metrics. As seen in the previous section, the services were already instrumented to include the monitoring agent as part of the docker-compose file. The collector was also running in the background (it can be included in the docker-compose file or deployed on its own) and, in order to extract the collected data from the MongoDB database, we can use MongoDB Compass1. Alternatively, the collector also provides some endpoints to extract the traces and metrics without manually querying the database. The collected data, extracted in CSV format, was then uploaded to an online Jupyter notebook for further processing. This notebook is available at [85], and it is shared with execute access. The notebook also contains examples of analysis operations using Smart LAMA. The analysis of the data in the notebook follows three steps: data extraction, exploratory analysis and capacity analysis. Data Extraction The data extraction process comprises two steps: data loading and preprocessing. The CSV files are loaded into the notebook using the pandas library2, and are stored in dataframes. Then, during the preprocessing step, the dataframes are cleaned through three actions: 1. Information that will not be used in the notebook is removed. This information may be useful for other types of analysis, though. 1https://www.mongodb.com/products/compass 2https://pandas.pydata.org/ 130 11.3. EXPERIMENT EXECUTION 2. Some columns are renamed for better readability, as they originally followed the OpenTelemetry naming conventions. 3. Columns that refer to the same concepts are merged, such as columns including information about response status codes and response text (e.g. 200 and OK). Exploratory Analysis This analysis includes the automated extraction of the topology of the LAMA. For this part, the notebook includes various functions and methods to draw diagrams to visualize the data. For example, there is a function to draw a cascade representation of a trace, which includes all of its spans and some information about them. Hovering the mouse over the spans will show this information. However, the most important diagram in this analysis is the one that shows the topology of the LAMA. The notebook includes an example that shows how to draw a DAG from the information of a trace. This DAG looks somewhat similar to the diagram in Fig. §11.1, and it does not include information about pricings, because they are not part of the topology itself. Fig. §11.2 shows the resulting diagram in the notebook. Figure 11.2: Inferred topology of the LAMA under analysis 131 CHAPTER 11. MONITORING MODEL Capacity Analysis Once we know the topology of the LAMA and the pricings of the external APIs, we can perform various capacity analysis operations. Although we did not need the GitHub pricing for the topology, we now need it for the capacity analysis of the LAMA. To solve these analysis operations, we used Smart LAMA, already presented in Section §10.2. However, Smart LAMA requires the LAMA to be described using LDL. For this reason, we included a function in the notebook that transforms the inferred topology to LDL, and we manually included the GitHub API pricing. We can automatically detect whether any service in the LAMA is an internal service or an external API by analyzing its IP address. Furthermore, we know which service is the entrypoint because its span happens first and has no parents. Using this information, along with the number of each edge, we can automatically generate the LDL of the LAMA. Listing 11.3 shows the resulting LDL of our example. Notice that the services and APIs are identified through their IP address and not actual names.   1Entry : 172 _19_0_1 2External services : [ ’140 _82_121_5 ’] 3Internal services : [ ’172 _19_0_4 ’ , ’172 _19_0_1 ’ , ’172 _19_0_2 ’ , ’172 _19_0_3 ’] 4Relationships : [{ ’from ’: ’172 _19_0_4 ’, ’to ’: ’172 _19_0_1 ’, ’value ’:1} , ...]   Listing 11.3: Automatically generated LDL of our example With the topology of the LAMA described using LDL and the external pricing manually added, we can now send requests to Smart LAMA to perform analysis operations. For example, some operations include: 1. What is the maximum number of requests that can be sent in 1 minute? Answer: 4,545 requests. 2. What is the maximum monthly quota we can offer to our customers without exceeding the limitations of the external APIs? Answer: 136,363 requests per month. 3. What is the minimum cost to send a certain number of requests? Answer: $0.00, because the GitHub API has no cost. 4. What is the minimum time to send 5 requests? Answer: 1 second, because the GitHub API does not specify any rate. 132 11.3. EXPERIMENT EXECUTION 5. What is the minimum time to send 300 requests with a maximum budget of $20? Answer: 1 second, because the GitHub API has no cost and the quota is not exceeded. While the analysis with the GitHub API may be somewhat limited (because it has no cost and no rates), it is enough to showcase that our monitoring framework is able to automatically infer the topology of the LAMA and then perform capacity analysis operations. Thanks to this automation, we can get the topology at any point during execution time, so we can be aware of unexpected changes in the topology that may produce service disruptions or unwanted charges. 133 CHAPTER 11. MONITORING MODEL 134 CONCLUSIONS Part V