HySOR: A Simulation Model for the Sharing of Risk in a Service Level Agreement-Aware Hybrid Cloud
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Seifert, Michael; Kuehnel, Stephan Article — Published Version HySOR: A Simulation Model for the Sharing of Risk in a Service Level Agreement-Aware Hybrid Cloud Business & Information Systems Engineering Provided in Cooperation with: Springer Nature Suggested Citation: Seifert, Michael; Kuehnel, Stephan (2024) : HySOR: A Simulation Model for the Sharing of Risk in a Service Level Agreement-Aware Hybrid Cloud, Business & Information Systems Engineering, ISSN 1867-0202, Springer Fachmedien Wiesbaden GmbH, Wiesbaden, Vol. 67, Iss. 4, pp. 495-510, https://doi.org/10.1007/s12599-024-00890-7 This Version is available at: https://hdl.handle.net/10419/330624 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
RESEARCH PAPER HySOR: A Simulation Model for the Sharing of Risk in a Service Level Agreement-Aware Hybrid Cloud Michael Seifert •Stephan Kuehnel Received: 3 July 2023 / Accepted: 17 April 2024 / Published online: 13 December 2024 The Author(s) 2024 Abstract More and more organizations are considering public cloud services for their business. Functional improvements, innovative features, and strategic factors are further driving demand. The adoption of public cloud services typically involves the integration into existing IT architectures and an established service structure that is ideally aligned with the non-functional requirements of the business processes to be supported. Hybrid cloud providers must be able to accommodate a variety of different public cloud providers while ensuring continuity of service or appropriate compensation prior to implementation. Existing literature focuses on the calculation and simulation of service availability, but less on service credit or business process outage costs of service compositions. In consequence, this paper presents a calculation and simulation model for the concept of ‘‘sharing of risk’’ in Service Level Agreement (SLA)-aware hybrid clouds (HySOR), focusing on the risk-sensitive simulation of the financial impact on hybrid cloud providers and customers. The model was implemented as an R-based application and evaluated with 12 leading experts in the field, yielding interesting implications for theory and practice. Keywords Risk simulation Cloud adoption Hybrid cloud Service level agreement 1 Introduction Public cloud usage has grown steadily in recent years and is forecast to continue to grow even further, with Software as a Service (SaaS) remaining the largest segment (Gartner Inc. 2024). As public cloud SaaS adoptions increase, their integration into existing IT architectures towards hybrid clouds must be properly managed (Sun et al. 2008). However, managing this integration poses a number of challenges for organizations, e.g., changes to contractual commitments, different definitions and formulations of Service Level Agreements (SLA), or the economic evaluation of SLA breaches (Seifert et al. 2023), to name just a few. The various SLAs of different cloud providers offer customers a contractual basis for assessing the service commitment (Aljoumah et al. 2015; Yuan et al. 2015). For this purpose, however, the continuity-relevant components must first be formalized (Seifert 2021) before being aggregated across the various vertical and horizontal integration patterns (Breiter and Naik 2013) and the several cloud architecture levels (Comuzzi et al. 2009) up to the hybrid cloud composition (Theilmann et al. 2010). An essential aspect of such hybrid cloud architectures is the one-to-many relationship between the (one) provider of public cloud SaaS and the (many) customers, including the provider of the hybrid cloud composition (Pan and Mitchell 2015; Seifert et al. 2023). The possibility of opportunistic behavior of several additional public cloud providers (Pan and Mitchell 2015) leads to increased risks in contractual commitment and partnership, which should be considered in the financial assessment of the desired hybrid cloud architecture (Seifert et al. 2023). Such an architecture often has to deal with different negotiability types of SLAs (Kritikos et al. 2016) and a resulting risk regarding the adequacy of the top-level SLA for the service composition Accepted after two revisions by O ´scar Pastor. M. Seifert (&)S. Kuehnel (&) Chair for Information Systems, esp. Business Information Management, Martin Luther University Halle-Wittenberg, Universitaetsring 3, 06108 Halle, Germany e-mail: [email protected] S. Kuehnel e-mail: [email protected] 123 Bus Inf Syst Eng 67(4):495–510 (2025) https://doi.org/10.1007/s12599-024-00890-7
(Comuzzi et al. 2013). This may further lead to risks in the execution of the business processes being supported. Assessing the financial risk of failure of external (cloud) services is already reflected in various risk calculations and simulations, as can be seen, for example, in Jiang et al. (2013), Yuan et al. (2015), and Wang and Franke (2020). However, existing models discussed in the literature lack the consideration of uncertainties that go beyond availability parameters, as described as a relevant issue by Franke et al. (2013) and Johnson et al. (2014). In addition, the consideration of business-critical SLA parameters in recent public cloud SLAs is considered a promising extension of risk calculation and simulation approaches (Seifert 2021). The focus of this paper is on the assessment of hybrid cloud service compositions prior to implementation, which is motivated by a functional or strategic decision (Seifert et al. 2023). The provider – whether internal or external – has to deal with the risk of newly added (‘‘incoming’’) public cloud services. The decisions on pricing or penalties for failing the top-level SLA must be made in a long-term partnership between the provider and the customer (Benlian et al. 2011; Goo et al. 2009), even if this leads to unacceptability or a ‘‘don’t do it’’ decision. In this context, this paper presents a new risk calculation and simulation model and examines the following research questions (RQ): •RQ1: What are the core elements and business-critical SLA parameters driving risk in hybrid cloud architectures? •RQ2: How can the impact of simulated IT outages of participating components be taken into account in the risk calculation of hybrid cloud architectures and how does this affect the expected costs for providers and customers? •RQ3: How can the financial impact for providers and customers be calculated and related to the availability from the service composition’s top-level SLA to provide decision support for the sharing of risk? The further structure of this paper is based on the Design Science Research (DSR) process of Peffers et al. (2007). Accordingly, we first present related work found within this multi-stage research project, addressing the ‘‘Objectives of a Solution’’ phase of DSR in Sect. 2. Next, the calculation model for the sharing of risk for SLA-aware hybrid clouds (HySOR) is presented in Sect. 3, which corresponds to the first part of the ‘‘Design & Development’’ phase. In addition, the simulation aspects are also presented here. The second part (with a stronger focus on development) is represented by the implementation of the HySOR model in an R-based application built on the ‘‘Shiny’’ library, along with the demonstration using a case study in Sect. 4. A survey of 12 leading experts in the field was conducted to obtain a summative evaluation, which is documented in Sect. 5and represents the ‘‘Evaluation’’ phase of DSR. A second part of this phase can be found in the subsequent Sect. 6, which contains implications for practice condensed from the follow-up expert interviews. The limitations of this work, as well as opportunities for further research, are presented in Sect. 7. The paper closes with a conclusion, answers to the research questions, and a presentation of the contributions to theory and practice in Sect. 8. 2 Theoretical Background and Related Work In this chapter, we draw on our own previous research (see Seifert et al. 2023), in which we conducted a comprehensive systematic literature review (SLR) of research on hybrid clouds resulting from the adoption of SaaS. The study identified six major challenges in dealing with hybrid clouds resulting from public cloud adoption, which we use as a basis for this conceptualization. It is assumed that further research is needed, especially in the area of modeling business-critical Qualities of Service (QoS), such as service commitment or service credit and their aggregation in hybrid cloud compositions (Seifert et al. 2023). Furthermore, the consideration of business and technological uncertainties as well as the changing commitment of the involved cloud providers serves as a research gap for the extension of the existing knowledge base (Seifert et al. 2023). For consistent application of the existing literature to the identified research gap, we conducted a second short SLR in preparation for this design cycle with a special focus on risk calculation and simulation of cloud architectures and service compositions. We included research with a focus on the following topics (inclusion criteria): (I) ‘‘service downtime’’, ‘‘service outage’’, and ‘‘service failure’’ to address availability simulation and aggregation. (II) ‘‘Service credit’’ and ‘‘service penalty’’ to reflect business process outages and for calculating the financial impact on the customer side as well. We deliberately do not consider the contractual distinction between penalty and credit, as we see both as a payment from the provider to the (subscription) customer. And (III), ‘‘cloud’’, to ensure that both the characteristics of the different depths of negotiability (risking the business continuity), and multiple cloud component providers involved in the composition are addressed. We excluded research with a focus on the following: (I) Optimization in terms of performance and cost (e.g., Gue ´rout et al. 2014; Mateo-Fornes et al. 2019), (II) dynamics and SLA (e.g., Al-Ghuwairi et al. 2016; Faniyi 123 496 M. Seifert, S. Kuehnel: HySOR: A Simulation Model for the Sharing of Risk..., Bus Inf Syst Eng 67(4):495–510 (2025)
et al. 2012), and (III) risk assessment based on transaction history (e.g., Hussain et al. 2010). Based on the results of the SLRs, the following subsections define and derive both terminology and requirements that are used to conceptualize the HySOR risk calculation and simulation model. The requirements are highlighted in italics at the end of each sub-chapter. 2.1 Cloud Service Components, Cloud Service Compositions, and the Hybrid Cloud Architecture When describing hybrid cloud architectures, a distinction must be made between individual components and different compositions of components involved in the provision of a cloud service. To this end, we build on the works of Comuzzi et al. (2009), Labidi et al. (2016), and Seifert and Kuehnel (2021), that use architecture modeling to design service compositions in order to formally map hybrid clouds and capture SLA-critical risk aspects (e.g., availability). The related business and technology perspectives are also promising for understanding the impact on economic risks (Comuzzi et al. 2013; Seifert and Kuehnel 2021). A service component is a single IT service that fulfills a specific function, is required for the execution of a business process, and has an SLA. Service components can occur in any cloud service or cloud deployment model. In contrast, a service composition can be described as a combination of several service components, whereby these are divided into horizontal and vertical components depending on the type of integration (Seifert et al. 2023). Horizontal integration is the interaction of IT services of the same type from different sources (Breiter and Naik 2013), such as two SaaS components, where one is obtained from an existing private cloud and one from a public cloud. In contrast, vertical integration means that one service depends on another or consumes it (Breiter and Naik 2013), such as when SaaS is built on infrastructure as a service (IaaS). To put it simply, service compositions refer to the various possible combinations of different clouds and can describe, for example, multi-clouds or the hybrid clouds relevant to this work. In addition, the simple use of public cloud services alongside the existing IT landscape usually already leads to hybrid clouds, whereby the individual components differ in terms of their negotiability (e.g., availability). Since our risk model is applicable to both types of cloud deployments, we use the term cloud service composition as a generic term. HySOR has to consider an architectural model consisting of (i) service components that may represent existing IT components used by an organization and (ii) incoming public cloud services that result in service compositions that support the business process. 2.2 Customer and Provider in Different Roles There are also organizational aspects of hybrid cloud architectures that need to be taken into account and can become a challenge, e.g., with regard to the different roles. ‘‘These challenges become even more complex when cloud providers take on the dual role of provider and customer, for example, when their own cloud service offerings build on the services of external cloud providers’’ (Seifert and Kuehnel 2021). Regardless of whether it is an internal or external cloud service provider, someone needs to be responsible for the top-level SLA and discuss the risk of business process outages with their customers (Seifert and Kuehnel 2021). Subscription is a very common pricing model for public cloud services (Mazrekaj et al. 2016). ‘‘With subscription pricing, users pay on a recurring basis to access software as an online service or to benefit from a service’’ (Mazrekaj et al. 2016). The role distinction in this model is crucial for determining who holds the subscription to the cloud (Zhang and Zhou 2009). Subscription in the open architecture of cloud computing consists, among other things, of a process and the roles (Zhang and Zhou 2009). Whoever subscribes to the cloud service is a contractual partner, which is also a decisive risk element for the calculation of penalties. HySOR has to consider (i) the hybrid cloud customer and (ii) the hybrid cloud provider as the top-level contracting parties, with (iii) one of them subscribing to each of the cloud services involved in the service composition. 2.3 Service Level Agreement and Operational Level Agreement Formalization A core facet of hybrid cloud architecture involves SLA and operational level agreement (OLA) aspects (Seifert et al. 2023). To quantify availability as a risk metric, we need to formalize SLA/OLA parameters, including their measurable and calculable parameters called QoS (Suakanto et al. 2012). Availability is one of the most important attributes of cloud service quality, and most popular public cloud services claim their availability promise (Baset 2012; Gulia and Sood 2013; Seifert 2021). Availability is often described not only by a number but also by different parameters. The categories of Yuan et al. (2015) describe this appropriately for our context. The first parameter is the measurement period, usually defined as one month. The service granularity defines which service scope is meant, and the time granularity is usually specified in the form of 1, 5, or 10 min. The coverage describes what must be running correctly and which services must be included. In 123 M. Seifert, S. Kuehnel: HySOR: A Simulation Model for the Sharing of Risk..., Bus Inf Syst Eng 67(4):495–510 (2025) 497
addition, exclusions of unavailability are commonly defined. Yuan et al. (2015) provide a useful formalization of the penalty function: a) ‘‘total charge ratio’’, b) ‘‘fixed value at different violation levels’’, or c) ‘‘downtime ratio’’. We can also see the public cloud penalty in the ‘‘service credit’’ category of Seifert (2021). Moreover, we can find the typical service penalty calculation for public cloud SaaS in Seifert (2021) as a combination of the methods from Yuan et al. (2015). An example is the software company SAP with the ratio of total charge and the ratio of downtime: ‘‘per 1% below availability (99.5) you get a credit of 2% of your monthly fee’’ (Seifert 2021). In addition, the maximum credit volume is given, which is also crucial for risk assessment. Another important aspect of the SLA is its negotiability (Comuzzi et al. 2013; Kritikos et al. 2016), which only has a secondary effect on the formalization, i.e., as a formulation of fixed parameters in public cloud SLAs (Seifert 2021). The negotiability of the SLA plays a decisive role in the description of the ‘‘sharing of risk’’ in Sect. 2.6. HySOR has to (i) simulate service continuity based on availability commitment parameters and (ii) enable integrated calculation of typical penalty functions from cloud SLAs so that these (iii) can be aggregated in the service composition. 2.4 Uncertainty as a Risk Driver Uncertainty has to be taken into account in hybrid cloud architectures (Johnson et al. 2014). In the category ‘‘uncertainty in technology and business’’, Seifert et al. (2023) distinguish three dimensions of risk or uncertainty in this context. First, tangible risks, such as availability (Paquette et al. 2010), can be represented as a probability distribution, e.g., from the QoS history (Johnson et al. 2014). Second, there are intangible risks when the business process depends on an external cloud element (Paquette et al. 2010). This may lead to opportunistic behavior by the public cloud provider due to a one-to-many relationship in the hybrid cloud context (Pan and Mitchell 2015). Third, uncertainty regarding the knowledge of the architecture supporting the cloud compositions and business processes (Franke et al. 2013; Johnson et al. 2014; Rockmann et al. 2014) is an intangible risk driver and must also be considered. HySOR has to consider (i) tangible risks, (ii) intangible risks arising from possible opportunistic behavior of the cloud component providers involved, and (iii) intangible risks arising from the business technology architecture. 2.5 Related Work on Risk Calculation and Simulation Within the SLRs, we found three related papers on risk calculation or simulation models based on architectures with strong relevance to our context. Yuan et al. (2015) motivate their approach with the lack of clarity in availability commitment and penalty for cloud consumers, and the business model for cloud providers to find the optimal penalty level. The tripartition of possible penalty functions, i.e., (I) ratio of total charge, (II) fixed value, and (III) ratio of downtime, seems promising for the development of risk calculation and simulation models. The financial impact of customer cost as provider price together with the impact of downtime appears reasonable. The important finding that the provider will reduce the penalty in order to compensate for higher availability requirements or claims leads us to the assumption made later in this paper that better risk sharing (e.g., more equally shared risks) is a gap in research to date. Due to the lack of consideration of the composition (and the composition provider), i.e., the combination of several components and their aggregation, this risk calculation is not adequate for our context with the above-mentioned concepts. Jiang et al. (2013) motivate a QoS-based risk approach combined with business-oriented target monitoring. The five-stage procedure for finding and parameterizing a suitable failure probability distribution confirms the necessity of modeling based on distributions. In particular, parameterization is an interesting aspect of risk sensitivity of assumed failure probabilities in order to test decision support in variants. The lack of focus on the provider’s revenue and loss of revenue (e.g., by focusing only on reducing usage) is an incentive for developing new methods. Wang and Franke (2020) present a model for analyzing IT service outages for individual organizations and supply chains. An important part of their model is the cost of business process downtime (i.e., economic impact on the customer side) in terms of three function types: constant, linear, and quadratic. Another useful aspect is the frequency of breakdowns/downtime, represented by a Poisson arrival model – as found in the literature. On this basis, the downtime duration is modeled with a lognormal distribution, as is often used when modeling downtimes. The lack of consideration of uncertainties arising from business and technology (e.g., support of business processes by IT, see Sect. 2.4) is a point that requires adjustment. Furthermore, the consideration of today’s public cloud penalty functions for compensation on the customer side and costs on the provider side can play an important role. HySOR needs to (i) consider different business process downtime cost functions and demonstrate the financial 123 498 M. Seifert, S. Kuehnel: HySOR: A Simulation Model for the Sharing of Risk..., Bus Inf Syst Eng 67(4):495–510 (2025)
implications for (ii) the hybrid cloud composition customer and (iii) the provider. 2.6 The Concepts ‘‘Sharing of Risk’’ and ‘‘Zone of Possible Agreements’’ Finally, our approach also differs from related work as we do not aim at an economic optimization model for cloud providers. For example, Wang and Franke (2020) state that their ‘‘intuition behind the model is that capital K can buy better hardware, thereby reducing the frequency of downtime’’. HySOR, however, does not primarily focus on investment opportunities to reduce downtime, as we assume a typical ‘‘take it or leave it’’ scenario for public cloud adoption, as described by Comuzzi et al. (2013). Combined with the assumption of a desired partnership between the hybrid cloud customer and provider, this leads to the need to appropriately judge the hybrid cloud architecture. The following model therefore focuses on the newly proposed concept ‘‘sharing of risk’’. This concept opens up a negotiation corridor that allows hybrid cloud customers and providers to account for tangible and intangible decision factors by varying risk parameters. The sharing of risk reflects the range of risky financial consequences for the provider and the customer of a cloud service composition. To this end, we rely on one of the most well-known descriptive negotiation concepts – Raiffa (1982)’s Zone of Possible Agreements (ZOPA) (see Fig. 1). With this concept, Raiffa (1982) represents a twoperson distributive bargaining problem that is bounded by the parties’ reservation prices (respectively their best cases/ alternatives) (Ahlert and Stra ¨ter 2016). If the buyer’s reservation price is greater than that of the seller, there is a ZOPA containing a possible agreement value (Ahlert and Stra ¨ter 2016; Raiffa 1982). If the reservation prices of both parties are the same, there is one possible point of agreement, whereby in this case the gains for both parties are zero. If the seller’s reservation price is higher than the buyer’s, there is no ZOPA (Ahlert and Stra ¨ter 2016). Applied to the context of this study, we imagine a situation in which a customer and a provider of a hybrid cloud service composition enter into negotiations about risky financial consequences (instead of prices). The situation is more complex than in Raiffa (1982)’s original model of price negotiation, as the risky financial consequences for both parties are influenced by various aspects. Two main aspects lead to increased risk and require sharing as part of a long-term partnership between hybrid cloud provider and customer. First, the non-negotiability of the incoming public cloud components involved in the composition increases the risk of an outage. Second, the underlying penalty functions of these components may be insufficient to compensate for the financial consequences in the worst case. For example, passing the subscription for a cloud component on to the provider could reduce the provider’s risk, while increasing the customer’s (financial) risk of process outages. However, this could be taken into account when negotiating the service credit or the pricing for the service composition between the hybrid cloud customer and the provider. HySOR must consider risk-sensitive simulation parameters to predict, evaluate, and compare variants of the desired hybrid cloud composition architecture. 3 The Risk Calculation and Simulation Model HySOR In the following, HySOR and its different elements are described concerning the calculation and the simulation model, which ensures the transparency and comprehensibility of our DSR project. Moreover, a detailed description of the artifact allows the fields of future research presented in Sect. 6and 7to be addressed in a targeted manner. For example, other scientists can exchange specific simulation components or integrate additional calculation modules. Fig. 1 Raiffa’s descriptive negotiation concept (illustration adapted from Ahlert and Stra ¨ter (2016) 123 M. Seifert, S. Kuehnel: HySOR: A Simulation Model for the Sharing of Risk..., Bus Inf Syst Eng 67(4):495–510 (2025) 499
3.1 Core Elements and Characteristics The core elements of HySOR are a service composition s consisting of service components ithat are involved and functionally necessary to support the business process and its execution. The model refers to mmonths of a contract duration M. The QoS of the participating service components is then simulated for all hours (n) of this contract duration and then aggregated to the QoS of the service composition. Core elements: I: nonempty finite set of service components i2I: a service component from the set of I sI: a service composition representing a subset of I (i.e., consisting of i= 1,2,3,...,q service components) m2f1,2,3;...;Mg: month (m) of a contract duration (M) of the service composition (s) n2f1;2;3;...;720 Mg: hour (n) of a contract duration (M) of the service composition (s) The basic characteristics of the service components include the subscription ownership, the monthly cost of the service, the availability per month guaranteed in the SLA, and the penalty amount for missing the SLA (see the listed service component characteristics below). The tangible risks are modeled using the risk parameters lambda for the Poisson arrival of the downtime and the expected value and variance of the lognormal distribution of the downtime duration. The specific characteristics of these parameters can either be derived through benchmarking and from historical data (e.g., QoS history, as described in Sect. 2.4) or must be estimated by experts. As also described in Sect. 2.4, insufficient knowledge of the business technology architecture and its relevance for business process support is an intangible risk driver that needs to be considered. Ambiguities in the dependency of the business process on a component are therefore taken into account as a risk factor and are included in the model as a security perception parameter. The service components also feature downtime simulation and monthly availability aggregation capabilities. Finally, there are different types of penalty functions that can be modeled for each component. Service component characteristics: oi: type of subscription ownership of the service component (i) ci: monthly costs for the service component (i) si: in the SLA guaranteed availability of the service component (i) per month pai: penalty amount for the service component (i) for missing the SLA ki: expected value and variance for the Poisson arrival of the downtime of a service component (i) li: expected value of the downtime duration of service component (i) ri: standard deviation of the downtime duration of service component (i) l¼ln l2 i ffiffiffiffiffiffiffiffiffiffi l2 iþr2 i p :lparameter for the lognormal distribution r2¼ln 1 þr2 i l2 i :r 2 parameter for the lognormal distribution idi¼0;if i dispensable 1;else : indispensability of the service component (i) for the business process rii2½0,1: risk factor regarding the indispensability of the service component (i) for the business process ei¼idið1riiÞ: security perception parameter for the indispensability of the service component (i) for the business process dai;n¼fðkiÞ: simulated downtime arrival of the service component (i) per hour (n) di;n¼fðdai;l;rÞ: simulated downtime duration of the service component (i) per hour (n) ai;m¼fðdi;mÞ: calculated availability of the service component (i) in percent per month (m) pi;m¼fðai;m;si;paiÞ: penalty function for the service component (i) for missing the SLA per month (m) The service composition has specific characteristics, such as the monthly cost for the customer, which is equivalent to the composition revenue for the provider. The service composition itself has a defined availability in the so-called top-level SLA, with which the aggregation of the availabilities of the components involved is later compared to measure SLA compliance. Analogous to the components, the composition has a penalty amount and a corresponding penalty function. Another essential dimension of the service composition characteristics is the modeling of business risk (and related costs). According to our concept of ‘‘sharing of risk’’, the risk of penalties on the provider’s side must be balanced with the risk of business process outages on the customer’s side, both of which are connected via the service composition and corresponding penalty and cost functions. The 123 500 M. Seifert, S. Kuehnel: HySOR: A Simulation Model for the Sharing of Risk..., Bus Inf Syst Eng 67(4):495–510 (2025)
business process outage cost function types for compositions are either constant, linear, or quadratic, as described in Sect. 2.5. Service composition characteristics: cs: monthly costs for the service composition (s) ss: in the (top-level) SLA guaranteed availability of the service composition (s) per month pas: penalty amount for the service composition (s) for missing the SLA ds;n¼fðdiÞ: aggregated downtime duration of the service composition (s) per hour (n) as;m¼fðdsÞ: calculated availability of the service composition (s) in percent per month (m) ps;m¼fðas;m;ss;pasÞ: penalty amount for service composition (s) for missing the SLA per month (m) btype n¼fðds;eiÞ: business process outage cost function type, depending on the unavailability of the service composition (s) per hour (n) bm¼fðbnÞ: business process outage costs caused by the unavailability of the service composition (s) per month (m) 3.2 Calculation and Simulation Procedure As already mentioned above, we assume a Poisson arrival process as the basis for the availability simulation of service components, in line with Wang and Franke (2020). The downtime duration is simulated separately for each service component using a lognormal distribution. Remark 1 We assume that service components having a dedicated SLA also have independent outages, as they each have differentiated architectures and resilience procedures (corresponding to the SLA offered). Downtime arrival per service component per hour: dai;n¼daki¼fnðÞ¼ðkÞn n!ek ; 8n2f1;2;3;...;720 Mg Downtime duration per service component per hour: di;n¼fda i;n ¼1 dairffiffiffiffiffiffi 2p pexp ln dai ðÞlðÞ 2 2r2 ! ; for dai;n[0; 8n2f1;2;3;...;720 Mg The aggregation of the service components into a service composition with regard to the downtime duration is determined by the component with the maximum downtime per hour (across all components). As part of the later availability aggregation, the values are then summed up for the month (with 720 hours per month). Remark 2 We assume that each month consists of 30 days of 24 hours each and that a standard contract year comprises 12 months. This is based on the assumption that business process outages have the same financial impact each month, as the associated costs are independent of the occurrence of downtime during the contract period. Downtime duration aggregation per service composition per hour: ds;n¼fd i;n ¼max ðdi;nÞ; 8i2f1;2;3;...;qg; 8n2f1;2;3;...;720 Mg The aggregation of the downtime duration per hour is also crucial for determining the amount of the business process outage costs. This is because two or more components can fail simultaneously, resulting in the same service composition outage. Remark 3 We have a limitation in the case where a downtime arrival coincides with the downtime duration of an earlier downtime arrival. Due to a moderate arrival rate per hour, we assume this to be negligible. The business process outage costs are calculated per hour of the contract duration and subsequently aggregated for each month. The penalties are calculated for each component and for the composition based on the respective availability aggregation. This paper illustrates the calculation of the monthly penalty. Availability aggregation per service component per month: ai;m¼fd i;n ¼720 Pdi;n 720 ; for 0\n\721 !m¼1;for 720\n\1441 !m ¼2;...;!m¼M Availability aggregation per service composition per month: as;m¼fd s;n ¼720 Pds;n 720 ; for 0\n\721 !m¼1;for 720 \n\1441 !m¼ 2;...;!m¼M Business process outage costs per hour: 123 M. Seifert, S. Kuehnel: HySOR: A Simulation Model for the Sharing of Risk..., Bus Inf Syst Eng 67(4):495–510 (2025) 501
bconstant da ¼fd s ðÞ¼ba;e i;if ds;n[0; blinear da ¼fd s ðÞ¼ds;nba;e i; bquadratic da ¼fd s ðÞ¼ds;n2ba;e i; 8n2f1;2;3;...;720 Mg Business process outage cost aggregation per month: bm¼fb da ðÞ¼ Xbda; for 0\n\721 !m¼1;for 720\n\1441 !m ¼2;...;!m¼M Penalty calculations per month: pfix k;m¼fa k;m;sk ¼ak;m\sk!pai;else !0; ppercentage k;m¼fa k;m;sk;ck ¼max 2ckdskak;me;ck ; pnone k;m¼fa k;m;sk;ck ¼0; k2fi;sg; 8m2f1,2,3;...;Mg Following the previous calculations at the levels of components, compositions, and business processes, the financial impact for the customer (totalcostscustÞand the provider (totalcostsprovÞof the service composition can be determined. This is done by distinguishing between fixed costs, which are incurred on a monthly basis regardless of the simulated availability of the components, and variable costs, which are composed of business process costs and/or the penalties for the service components and composition. Fixed costs per month: fixcostscust;m¼fc s;ci ðÞ¼csþX n i¼0 ci; 8oi2fcustomer public cloud subscriptiong; fixcostsprov;m¼fc s;ci ðÞ¼csþXn i¼0ci; 8oi2fprovider public cloud subscription;provider private cloud componentg; 8m2f1,2,3;...;Mg Remark 4 For reasons of comprehensibility, the fixed costs of the composition ðcsÞthat a customer pays to the provider were modeled in the provider’s calculation with the same variable, but with the mathematical sign reversed, i.e. as negative costs ðcsÞ Variable costs per month: varcostscust;m¼fp i;m;ps;m;bm ¼pi;mps;mþbm; 8oi2fcustomer public cloud subscription;provider private cloud componentg varcostsprov;m¼fp i;m;ps;m ¼pi;mþps;m; 8oi2fprovider public cloud subscriptiong; 8m2f1,2,3;...;Mg Remark 5 For reasons of comprehensibility, penalty payments for the service composition ðps;mÞthat a provider pays to the customer for missing the SLA were modeled in the customer’s calculation with the same variable, but with the mathematical sign reversed, i.e. as negative variable costs ðps;mÞ: Financial impact calculations per month: totalcostscust;m¼fixcostscust;mþvarcostscust;m; totalcostsprov;m¼fixcostsprov;mþvarcostsprov;m; 8m2f1,2,3;...;Mg The financial impact per month is the sum of the monthly variable and fixed costs. The implicit connection between the aggregated availability of the service composition and the financial impact on the customer and the provider serves as a basis for decision support regarding the sharing of risk. Using the fitted total cost graph for the customer and the provider, interesting considerations arise concerning the decision-making and negotiation options of both parties, as the following case study shows. 4 Development and Demonstration HySOR was implemented using the integrated development environment RStudio/2023.03.1, with which an R-based application was developed that builds on the library ‘‘Shiny’’. The modeling of the Poisson arrival of service outages was realized by the R function rpois, the lognormal distribution of the downtime duration by rlnorm. The input fields are distributed over three areas of the user interface, as shown in Fig. 2. In the first, the ‘‘Service Composition’’ panel (see Fig. 2[A]), the characteristics of the service composition are represented by business process variables and SLA parameters at the top level. In the second area, the ‘‘ \Type [Component’’ panel, the characteristics of each service component involved are parameterized, as shown in Fig. 2[B] using the example of a private cloud component. The four ‘‘Add Component’’ buttons (see Fig. 2[C]) are used to add an additional component to the composition as a third input area. The addition of a typical public cloud component was a requirement from an early test phase of the prototype and was implemented to strengthen the demonstrability of the model. SAP, Microsoft, and Salesforce were chosen as suitable examples for the expected case studies due to the 123 502 M. Seifert, S. Kuehnel: HySOR: A Simulation Model for the Sharing of Risk..., Bus Inf Syst Eng 67(4):495–510 (2025)
describe the planned unavailability of cloud services. In addition, the expert interviews revealed – contrary to remark 2 – that business process outage costs in practice vary greatly depending on the time frame. This is immediately apparent at the weekends, during which many companies’ business processes do not incur any outage costs. A mapping of different outage cost functions in relation to time frames can lead to further interesting results, especially when considering scheduled maintenance periods. A second interesting direction for further research is the development of suitable estimation methods and, in connection with this, the examination of historical data for modeling network, interface, or other unknown risks which could influence the availability of service compositions (e.g., depending on the number of service providers and service components involved). This could be factored into the simulation model without the experts having to use their intuition to make estimates. 8 Conclusion This paper presents a risk calculation and simulation model that was developed considering artifacts of the existing knowledge base in a rigorous DSR process iteration. The core elements of service components, service compositions, associated SLAs/OLAs as well as their owners (in different roles), supported business processes, and various risks were derived from the literature, defined, and conceptualized. The areas of service commitment and service credit, which were identified as financially relevant in connection with the SLAs, complement this to answer the first research question (RQ1). Using the service component and service composition characteristics of HySOR, as well as the corresponding functions, availability simulation and aggregation, business process outage calculation and aggregation, and penalty calculation and aggregation, the impact of IT downtimes of components was implemented in the simulation of customer and provider costs to answer the second research question (RQ2). The third research question (RQ3) is answered on the basis of the results of simulating and calculating the aggregated availability of the service composition and the resulting financial impact for the customer and the provider in the best, mean, and worst case. Finally, the visualization of the financial impact depending on the top-level availability in the form of bar charts, supplemented by linear trend lines, was considered helpful by the leading experts surveyed. The results of this work illustrate the relevance of the topic and confirm the different perceptions of transparency in hybrid cloud architectures. HySOR provides useful impulses for the negotiation of hybrid cloud service compositions in the context of the sharing of risk. The contribution of this paper is twofold. On the empirical side, we address a practically important topic with an effective artifact implemented. The evidence from the evaluation of leading experts provides promising directions for further research. On the conceptual side, the description of the calculation and simulation elements of the HySOR model/artifact can serve as an evaluated foundation for other researchers. Funding Open Access funding enabled and organized by Projekt DEAL. References Ahlert M, Stra ¨ter KF (2016) Refining Raiffa – aspiration adaptation within the zone of possible ag. Ger Econ Rev 17:298–315. https://doi.org/10.1111/geer.12096 Al-Ghuwairi A-R, Khalaf MN, Al-Yasen L, Salah Z, Alsarhan A, Baarah AH (2016) A dynamic model for automatic updating cloud computing SLA (DSLA). In: Proceedings of the international conference on internet of things and cloud computing. ACM, New York, pp 1–7. https://doi.org/10.1145/2896387. 2896442 Aljoumah E, Al-Mousawi F, Ahmad I, Al-Shammri M, Al-Jady Z (2015) SLA in cloud computing architectures: a comprehensive study. Int J Grid Distrib Comput 8:7–32. https://doi.org/10. 14257/ijgdc.2015.8.5.02 Baset SA (2012) Cloud SLAs: present and future. SIGOPS Oper Syst Rev 46:57–66. https://doi.org/10.1145/2331576.2331586 Benlian A, Koufaris M, Hess T (2011) Service quality in software-asa-service: developing the SaaS-Qual measure and examining its role in usage continuance. J Manag Inf Syst 28:85–126. https:// doi.org/10.2753/MIS0742-1222280303 Breiter G, Naik VK (2013) A framework for controlling and managing hybrid cloud service integration. In: 2013 IEEE international conference on cloud engineering, pp 217–224. https://doi.org/10.1109/IC2E.2013.48 Comuzzi M, Kotsokalis C, Rathfelder C, Theilmann W, Winkler U, Zacco G (2009) A framework for multi-level SLA management. In: Dan A, et al (eds): Service-oriented computing. Springer, Heidelberg, pp 187–196. https://doi.org/10.1007/978-3-64216132-2_18 Comuzzi M, Jacobs G, Grefen P (2013) Clearing the sky: understanding SLA elements in cloud computing. BETA publicatie: working papers vol. 412, Eindhoven University of Technology, Eindhoven, pp 1–25. https://research.tue.nl/en/publications/clear ing-the-sky-understanding-sla-elements-in-cloud-computing. Accessed 20 Jul 2024 Faniyi F, Bahsoon R, Theodoropoulos G (2012) A dynamic datadriven simulation approach for preventing service level agreement violations in cloud federation. Procedia Comput Sci 9:1167–1176. https://doi.org/10.1016/j.procs.2012.04.126 Franke U, Johnson P, Ko ¨nig J (2013) An architecture framework for enterprise IT service availability analysis. Softw Syst Model 13:1417–1445. https://doi.org/10.1007/s10270-012-0307-3 Goo J, Kishore R, Rao HR, Nam K (2009) The role of service level agreements in relational management of information technology 123 M. Seifert, S. Kuehnel: HySOR: A Simulation Model for the Sharing of Risk..., Bus Inf Syst Eng 67(4):495–510 (2025) 509
outsourcing: an empirical study. MIS Q 33:119–145. https://doi. org/10.2307/20650281 Gue ´rout T, Medjiah S, Da Costa G, Monteil T (2014) Quality of service modeling for green scheduling in clouds. Sustain Comput Inform Syst 4:225–240. https://doi.org/10.1016/j.suscom.2014. 08.006 Gulia P, Sood S (2013) Comparative analysis of present day clouds using service level agreements. Int J Comput Appl 71:1–8. https://doi.org/10.5120/12335-8603 Hussain O, Dong H, Singh J (2010) Semantic similarity model for risk assessment in forming cloud computing SLAs. In: Meersman R et al (eds) On the move to meaningful internet systems, OTM 2010. Springer, Heidelberg, pp 843–860. https://doi.org/10. 1007/978-3-642-16949-6_12 Iivari J, Hansen M, Haj-Bolouri A (2021) A proposal for minimum reusability evaluation of design principles. Eur J Inf Syst 30:286–303. https://doi.org/10.1080/0960085X.2020.1793697 Gartner Inc. (2024) Gartner Forecasts Worldwide Public Cloud EndUser Spending to Surpass $675 Billion in 2024. In: Gartner Newsroom, Information Technology, Press Release, https:// www.gartner.com/en/newsroom/press-releases/2024-05-20-gart ner-forecasts-worldwide-public-cloud-end-user-spending-to-sur pass-675-billion-in-2024. Accessed July 20, 2024. Jiang M, Byrne J, Molka K, Armstrong D, Djemame K, Kirkham T (2013) Cost and risk aware support for cloud SLAs. 2184–5042. https://doi.org/10.5220/0004377302070212 Johnson P, Ullberg J, Buschle M, Franke U, Shahzad K (2014) An architecture modeling framework for probabilistic prediction. Inf Syst E-Bus Manag 12:595–622. https://doi.org/10.1007/s10257014-0241-8 Kritikos K, Plexousakis D, Plebani P (2016) Semantic SLAs for services with Q-SLA. Procedia Comput Sci 97:24–33. https:// doi.org/10.1016/j.procs.2016.08.277 Labidi T, Mtibaa A, Brabra H (2016) CSLAOnto: a comprehensive ontological SLA model in cloud computing. J Data Semant 5:179–193. https://doi.org/10.1007/s13740-016-0070-7 Mateo-Fornes J, Solsona-Tehas F, Vilaplana-Mayoral J, TeixidoTorrelles I, Rius-Torrento J (2019) CART, a decision SLA model for SaaS providers to keep QoS regarding availability and performance. IEEE Access 7:38195–38204. https://doi.org/10. 1109/ACCESS.2019.2905870 Mazrekaj A, Shabani I, Sejdiu B (2016) Pricing schemes in cloud computing: an overview. International Journal Advance Computer Science Applications 7. https://doi.org/10.14569/IJACSA. 2016.070211 Pan W, Mitchell G (2015) Software as a service (SaaS) quality management and service level agreement. Infuture 26:225–234. https://doi.org/10.17234/INFUTURE.2015.26 Paquette S, Jaeger PT, Wilson SC (2010) Identifying the security risks associated with governmental use of cloud computing. Gov Inf Q 27:245–253. https://doi.org/10.1016/j.giq.2010.01.002 Peffers K, Tuunanen T, Rothenberger MA, Chatterjee S (2007) A design science research methodology for information systems research. J Manag Inf Syst 24:45–77. https://doi.org/10.2753/ MIS0742-1222240302 Raiffa H (1982) The Art and Science of Negotiation: How to resolve conflicts and get the best out of bargaining. Harvard Univ. Press, Cambridge Rockmann R, Weeger A, Gewald H (2014) Identifying organizational capabilities for the enterprise-wide usage of cloud computing. In: PACIS 2014 Proceedings. http://aisel.aisnet.org/pacis2014/355 Seifert M, Kuehnel S, Sackmann S (2023) Hybrid clouds arising from software as a service adoption: challenges, solutions, and future research directions. ACM Comput. Surv. Vol. 55, No. 11. Article 228:1–35. https://doi.org/10.1145/3570156 Seifert M (2021) Analysis of public cloud service level agreements - an evaluation of leading software as a service provider. In: Ku ¨hnel S, Sackmann S, Trang S (eds): Proceedings of the First International Workshop on Current Compliance Issues in Information Systems Research (CIISR’21), Co-located with the 16th International Conference on Wirtschaftsinformatik (WI’21), Online (initially located in Duisburg-Essen, Germany), March 9th, 2021. CEUR Workshop Proceedings 2966, pp 22–35. https://ceur-ws.org/Vol-2966/paper2.pdf Seifert M, Kuehnel S (2021) ‘‘HySLAC’’ - a conceptual model for service level agreement compliance in hybrid cloud architectures. In: Reussner RH, Koziolek A, Heinrich R (eds): INFORMATIK 2020, Lecture Notes in Informatics (LNI), Gesellschaft fu ¨r Informatik, Bonn 2021 205–218. https://doi. org/10.18420/inf2020_19 Suakanto S, Supangkat SH, Suhardi, Saragih R (2012) Performance measurement of cloud computing services. International Journal Cloud Computing Services Architecture 2 2 9 20. https://doi.org/ 10.5121/ijccsa.2012.2202 Sun W, Zhang X, Guo CJ, Sun P, Su H (2008) Software as a service: configuration and customization perspectives. In: 2008 IEEE congress on services part II, pp 18–25. https://doi.org/10.1109/ SERVICES-2.2008.29 Theilmann W, Happe J, Kotsokalis C, Edmonds A, Kearney K, Lambea J (2010) A reference architecture for multi-level SLA management. J Internet Eng 4:289–298. https://doi.org/10. 21256/zhaw-1757 Venable J, Pries-Heje J, Baskerville R (2016) FEDS: a framework for evaluation in design science research. Eur J Inf Syst 25:77–89. https://doi.org/10.1057/ejis.2014.36 Wang SS, Franke U (2020) Enterprise IT service downtime cost and risk transfer in a supply chain. Oper Manag Res 13:94–108. https://doi.org/10.1007/s12063-020-00148-x Yuan X, Li Y, Jia T, Liu T, Wu Z (2015) An analysis on availability commitment and penalty in cloud SLA. In: 2015 IEEE 39th annual computer software and applications conference, pp 914–919. https://doi.org/10.1109/COMPSAC.2015.39 Zhang L-J, Zhou Q (2009) CCOA: cloud computing open architecture. In: 2009 IEEE international conference on web services, pp 607–616. https://doi.org/10.1109/ICWS.2009.144 Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law. 123 510 M. Seifert, S. Kuehnel: HySOR: A Simulation Model for the Sharing of Risk..., Bus Inf Syst Eng 67(4):495–510 (2025)