Modelling and predicting enterprise-level cyber risks in the context of sparse data availability
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Zängerle, Daniel; Schiereck, Dirk Article — Published Version Modelling and predicting enterprise-level cyber risks in the context of sparse data availability The Geneva Papers on Risk and Insurance - Issues and Practice Provided in Cooperation with: Springer Nature Suggested Citation: Zängerle, Daniel; Schiereck, Dirk (2022) : Modelling and predicting enterpriselevel cyber risks in the context of sparse data availability, The Geneva Papers on Risk and Insurance - Issues and Practice, ISSN 1468-0440, Palgrave Macmillan, London, Vol. 48, Iss. 2, pp. 434-462, https://doi.org/10.1057/s41288-022-00282-6 This Version is available at: https://hdl.handle.net/10419/309580 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
Vol:.(1234567890) The Geneva Papers on Risk and Insurance - Issues and Practice (2023) 48:434–462 https://doi.org/10.1057/s41288-022-00282-6 Modelling andpredicting enterprise‑level cyber risks inthecontext ofsparse data availability DanielZängerle1 · DirkSchiereck1 Received: 1 April 2022 / Accepted: 1 December 2022 / Published online: 10 December 2022 © The Author(s) 2022 Abstract Despite growing attention to cyber risks in research and practice, quantitative cyber risk assessments remain limited, mainly due to a lack of reliable data. This analysis leverages sparse historical data to quantify the financial impact of cyber incidents at the enterprise level. For this purpose, an operational risk database—which has not been previously used in cyber research—was examined to model and predict the likelihood, severity and time dependence of a company’s cyber risk exposure. The proposed model can predict a negative time correlation, indicating that individual cyber exposure is increasing if no cyber loss has been reported in previous years, and vice versa. The results suggest that the probability of a cyber incident correlates with the subindustry, with the insurance sector being particularly exposed. The predicted financial losses from a cyber incident are less extreme than cited in recent investigations. The study confirms that cyber risks are heavy-tailed, jeopardising business operations and profitability. Keywords Cyber risk modelling· Cyber risk management· Cyber insurance· Vine copula· Sparse time series Introduction Cyber risks are one of the greatest threats of the twenty-first century (WEF 2021). Originally arising from the use of information technology (IT), cyber risks have since increased in both number and financial impact, especially due to rapidly progressing digitisation, worldwide interconnection and the introduction of new digital * Daniel Zängerle daniel.zaenger[email protected]mstadt.de Dirk Schiereck dirk.sc[email protected] 1 Department ofCorporate Finance, Technical University ofDarmstadt, Hochschulstr. 1, 64289Darmstadt, Germany
435 Modelling andpredicting enterprise‑level cyber risks in… products and services (Njegomir and Marović 2012; Rakes etal. 2012; Aldasoro etal. 2020). The cost of cyber incidents is estimated at more than USD 1 trillion (McAfee 2020) globally. Cyber incidents not only jeopardise private customers but also pose new challenges for companies and organisations (Njegomir and Marović 2012; Choudhry 2014; Bendovschi 2015; Wrede etal. 2018; Aldasoro etal. 2020). Despite the high awareness of cyber risk among corporate decision makers (Smidt and Botzen 2018) and insurance companies (Pooser etal. 2018), enterprise risk management (ERM) still neglects the associated risks, with some industries and firms even adopting a passive stance (Ashby etal. 2018; Pooser etal. 2018). Effective cyber risk management should be comprehensively incorporated into ERM rather than analysed in an isolated manner, such as exclusively in IT departments (Marotta and McShane 2018; Shetty etal. 2018; Poyraz etal. 2020). Furthermore, there is evidence that cyber risk management processes are generally qualitative and are missing quantitative findings (Palsson etal. 2020). The usual method to quantify cyber risk is through an analysis of historical cyber incidents from verifiable sources and the performance of empirical, statistical and actuarial examinations to determine the financial impact and likelihood of a cyber incident in a specific organisation (Smidt and Botzen2018; Palsson etal. 2020). However, the lack of data restrains the quality of such assessments and constitutes the main research gap in the cyber risk literature (Eling and Schnell 2016; Marotta etal. 2017; Boyer 2020). To address this, we quantitatively assess the financial impact of cyber risks at the enterprise level using sparse historical data. Our analysis is based on the Öffentliche Schadenfälle OpRisk (ÖffSchOR) database—an operational risk (OpRisk) database on publicly disclosed loss events in the European financial sector—which has not been adapted to cyber risk research. We apply advanced modelling techniques suggested by Shi and Yang (2018), Eling and Wirfs (2019) and Fang etal. (2021) to predict the likelihood and loss exposure of a potential cyber incident. We specifically use statistical dependence, modelled by a D-vine copula structure, to cope with the sparsity of events in a multivariate time series setting. In doing so, we provide new empirical evidence and quantitative results on actual cyber risk losses at the company level. Our findings suggest that cyber risks are less severe than recent studies claim and that subindustries must be separately modelled. Additionally, our results support the insight that cyber risks are heavy tailed, with an extreme cyber incident as a worst-case scenario that would seriously harm or default a company (Eling and Wirfs 2019; Wheatley etal. 2021). The results of this study provide one of the first quantitative insights on the nature of cyber risks and introduce a new dataset to cyber research. The outlined methodology allows researchers and practitioners, in particular cyber insurers, to assess cyber risks despite the lack of larger datasets and to combine with existing pricing tools in order to evaluate risk-based premiums (Nurse etal. 2020; Cremer etal. 2022). Our study, thus, contributes to the limited research available on the empirical quantification of cyber risks and to abetter understanding within the field. The remainder of this paper is structured as follows. The next section provides a summary of the most relevant literature. Then, we introduce the dataset and methodology. The fourth section presents the results of our analysis. The final section
436 D.Zängerle, D.Schiereck concludes with a discussion of the findings and limitations of the study as well as future research possibilities. Literature review Compared to the prevailing research on operational risk modelling (see e.g. Cox 2012; MacKenzie 2014), cyber risk analyses are still very limited (Eling 2020).1 This lack of research is often linked to the limited availability of cyber loss data (Maillart and Sornette 2010; Biener etal. 2015), which is typically not disclosed by organisations in an effort to avoid reputational damage (Giudici and Raffinetti 2020). Despite several public and private initiatives to form databases (see the next section), companies have little incentive to share loss information in a public or consortium repository (Palsson etal. 2020). Initiatives such as the introduction of new reporting requirements for cyber incidents and data breaches—in the U.S. by the National Conference of State Legislatures (NCSL 2016) and in Europe by the European Union (EU 2016)—might improve modelling techniques (Eling and Wirfs 2019) but are still incapable of delivering new insights. Further, the introduction of individual cyber risk definitions leads to a maze of terms rather than a comprehensive and unified terminology and understanding of cyber risks (Zängerle and Schiereck 2022). The lack of cyber risk data has also been addressed in recent publications. In particular, Cremer et al. (2022) conduct a comprehensive and systematic review of cyber data availability, identifying only 79 datasets from a preliminary 5,219 peer-reviewed cyber studies. Furthermore, most of these databases focus on technical cybersecurity aspects, such as intrusion detection and machine learning, with only a fraction of available datasets on cyber risks. The authors find that the lack of available data on cyber risks is a serious problem for stakeholders that undermines collective efforts to better manage these risks. This interpretation is supported by Romanosky etal. (2019), who show that (cyber) insurers in the U.S. have no historic or credible data to assess the loss expectation of cyber insurance coverages. Due to the scarcity of cyber loss information, data breaches, mainly in the U.S., have received the most attention in empirical research (see e.g. Maillart and Sornette 2010; Edwards etal. 2016; Wheatley etal. 2016; Eling and Loperfido 2017; Eling 2018; Xu etal. 2018; Wheatley etal.2021). Attempts have also been made to assess the monetary impact of data breaches (see e.g. Layton and Watters 2014; Romanosky 2016; Ruan 2017; Poyraz etal. 2020). However, as addressed by Woods and Böhme (2021), these studies have produced contradictory results, depending on the dataset and methodology applied. Furthermore, Eling and Wirfs (2019) found that data breaches account for just 25% of all cyber events, and the estimated distribution of breached records does not align with that of the actual financial cost of cyber 1 For a comprehensive review and the status quo of cyber risk research, we refer to Eling and Schnell (2016), Marotta etal. (2017), Eling (2020) and Woods and Böhme (2021).
437 Modelling andpredicting enterprise‑level cyber risks in… incidents. To this day, only a few studies have assessed the financial impact of cyber incidents in a comprehensive way. Biener etal. (2015) analyse cyber losses from the SAS operational loss database and emphasise the distinct characteristics of cyber risks, including the lack of data, information asymmetries and highly interrelated losses. However, the authors focus on insurability rather than the modelling and prediction of cyber losses. Romanosky (2016) provides the first quantitative insights from actual loss information based on the Advisen dataset but concentrates on descriptive statistics.2 Later, Palsson etal. (2020) use the same database to model the financial cost of different cyber event types by applying a random forest algorithm. Although the data are not sufficiently detailed to construct a predictive model with high accuracy, the researchers identify relevant factors affecting the expenses of such incidents. Similar to our examination, Eling and Wirfs (2019) analyse the actual costs of cyber incidents from the SAS loss database with statistical and actuarial methods. By applying the peaksover-threshold (POT) method from extreme value theory (EVT), they find that cyber risks are distinct from other risk categories and argue that researchers must distinguish between ‘cyber risks of daily life’ and ‘extreme cyber risks’. In addition, they present a simulation study for practical application. We apply techniques similar to those of Eling and Wirfs (2019), who focus on monthly aggregated observations from all entities available, treated as one sample from a single distribution. We, however, utilise enterprise-level sparse time series data from a database that has not yet been used in the context of cyber risk modelling. A second research stream focusing on the modelling of dependence structures has recently emerged (Eling 2020). In particular, the application of copula theory is widely accepted due to the ability to use any marginal distribution, which is essential for diverse cyber risk classes, and to address non-linear dependencies (see e.g. Böhme and Kataria 2006; Herath and Herath 2011; Mukhopadhyay etal. 2013). Further studies have extended these approaches to multivariate settings using vine copulas (see e.g. Joe 1997; Bedford and Cooke 2002; Kurowicka and Cooke 2006; Aas etal. 2009), which generate a multivariate copula based on iterative and bivariate pairwise copula constructions (PCC). The D-vine, a distinct vine copula, is particularly structured and simple to interpret in the time series context (Zhao etal. 2020). For example, Peng etal. (2016) use honeypot data to model multivariate and extreme cyber risks with marked point processes and vine copulas, later progressing with a vine copula GARCH model (Peng etal. 2018). Shi and Yang (2018) analyse the temporal dependence in longitudinal data by a D-vine copula. Xu etal. (2018) model the interarrival times of data breaches by ARMA-GARCH and joint density with copula. Eling and Jung (2018) apply the Privacy Rights Clearinghouse (PRC) dataset and model the cross-sectional dependence of data breaches. They find that vine structures exhibit a better fit than simple elliptical or Archimedean copulas. Fang etal. (2021) also study the same dataset, but in a multivariate time series setting with sparse observations at the enterprise level. Therefore, they propose 2 The author uses a logistic regression model to analyse the actual costs of data breaches only. A new approach for assessing the monetary impact of mega data breaches is suggested by Poyraz etal. (2020).
438 D.Zängerle, D.Schiereck a D-vine copula to model the serial trend. We adopt this framework to model the financial impact of actual cyber incidents rather than data breaches alone. The current emergence of network models also offers a new, more appealing path for cyber risk modelling (see e.g. Fahrenwaldt etal. 2018; Jevtić and Lanchier 2020; Wu etal. 2021). However, these advanced predictive models are currently limited to simulation studies, as applying such methods to real-world data requires a vast amount of unfiltered data points in order to provide accurate predictions (Tavabi etal. 2020). These techniques are consequently not applicable to our setting due to the lack of sufficient data. Data andmethodology In addition to the fact that information on cyber risks is typically not publicly available, the systematic collection of known cyber incidents poses further challenges (Eling and Wirfs 2016b). As Romanosky (2016) illustrates, only a fraction of actual cyber incidents is recorded in associated loss databases. A limited number of cyber databases (see Table1) do exist, mainly established by private and public companies and consortia. Nevertheless, it is challenging to gain access to them, and there is no standard practice in the recording and collection of cyber incidents. OpRisk databases from the U.S. have been primarily used to model cyber risks in the existing literature, including Advisen (see e.g. Romanosky 2016; Kesan and Zhang 2019; McShane and Nguyen 2020; Palsson etal. 2020) and the SAS OpRisk database (see e.g. Biener et al. 2015; Eling and Wirfs 2016a, 2019). Furthermore, other organisations and consortia collaborate and share data on operational and cyber risks to build systematic databases. Specific databases focusing on data breaches in the U.S. (e.g. Privacy Rights Clearinghouse) and private initiatives (e.g. Hackmaggedon) have emerged. However, only some of the above-mentioned initiatives provide information on the economic loss of reported cyber incidents. For our analyses, we use the German Öffentliche Schadenfälle OpRisk database due to the following reasons. First, the database is rather small, which emphasises the introduced motivation of sparse cyber risk modelling. Second, ÖffSchOR focuses on OpRisk losses from the financial sector in Europe, providing some of the first insights both from this important industry in the European Union and from Europe overall. In particular, the size of the recorded losses and relative number of cyber incidents are comparable to previous studies (see e.g. Eling and Wirfs 2019). Third, ÖffSchOR provided free access to the database to conduct this research project and to promote quantitative cyber research. Fourth, and to the best of our knowledge, this is the first scientific analysis based on ÖffSchOR in the context of cyber risk research.3 3 A few other studies have previously used ÖffSchOR, mainly in OpRisk research (see e.g. Sturm 2013; Kaspereit etal. 2017; Eckert etal. 2020).
439 Modelling andpredicting enterprise‑level cyber risks in… Table 1 Overview of relevant cyber risk databases References from OpRisk research are not included in this overview Database Scope Focus Sample Selected references from cyber risk literature A. Commercial Advisen OpRisk U.S > 90,000 Romanosky (2016), Kesan and Zhang (2019), McShane and Nguyen (2020), Palsson etal. (2020) IBM FIRST Risk Case Studies OpRisk Global > 17,000 ./. Öffentliche Schadenfälle OpRisk (ÖffSchOR) OpRisk DACH > 3,300 ./. SAS OpRisk OpRisk U.S > 32,000 Biener etal. (2015), Eling and Wirfs (2016a), Eling and Wirfs (2019) B. Consortium Associazione Italiana per la Sicurezza Informatica (Clusit) Cyber Global > 6,800 Giudici and Raffinetti (2020) Datenkonsortium OpRisk (DakOR) OpRisk Germany > 38.000 ./. Italian Database of Operational Losses (DIPO) OpRisk Italy ./. ./. Operational Riskdata eXchange Association (ORX) OpRisk Global > 800,000 Bouveret (2018) C. Public Hackmaggedon Cyber Global > 10,000 ./. Privacy Rights Clearinghouse (PRC) Data breaches U.S > 8,500 Edwards etal. (2016), Wheatley etal. (2016), Eling and Loperfido (2017), Eling and Jung (2018), Xu etal. (2018), Eling and Wirfs (2019), Jung (2019), Fang etal. (2021), Kamiya etal. (2021), Wheatley etal. (2021) Vocabulary for Event Recording and Incident Sharing (VERIS) Data breaches U.S > 8,000 ./.
440 D.Zängerle, D.Schiereck ÖffSchOR database ÖffSchOR is an information database on publicly disclosed loss events of operational risks in the financial sector. The database is operated by VÖB-Service GmbH, a subsidiary of the Federal Association of Public Banks (Bundesverband Öffentlicher Banken Deutschlands, VÖB) in Germany. In general, losses of a gross amount of EUR100,000 or more are recorded in the database, including reputational risks and risk scenarios. The industry focus is on financial services and insurance companies in Europe. In addition, interesting loss events can be examined from other economic sectors or regions. ÖffSchOR uses print and online media services to collect data. All loss events are categorised according to the Capital Requirements Regulation (CRR) specifications (EU 2013). Loss incidents are assigned to different subcategories, such as conduct risk, legal risk, information and communication technology (ICT) risk or sustainability risk. To date, however, there is no unique identifier for cyber risk in the ÖffSchOR database. Therefore, subcategories distinguishing cyber and non-cyber events are necessary. Methodology Motivated by the framework of Fang etal. (2021), the methodology of this study is organised into six key components: (1) data preparation, analysis and transformation; (2) marginal model; (3) modelling frequency; (4) modelling severity; (5) modelling temporal dependence and (6) predicting the next time period. Data preparation, explorative data analysis anddata transformation As of 30 September 2021, the ÖffSchOR database consists of 3,261 operational loss events between 2002 and 2021. Given that the database does not categorise cyber events, it is first necessary to allocate the sample to cyber and non-cyber incidents. Cyber risk is defined as “any risk emerging from the use of ICT that compromises the confidentiality, availability, or integrity of data or services […]. Cyber risk is either caused by natural disasters or is man-made where the latter may emerge from human failure, cyber criminality (e.g. extortion, fraud), cyber war or cyber terrorism” (Eling etal. 2016). Based on this definition, which has been suggested as the most comprehensive in the cyber risk literature (Strupczewski 2021), Tables2 and 3 present the search strategy employed to identify 341 cyber events in the ÖffSchOR database. The strategy combines both systematic and manual search steps to maximise and validate the categorisation of cyber events. In order to gain preliminary insights from the data, a descriptive analysis is conducted. The dataset is then transformed into a time series, where yit is the amount of all cyber losses of company i in year t, n is the number of companies in the data and T is the time horizon. Non-cyber incidents and companies without a single cyber
441 Modelling andpredicting enterprise‑level cyber risks in… Table 2 Search and identification of cyber events in the ÖffSchOR database and remaining data points (in bold) Values in italics used to highlight the difference or add-on of data points Step Task Data points 0 Extraction of ÖffSchOR database as of 30 September 2021 3,261 1 Systematic and manual search of cyber events according to the definition of Eling etal. (2016) by identifying cyber keywords in the event description, according to Table3(−2898) 363 2 Manual review of all tagged cyber incidents in terms of validity and consistency including recategorisation, if necessary (−38) 325 3 Manual review of randomly selected non-cyber incidents in terms of validity and consistency including recategorisation, if necessary (+ 21) 341
448 D.Zängerle, D.Schiereck Table 5 Results of the logistic regression M M1 M2 M3 Est SD Est SD Est SD Est SD 𝛽0 3.724 1.475* 3.639 0.706*** 0.429 0.427 1.623 0.329*** 𝛽1 −0.430 0.479 −0.659 0.146*** 0.386 0.208 −0.340 0.107** 𝛽2 0.038 0.037 0.076 0.009*** −0.006 0.020 0.073 0.009*** 𝛽3 3.372 0.666*** 2.471 0.436*** 4.047 0.603*** 2.841 0.409*** 𝛽4 4.501 1.135*** 2.663 0.564*** 5.898 1.022*** 2.937 0.496*** 𝛽5 14.025 6.635* 7.488 1.436*** 14.518 6.497* 7.632 1.416*** 𝛽6 −1.473 1.495 −1.604 0.630* – – – – 𝛽7 −3.834 1.533* −1.742 0.704* – – – – 𝛽8 −4.124 1.468** −2.330 0.658*** – – – – 𝛽9 −1.095 0.277*** −0.528 0.091*** −1.309 0.254*** −0.579 0.087*** 𝛽10 −1.452 0.410*** −0.577 0.105*** −1.979 0.367*** −0.607 0.096*** 𝛽11 −3.139 1.605 −1.169 0.181*** −3.352 1.571* −1.185 0.178*** 𝛽12 0.129 0.459 0.239 0.086** – – – – 𝛽13 1.058 0.480* 0.269 0.098** – – – – 𝛽14 1.086 0.461* 0.338 0.095*** – – – – 𝛽15 0.064 0.025* – – 0.079 0.023*** – – 𝛽16 0.087 0.034* – – 0.130 0.031*** – – 𝛽17 0.147 0.096 – – 0.166 0.094 – – 𝛽18 0.012 0.034 – – – – 𝛽19 −0.062 0.035 – – – – 𝛽20 −0.063 0.034 – – – – AIC 1903.3 1926.1 1917.3 1930.5 LL −930.6 −948.1 −946.6 −956.3 R2 0.38 0.30 0.34 0.28
449 Modelling andpredicting enterprise‑level cyber risks in… Est: Estimated parameter, SD: Standard deviation, AIC: Akaike information criterion, LL: Log-likelihood, R2 : Pseudo-measure according to McKelvey and Zavoina (1975), AUC: Area under curve of the ROC curve, HLT: Hosmer–Lemeshow test, ***: p value ≤0.001 , **: p value ∈(0.001;0.01] , *: p value ∈(0.01;0.05 ] Table 5 (continued) M M1 M2 M3 Est SD Est SD Est SD Est SD AUC 0.72 0.71 0.70 0.69 HLT 13.0 26.1 17.5 18.7
450 D.Zängerle, D.Schiereck are further assessed in terms of MAE and MSE, as reflected in Table6. Both the MAE and the MSE are low, with 2.0% and 0.1% respectively. Furthermore, M2 has the lowest MAE and MSE regarding the subcategories municipal bank(MB) and insurance(I). Hence, the adequacy and accuracy of model M2 can be sufficiently confirmed. Based on M2, the probability 1−pi,T+1 of a cyber event for company i will be predicted. Modelling severity We subsequently model the severity of cyber losses according to the proposed mixed model from Eq.(3). As described previously, only 207 cyber losses with a loss amount yit >0 are used in the following analysis. Due to the very small dataset and the fact that six parameters of the vector Θ={𝜇,𝜎𝜇,𝜉,𝜙𝜇,𝜇G,𝜎G} need to be estimated, separated modelling by subindustry—analogous to the probability of occurrence—cannot be conducted in order to ensure convergence and robust estimation. Figure1 depicts the plotted log-transformed cyber losses and the fitted mixed model, which exhibits a good overall fit to the data. In greater detail, Table7 presents the estimated values and standard deviations of the parameter vector Θ for T=2018 and the estimation results with a truncated period from t0=2005 to T=2013, …, 2017 . The truncated analysis is performed to affirm the overall robustness of the model due to data scarcity. For T=2018 , we observe a log threshold μ=7.12 , which nominally is equal to EUR13.2million (107.12). Regarding the normal distribution below the threshold, the log-expected value is 𝜇G=5.62 (nominally EUR 417,000) with a standard deviation of log(𝜎G)=0.67 . The scaling parameter 𝜎𝜇 of the generalised Pareto distribution is equal to 0.06 and the shape parameter 𝜉=1.56 . With 𝜙𝜇=0.21 , approximately every fifth cyber loss is above the threshold value µ. Furthermore, Table7 presents that there are no significant changes in the estimated parameters while truncating the time period, indicating a very robust estimation of the parameter vector Θ and a robust mixed-model approach. Table 6 Mean absolute error (MAE) and mean squared error (MSE) of the four regression models Σ: Total, B: Bank, MB: Municipal bank, I: Insurance, O: Other Bold highlights the selected modelModel M2 MAE MSE Σ B MB I O Σ B MB I O M 2.1% 2.9% 4.0% 6.5% 4.0% 0.1% 0.1% 0.3% 1.0% 0.3% M1 1.8% 2.8% 4.4% 7.0% 5.6% 0.1% 0.1% 0.3% 1.2% 0.6% M2 2.0% 2.8% 4.0% 6.4% 4.1% 0.1% 0.1% 0.3% 1.0% 0.3% M3 1.8% 2.8% 4.5% 7.0% 5.6% 0.1% 0.1% 0.3% 1.2% 0.6%
451 Modelling andpredicting enterprise‑level cyber risks in… Modelling temporal dependence Following the methodology, we next model the serial trend based on the D-vine copula from Eq.(5). Regarding the pair-copula construction, six bivariate copulas are considered Ω = {Independent,Gaussian,Clayton,Frank,Gumbel,Joe} . Given the time frame T−t0=13 , a maximum of 13 trees could be estimated. However, due to the sparse information on serial trends in the database, we decide to limit the estimation to five trees, meaning that the temporal dependence of the last six years is Fig. 1 Histogram of the logtransformed cyber losses and plot of the estimated mixed model (red line) Table 7 Estimated parameter values and standard deviations (SD) of Θ while using different time intervals t0=2005 and T=2013, …, 2018 [bold] highlights the selected year - 2018 is selected 𝝁 SD( 𝝁 ) 𝝈𝝁 SD( 𝝈𝝁) 𝝃 SD( 𝝃 ) 𝝓𝝁 SD( 𝝓𝝁 ) 𝝁G SD( 𝝁G ) 𝝈G SD( 𝝈G ) 2013 7.12 0.00 0.02 0.01 1.29 0.32 0.21 0.04 5.56 0.06 0.63 0.05 2014 7.12 0.00 0.03 0.01 1.56 0.46 0.17 0.03 5.50 0.05 0.60 0.04 2015 7.12 0.00 0.05 0.02 1.59 0.44 0.19 0.03 5.58 0.06 0.67 0.05 2016 7.12 0.00 0.08 0.04 1.45 0.56 0.21 0.03 5.58 0.05 0.66 0.04 2017 7.12 0.00 0.08 0.04 1.47 0.54 0.20 0.03 5.61 0.05 0.66 0.04 2018 7.12 0.00 0.06 0.03 1.56 0.46 0.21 0.03 5.62 0.05 0.67 0.04
452 D.Zängerle, D.Schiereck taken into account in the copula model. For each tree Tr1,…,Tr5 , the bivariate linking copula with the lowest AIC is chosen. As reflected in the results in Table8 (Panel A), the Frank copula demonstrates the lowest AIC for all trees, which is why the Frank copula is selected to represent the pairwise serial trend. It is important to note that for the Gumbel and Joe copula 𝜂 ≈1 and for the Clayton copula 𝜂 ≈0 regarding the trees Tr1,…,Tr5 , suggesting very little to no temporal dependence. However, the log-likelihood and AIC are significantly less favourable in comparison to the Frank and Gaussian copula. Panel B provides the estimated parameter value, 𝜂 , of the selected bivariate linking copula (Frank), its standard deviation and the Kendall rank correlation coefficient 𝜏 for Tr1,…,Tr5 . The parameter 𝜂 of the Frank copula is negative for all trees, indicating a negative temporal dependence. This result suggests that if a company has not yet experienced a cyber loss, it is relatively likely that a loss will occur in the next time period. However, if there has been a previous cyber incident, it is relatively unlikely that another cyber incident will occur within the next five years. This negative dependence may be a result of the fact that (external) attackers are not interested in breaching the same company twice. In the aftermath of an attack, companies tend to close security gaps and invest in their cyber risk management (Kamiya etal. 2021). Further assessment of the goodness of fit reveals that the RPS of the mixed D-vine with Frank is the lowest at 0.219, followed by the mixed D-vine with Gauss (0.251) and independence copula (0.311). Therefore, the mixed D-vine with Frank provides the best fit and is chosen to represent the serial trend. Predicting thenexttime period Finally, we predict the frequency and severity of an enterprise-level cyber event with respect to the next time period T+1=2019 . Table9 summarises the key statistical values derived from the distribution Yi,T+1∣yi for randomly selected companies in the four industry categories. Values for the maximum and tail value at risk (TVaR) are not presented because the shape parameter 𝜉>1 (i.e. we deal with infinite mean models with extreme uncertainties in very high quantiles; see e.g. Chavez-Demoulin etal. 2016; Eling and Wirfs 2019). Regarding the probability of occurrence 1−pi,T+1 , the chance of a cyber incident in the next year is predicted to be 0.6% for a selected municipal bank, 2.0% for a bank, 8.7% for an insurer and 1.4% for any other financial services company. Under the condition that a cyber incident does occur in the next year, the minimum loss value is estimated to be around EUR1,000–3,000, while the median loss ranges between EUR455,000 (insurance) and EUR585,000 (other). With respect to value at risk (VaR) measures, the VaR(90%) is equal to EUR14.6million for the municipal bank and EUR15.4million for the selected bank, while the VaR(95%) is approximately EUR18.5–22.3million. At a higher confidence level, the VaR(99%) ranges from EUR69.5million (municipal bank) to EUR543million (insurance). Even more extreme values are observed for the VaR(99.5%), indicating a worst-case incident that can cause the collapse of a company.
453 Modelling andpredicting enterprise‑level cyber risks in… Table 8 Results of the linking copula selection for the five-dimensional D-vine structure Bold highlights the selected parameter - Frank copula is selected Tr1 Tr2 Tr3 Tr4 Tr5 LL AIC LL AIC LL AIC LL AIC LL AIC A. Copula selection Independence −27.34 56.68 −16.84 35.67 −7.97 17.93 −0.16 2.32 2.45 −2.90 Gaussian −18.73 39.45 −17.92 37.85 −6.79 15.57 1.71 −1.43 3.23 −4.46 Clayton −27.34 56.67 −16.83 35.66 −7.96 17.93 −0.16 2.32 2.45 −2.90 Frank −18.30 38.60 −15.96 34.93 −5.59 13.17 2.67 −3.34 4.28 −6.56 Gumbel −27.34 56.68 −16.84 35.67 −7.94 17.87 −0.13 2.25 2.97 −3.93 Joe −27.34 56.68 −16.84 35.67 −7.88 17.76 −0.11 2.22 3.06 −4.12 B. Statistics of selected copula 𝜂 −2.2565 −0.9591 −1.4439 −1.6284 −1.1577 SD 0.5588 0.5103 0.5890 0.5890 0.7529 Kendall τ−0.2383 −0.1054 −0.1566 −0.1759 −0.1268
454 D.Zängerle, D.Schiereck Discussion andconclusion This study provides new insights on the empirical nature and prediction of cyber risks at the enterprise level under data scarcity. We introduced the ÖffSchOR database to cyber risk research and applied advanced modelling techniques adapted from the work of Shi and Yang (2018), Eling and Wirfs (2019) and Fang etal. (2021) to predict the frequency, severity and serial trend of enterprise cyber risks. Our findings first suggest that cyber risks are indeed different from operational risks. In particular, we found that cyber risks are lower on average, less skewed and less extreme compared to non-cyber risks in the dataset (Biener etal. 2015; Woods and Böhme 2021). Second, the industry subcategories exhibited different probabilities of occurrence, a finding which has not yet been addressed in previous studies. Third, in modelling the impact of the log-transformed cyber incidents, the POT model with a normal distribution below the threshold demonstrated a satisfying fit, which is in line with previous empirical results and supports the differentiation of daily and extreme cyber risks (Eling and Loperfido 2017; Eling and Wirfs 2019). Due to the limited data, a separate loss modelling for each subcategory was not possible. By leveraging the serial dependence, the D-vine copula was able to predict the impact of a potential cyber incident in the next time period with a negative correlation over time (Fang etal. 2021). The prediction results provide some of the first quantitative insights on the financial impact of a cyber incident at the enterprise level based on historic data. Our results underline that high-level descriptive statistics from commercial datasets might be misleading for enterprise risk managers due to information asymmetry and interdependence of loss events (Eling and Wirfs 2016a; Marotta etal. 2017; Zeller and Scherer 2021). In particular, our model predicted a median enterpriselevel loss amount of EUR455,000− 585,000, only a fraction of the millions of dollars often cited in surveys (e.g. USD3.86million; IBM Security 2020). In a U.K. survey, the maximum loss is around GBP310,000 (~ EUR370,000; Heitzenrater and Simpson 2016), while Romanosky (2016) estimates the average data breach loss to be even lower, at USD 200,000 (~ EUR 180,000), bearing in mind that data breaches only account for 25% of cyber events and that the transfer from data Table 9 Predicted probability of occurrence 1−pi,T+1 and statistical values (in EUR thousand) of the distribution Yi,T+1∣ y i > 0 for randomly selected companies Bank Municipal bank Insurance Other 1 − pi,T + 1 2.0% 0.6% 8.7% 1.4% Minimum 1 1 1 3 25% quantile 182 123 147 205 50% quantile 533 465 455 585 75% quantile 2,218 2,033 1,707 2,679 VaR(90%) 15,373 14,611 14,244 14,971 VaR(95%) 22,309 18,446 18,957 20,326 VaR(99%) 351,255 69,496 542,616 292,680 VaR(99.5%) 490,671 107,435 1,902,275 720,306
455 Modelling andpredicting enterprise‑level cyber risks in… breaches to actual costs is misrepresentative (Eling and Wirfs 2019). Moreover, the estimated extreme losses of VaR(99%) and VaR(99.5%) can be compared to mega breaches such as those of Home Depot (USD340million), Anthem (USD407million) and Yahoo (USD502million) in the U.S. (Poyraz etal. 2020) or to General Data Protection Regulation (GDPR) fines on Whatsapp (EUR 225 million) and Amazon (EUR745million) in Europe (CNPD 2021; EDPB 2021). The estimated VaR(95%) can be interpreted as the lower limit of a GDPR penalty, at a minimum of EUR20million (Poyraz etal. 2020). This information supports the impression that our estimates are reasonable in size. Furthermore, our findings suggest that cyber risks are less heavy tailed than often anticipated. For example, Eling and Wirfs (2019) simulate VaR measures for a small bank with 5,000 employees, which is comparable to our municipal bank category. Our estimated figures are significantly lower, such as EUR18.5million vs. EUR48million (USD55million) for the VaR (95%) and EUR69.5million vs. EUR422million (USD480million) for the VaR (99%). One conclusion from these findings is that cyber risks are just not that harmful (Woods and Böhme 2021). Another reason suggested by practitioners is that attackers have focused on easier targets while the financial services industry is comparably well protected due to regulated risk management and anti-money laundering systems. However, cyber risks are still heavy tailed and extreme. With every fifth cyber incident above the threshold of EUR13.2million, there is still a (small) chance of a devastating cyber event seriously harming an individual company (Eling etal. 2016; Wheatley etal. 2021). Comparing the four subcategories, our findings imply that bigger banks suffer from a higher potential loss than smaller (municipal) banks, indicating that the loss amount might be correlated to the company size (i.e. revenue or number of employees; Poyraz etal. 2020). Furthermore, the selected insurance company exhibited a four-times greater chance of a cyber incident, with the highest estimated risk measures for VaR (99%) and VaR (99.5%). Similar heavy tails were observed for the category other consisting of payment providers, stock exchanges and other financial services providers, which seems plausible due to their high interconnectivity to other companies. In practice, most risk and expert assessments are solely qualitative due to the limited data available on cyber incidents. For example, the Operationally Critical Threat, Asset, and Vulnerability Evaluation (OCTAVE) provided one of the first frameworks to identify and manage information security risks by analysing a company’s asset, threat and vulnerability information (Alberts etal. 1999). Further information security risk assessment (ISRA) methods have been developed, with the Core Unified Risk Framework (CURF) being the most comprehensive and allinclusive approach (Wangen etal. 2018). A specific cyber risk classification framework named Quantitative Bow-tie (QBowTie) has been suggested by Sheehan etal. (2021) combining proactive and reactive barriers to reduce a company’s risk exposure and quantify the risk. However, all these (qualitative) methods are generally based on the assessment of probability of occurrence and of the associated consequence of an event, i.e. requiring a quantification of the (cyber) risk. Compared to that, our analysis provides a helpful tool in the ongoing quantification of cyber risks. Nevertheless, the method also comes with limitations. First,
456 D.Zängerle, D.Schiereck researchers have argued that the rapidly changing cyber risk environment may render historic data useless (CRO Forum 2014; Eling and Schnell 2016). Given ongoing digitisation and in times of a global pandemic, the usefulness of historic data can be questioned. A further limitation is the assumption of independence between entities. Particularly for extreme cyber risks and mega breaches, there is a high correlation between companies (Biener etal. 2015). However, our framework could be extended to model both the serial and cross-company dependence, as conceptually shown by Acar etal. (2019) and Zhao etal. (2020) for dense data. Further limitations arise due to the use of the ÖffSchOR database. In particular, ÖffSchOR relies on print and online media to detect operational risk events which could bias the recorded loss events and in turn the modelling results. The latter could also be influenced by the historical (log-normal) distribution of the cyber loss severity. Furthermore, the total number of identified cyber events is rather small compared to other studies within the research (e.g. 1,579 cyber incidents are analysed by Eling and Wirfs 2019), challenging the robustness of our results. Finally, due to the limited dataset, we did not distinguish between different cyber risks or loss categories even though different types of cyber risks follow different distributions (Eling and Loperfido 2017; Eling and Jung 2018) and cyber risks do not only cause economic losses, but also intangible losses, including reputational damage (Xie etal. 2020). Despite these limitations and the dynamic nature of cyber risks (Boyer 2020), this study contributes to the literature on cyber risk measurement and can help practitioners such as risk managers, insurers and policymakers by providing a quantitative and data-driven cyber risk assessment. Insurance stakeholders particularly face a major challenge in assessing and understanding cyber risk due to the lack of historical data (Cremer etal. 2022). We believe that the provided methodology could be combined and integrated with existing pricing tools and factors from cyber insurers to better evaluate cyber risk and the required risk-based premiums at the enterprise level (Nurse etal. 2020). There are plenty of future research opportunities to further develop quantitative approaches. With better and more data, more accurate models can be designed, for example by including both cyber incident data and corporate financial data as proposed by Palsson etal. (2020) or by using network models as seen in Fahrenwaldt etal. (2018), Jevtić and Lanchier (2020) and Wu etal. (2021). The integration of different approaches from diverse disciplines poses extensive future opportunities in the field of cyber risk measurement (Falco etal. 2019). Appendix Additional formulas See Fang etal. (2021), Shi and Yang (2018), and Smith (2015) for further technical details.
457 Modelling andpredicting enterprise‑level cyber risks in… Funding Open Access funding enabled and organized by Projekt DEAL. Declarations Conflict of interest On behalf of all authors, the corresponding author states that there is no conflict of interest. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http:// creat iveco mmons. org/ licen ses/ by/4. 0/. References Aas, Kjersti, Claudia Czado, Arnoldo Frigessi, and Henrik Bakken. 2009. Pair-copula constructions of multiple dependence. Insurance: Mathematics and Economics 44 (2): 182–198. https:// doi. org/ 10. 1016/j. insma theco. 2007. 02. 001. Acar, Elif F., Claudia Czado, and Martin Lysy. 2019. Flexible dynamic vine copula models for multivariate time series data. Econometrics and Statistics 12: 181–197. https:// doi. org/ 10. 1016/j. ecosta. 2019. 03. 002. (A.1) fi,s,t∣(s+1)∶(t−1) � ys,yt∣y(s+1)∶(t−1) � = ⎧ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎩ Cs,t;(s+1)∶(t−1)(Fis∣(s+1)∶(t−1)(0�y(s+1)∶(t−1)),Fit∣(s+1)∶(t−1)(0�y(s+1)∶(t−1))) Fis∣(s+1)∶(t−1)(0�y(s+1)∶(t−1))Fit∣(s+1)∶(t−1)(0�y(s+1)∶(t−1)),ys=0, yt=0, c1,s,t;(s+1)∶(t−1)(Fis∣(s+1)∶(t−1)(ys�y(s+1)∶(t−1)),Fit∣(s+1)∶(t−1)(0�y(s+1)∶(t−1))) Fit∣(s+1)∶(t−1)(0�y(s+1)∶(t−1)),ys>0, yt=0, c2,s,t;(s+1)∶(t−1)(Fis∣(s+1)∶(t−1)(0�y(s+1)∶(t−1)),Fit∣(s+1)∶(t−1)(yt�y(s+1)∶(t−1))) Fis∣(s+1)∶(t−1)(0�y(s+1)∶(t−1)),ys=0, yt>0, c s,t;(s+1)∶(t−1) (F is∣(s+1)∶(t−1) (y s� y (s+1)∶(t−1) ),F it∣(s+1)∶(t−1) (y t� y (s+1)∶(t−1) )),y s >0, y t > 0 (A.2) F is∣(s+1)∶(t−1) � ys∣y(s+1)∶(t−1) � = ⎧ ⎪ ⎨ ⎪ ⎩ Cs,t−1;(s+1)∶(t−2)(Fis∣(s+1)∶(t−2)(ys�y(s+1)∶(t−2)),Fi(t−1)∣(s+1)∶(t−2)(0�y(s+1)∶(t−2))) Fi(t−1)∣(s+1)∶(t−2)(0�y(s+1)∶(t−2)),yt−1=0, c2,s,t−1;(s+1)∶(t−2)�Fis∣(s+1)∶(t−2)�ys � y(s+1)∶(t−2)�,Fi(t−1)∣(s+1)∶(t−2)�yt−1 � y(s+1)∶(t−2)��,yt−1> 0. (A.3) F it∣(s+1)∶(t−1) ( yt∣y(s+1)∶(t−1) ) = {Ct,s+1;(s+2)∶(t−1)(Fit∣(s+2)∶(t−1)(yt|y(s+2)∶(t−1)),Fi(s+1)∣(s+2)∶(t−1)(0|y(s+2)∶(t−1))) Fi(s+1)∣(s+2)∶(t−1)(0 | y(s+2)∶(t−1)),ys+1=0, c2, t , s+ 1; (s+ 2 )∶(t− 1 )( F it∣(s+ 2 )∶(t− 1 )( y t| y (s+ 2 )∶(t− 1 )) ,F i(s+ 1 )∣(s+ 2 )∶(t− 1 )( y s+ 1 | y (s+ 2 )∶(t− 1 ))) ,y s+ 1> 0.