scieee AI-readable full text Open interactive document viewer

Assessing asset-liability risk with neural networks

Cheridito, Patrick,Ery, John,Wüthrich, Mario V.

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Cheridito, Patrick; Ery, John; Wüthrich, Mario V. Article Assessing asset-liability risk with neural networks Risks Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Cheridito, Patrick; Ery, John; Wüthrich, Mario V. (2020) : Assessing asset-liability risk with neural networks, Risks, ISSN 2227-9091, MDPI, Basel, Vol. 8, Iss. 1, pp. 1-17, https://doi.org/10.3390/risks8010016 This Version is available at: https://hdl.handle.net/10419/257971 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ risks Article Assessing Asset-Liability Risk with Neural Networks Patrick Cheridito ∗, John Ery and Mario V. Wüthrich RiskLab, ETH Zurich, 8092 Zurich, Switzerland; [email protected] (J.E.); [email protected] (M.V.W.) *Correspondence: [email protected] Received: 30 November 2019; Accepted: 4 February 2020; Published: 9 February 2020   Abstract: We introduce a neural network approach for assessing the risk of a portfolio of assets and liabilities over a given time period. This requires a conditional valuation of the portfolio given the state of the world at a later time, a problem that is particularly challenging if the portfolio contains structured products or complex insurance contracts which do not admit closed form valuation formulas. We illustrate the method on different examples from banking and insurance. We focus on value-at-risk and expected shortfall, but the approach also works for other risk measures. Keywords: asset-liability risk; risk capital; solvency calculation; value-at-risk; expected shortfall; neural networks; importance sampling 1. Introduction Different financial risk management problems require an assessment of the risk of a portfolio of assets and liabilities over a given time period. Banks, mutual funds and hedge funds usually calculate value-at-risk numbers for different risks over a day, a week or longer time periods. Under Solvency II, insurance companies have to calculate one-year value-at-risk, whereas the Swiss Solvency Test demands a computation of one-year expected shortfall. A determination of these risk figures requires a conditional valuation of the portfolio given the state of the world at a later time, called risk horizon. This is particularly challenging if the portfolio contains structured products or complicated insurance policies which do not admit closed form valuation formulas. In theory, the problem can be approached with nested simulation, which is a two-stage procedure. In the outer stage, different scenarios are generated to model how the world could evolve until the risk horizon, whereas in the inner stage, cash flows occurring after the risk horizon are simulated to estimate the value of the portfolio conditionally on each scenario; see, e.g., Lee (1998), Glynn and Lee (2003), Gordy and Juneja (2008) , Broadie et al. (2011) or Bauer et al. (2012). While nested simulation can be shown to converge for increasing sample sizes, it is often too time-consuming to be useful in practical applications. A more pragmatic alternative, usually used for short risk horizons, is the delta-gamma method, which approximates the portfolio loss with a second order Taylor polynomial; see, e.g., Rouvinez (1997), Britten-Jones and Schaefer (1999) or Duffie and Pan (2001). If first and second derivatives of the portfolio with respect to the underlying risk factors are accessible, the method is computationally efficient, but its accuracy depends on how well the second order Taylor polynomial approximates the true portfolio loss. Similarly, the replicating portfolio approach approximates future cashflows with a portfolio of liquid instruments that can be priced efficiently; see e.g., Wüthrich (2016), Pelsser and Schweizer (2016), Natolski and Werner (2017)orCambou and Filipovi´c (2018). Building on work on American option valuation (see e.g., Carriere 1996;Longstaff and Schwartz 2001;Tsitsiklis and Van Roy 2001) , Broadie et al. (2015) as well as Ha and Bauer (2019) have proposed to regress future cash flows on finitely many basis functions depending on state variables known at the risk horizon. This gives good results in a number of applications. However, typically, for it to work well, the basis functions have to be chosen well-adapted to the problem. Risks 2020,8, 16; doi:10.3390/risks8010016 www.mdpi.com/journal/risks Risks 2020,8, 16 2 of 17 In this paper we use a neural network approach to approximate the value of the portfolio at the risk horizon. Since our goal is to estimate tail risk measures such as value-at-risk and expected shortfall, we employ importance sampling when training the networks and estimating the risk figures. In addition, we try different regularization techniques and test the adequacy of the neural network approximations using the defining property of the conditional expectation. Neural networks have also been used for the calculation of solvency capital requirements by Hejazi and Jackson (2017), Fiore et al. (2018) and Castellani et al. (2019). But they all exclusively focus on value-at-risk, do not make use of importance sampling and apply neural networks in a slightly different way. Hejazi and Jackson (2017) develop a neural network interpolation scheme within a nested simulation framework. Fiore et al. (2018) and Castellani et al. (2019) both use reduced-size nested Monte Carlo to generate training samples for the calibration of the neural network. Here, we directly regress future cash-flows on a neural network without producing training samples. The remainder of the paper is organized as follows. In Section 2, we set up the underlying risk model, introduce the importance sampling distributions we use in the implementation of our method and recall how value-at-risk and expected shortfall can be estimated from a finite sample of simulated losses. Section 3discusses the training, validation and testing of our network approximation of the conditional valuation functional. In Section 4we illustrate our approach on three typical risk calculation problems: the calculation of risk capital for a single put option, a portfolio of different call and put options and a variable annuity contract with guaranteed minimum income benefit. Section 5 concludes the paper. 2. Asset-Liability Risk We denote the current time by 0 and are interested in the value of a portfolio of assets and liabilities at a given risk horizon τ> 0. Suppose all relevant events taking place until time τ are described by a d -dimensional random vector X= (X1 , . . . , Xd) defined on a measurable space (Ω , F) that is equipped with two equivalent probability measures, P and Q . We think of P as the real-world probability measure and use it for risk measurement. Q is a risk-neutral probability measure specifying the time-τvalue of the portfolio through V=v(X) + EQ"I ∑ i=1 Nτ Nti CtiX#, where v is a measurable function from Rd to R describing the part of the portfolio whose value is directly given by the underlying risk factors X1 , . . . , Xd ; Ct1 , . . . , CtI are random cash flows occurring at times τ<t1<··· <tI ; and Nτ , Nt1 , . . . , NtI model the evolution of a numeraire process used for discounting. Our goal is to compute ρ(L) for a risk measure ρ and the time τ net liability L=−V , which can be written as L=EQ[Y|X] for Y=−v(X)− I ∑ i=1 Nτ Nti Cti. Our approach works for a wide class of risk measures ρ . But for the sake of concreteness, we concentrate on value-at-risk and expected shortfall. We follow the convention of McNeil et al. (2015) and define value-at-risk at a level α∈(0, 1), such as 0.95, 0.99 or 0.995, as the left α-quantile VaRα(L):=min {x∈R:P[L≤x]≥α}=sup {x∈R:P[L≥x]>1−α}(1) Risks 2020,8, 16 3 of 17 and expected shortfall as ESα(L):=1 1−αZ1 αVaRu(L)du. (2) There exist different definitions of VaR and ES in the literature. But if the distribution function FL of Lis continuous and strictly increasing, they are all equivalent1. 2.1. Conditional Expectations as Minimizing Functions Let P⊗Qbe the probability measure on Fgiven by P⊗Q[A]:=ZRdQ[A|X=x]π(dx), where π is the distribution of X under P and Q[A|X=x] is a regular conditional version of Q given X . We assume that Y belongs to L2(Ω , F , P⊗Q) . Then L is of the form L=l(X) for a measurable function l:Rd→Rminimizing the mean squared distance EP⊗Qh(l(X)−Y)2i; (3) see, e.g., Bru and Heinich (1985). Note that lminimizes (3) if and only if l(x) = arg min u∈RZR(u−y)2Q[Y∈dy |X=x]for π-almost all x∈Rd, where Q[Y∈dy |X=x] is a regular conditional Q -distribution of Y given X . This shows that l is unique up to π-almost sure equality and can alternatively be characterized by l(x) = arg min u∈RZR(u−y)2Q[Y∈dy |X=x]for ν-almost all x∈Rd for any probability measure ν on Rd that is equivalent to π . In particular, if Pν is the probability measure on σ(X) under which X has distribution ν , and Y is in L2(Ω , F , Pν⊗Q) , then l can be determined by minimizing EPν⊗Qh(l(X)−Y)2i(4) instead of (3) . Since we are going to approximate the expectation (4) with averages of Monte Carlo samples, this will give us some flexibility in simulating (X,Y). 2.2. Monte Carlo Estimation of Value-at-Risk and Expected Shortfall Let X1 , . . . , Xn be independent P -simulations of X . By X(1) , . . . , X(n) we denote the same sample reordered so that L(1)=l(X(1))≥ ··· ≥ L(n)=l(X(n)). To obtain P -simulation estimates of VaRα(L) and ESα(L) , we apply the VaR and ES definitions (1)–(2) to the empirical measure 1 n∑n i=1δL(i). This yields d VaRα(n):=L(j)and c ESα(n):=1 1−α j−1 ∑ i=1 L(i) n+1−j−1 (1−α)nL(j), (5) 1 More precisely, VaRα is usually defined as an α - or (1 −α) -quantile depending on whether it is applied to L or −L . So up to the sign convention, all VaR definitions coincide if FL is strictly increasing. Similarly, different definitions of ES are equivalent if FLis continuous. Risks 2020,8, 16 4 of 17 where j=min {i∈{1, . . . , n}:i/n>1−α}; see also McNeil et al. (2015). It is well known that if FL is differentiable at VaRα(L) with F0 L(VaRα(L)) >0 , then d VaRα(n) is an asymptotically normal consistent estimator of order n−1/2 ; see, e.g., David and Nagaraja (2003). If in addition, L is square-integrable, the same is true for c ESα(n) ; see Zwingmann and Holzmann (2016). 2.3. Importance Sampling Assume now that X is of the form X=h(Z) for a transformation h:Rk→Rd and a k -dimensional random vector Z with density f:Rk→R+ . To give more weight to important outcomes, we introduce a second density g on Rk satisfying g(z)> 0 for all z∈Rk , where f(z)> 0. Let Zg be a k -dimensional random vector with density g and Zg,1 , . . . , Zg,n independent simulations of Zg . We reorder the simulations such that Lg,(1)=l◦h(Zg,(1))≥ ··· ≥ Lg,(n)=l◦h(Zg,(n)) and consider the random weights w(i)=f(Zg,(i)) ng(Zg,(i)). Since f(Zg)/g(Zg) is integrable, one obtains from the law of large numbers that for every threshold x∈R,n ∑ i=1 1{l◦h(Zg,(i))≥x}w(i)(6) is an unbiased consistent estimator of the exceedance probability P[L≥x] = P[l◦h(Z)≥x]=E1{l◦h(Zg)≥x} f(Zg) g(Zg). If in addition, g can be chosen so that f(Zg)/g(Zg) is square-integrable, it follows from the central limit theorem that (6) is asymptotically normal with a standard deviation of order n−1/2. In any case, ∑n i=1w(i) converges to 1 and the random measure ∑n i=1w(i)δl◦h(Zg,(i)) approximates the distribution of L . Accordingly, we adapt the P -simulation estimators (5) by replacing i/n with ∑i m=1w(m) . This yields the g-simulation estimates d VaRg α(n):=Lg,(j)and c ESg α(n):=1 1−α j−1 ∑ i=1 w(i)Lg,(i)+ 1−1 1−α j−1 ∑ i=1 w(i)!Lg,(j), where j=min (i∈{1, . . . , n}: i ∑ m=1 w(m)>1−α). If the α -quantile lα of L=l◦h(Z) were known, the exceedance probability P[L≥lα] could be estimated by means of (6) with x=lα . To make the procedure efficient, one would try to find a density gon Rkfrom which it is easy to sample and such that the variance Var 1{l◦h(Zg)≥lα} f(Zg) g(Zg)=E1{l◦h(Zg)≥lα} f(Zg)2 g(Zg)2−E1{l◦h(Zg)≥lα} f(Zg) g(Zg)2 , Risks 2020,8, 16 5 of 17 becomes as small as possible. Since E1{l◦h(Zg)≥lα} f(Zg) g(Zg)=Eh1{l◦h(Z)≥lα}i=P[L≥lα] does not depend on g, it can be seen that gis a good importance sampling (IS) density if E1{l◦h(Zg)≥lα} f(Zg)2 g(Zg)2=E1{l◦h(Z)≥lα} f(Z) g(Z)(7) is small. We use the same criterion as a basis to find a good IS density for estimating VaRα and ESα ; see, e.g., Glasserman (2003) for more background on importance sampling. 2.4. The Case of an Underlying Multivariate Normal Distribution In many applications X can be modelled as a deterministic transformation of a random vector with a multivariate normal distribution. Examples are included in Section 4below. In this case, it can be written as X=u(AZ)(8) for a function u:Rp→Rd , a p×k -matrix A and a k -dimensional random vector Z with a standard normal density f . To keep the problem tractable, we look for a suitable IS density g in the class of Nk(m , Ik) -densities fm with different mean vectors m∈Rk and covariance matrix equal to the k -dimensional identity matrix Ik . To determine a good choice of m , let v∈Rp be the vector with components vi=       1 if l◦uis increasing in yi −1 if l◦uis decreasing in yi 0 else. If ATv= 0, we choose m= 0. Otherwise, vTAZ/kATvk2 is standard normal. Denote its α -quantile by zα. Then Zfalls into the region D=nz∈Rk:vTAz/kATvk2≥zαo with probability 1 −α , and l◦u(Az) tends to be large for z∈D . We choose m as the maximizer of f(z) subject to z∈D, which leads to m=ATv kATvk2 zα. It is easy to see that this yields f(z)<fm(z)for all z∈D. So, if Dis sufficiently similar to the region nz∈Rk:l◦u(Az)≥lαo, it can be seen from (7) that this choice of IS density will yield a reduction in variance. 3. Neural Network Approximation Usually, the distribution of the risk factor vector X is assumed to be known. For instance, in all our examples in Section 4, they are of the form (8) . On the other hand, in many real-world applications, there is no closed form expression for the loss function l:Rd→R mapping X to L . Therefore, we approximate L=l(X) with lθ(X) for a neural network lθ:Rd→R ; see e.g., Goodfellow et al. (2016). We concentrate on feedforward neural networks of the form lθ=ψ◦aθ J◦ϕ◦aθ J◦···◦ ϕ◦aθ 1, Risks 2020,8, 16 6 of 17 where •q0=d,qJ=1, and q1, . . . , qJ−1are the numbers of nodes in the hidden layers 1, . . . , J−1; •aθ j are affine functions of the form aθ j(x) = Ajx+bj for matrices Aj∈Rqj×qj−1 and vectors bj∈Rqj , for j=1, . . . , J; •ϕ is a non-linear activation function used in the hidden layers and applied component-wise. In the examples in Section 4we choose ϕ=tanh; •ψ is the final activation function. For a portfolio of assets and liabilities a natural choice is ψ=id . To keep the presentation simple, we will consider pure liability portfolios with loss L> 0 in all our examples below. Accordingly, we choose ψ=exp. The parameter vector θ consists of the components of the matrices Aj and vectors bj , j= 1, . . . , J . So it lives in Rq for q=∑J j=1qj(qj−1+ 1 ) . It remains to determine the architecture of the network (that is, J and q1 , . . . , qJ−1 ) and to calibrate θ∈Rq . Then VaR and ES figures can be estimated as described in Section 2by simulating lθ(X). 3.1. Training and Validation In a first step we take the network architecture ( J and q1 , . . . , qJ−1) as given and try to find a minimizer of θ7→ E(lθ(X)−Y)2 , where the expectation is either with respect to P⊗Q or Pν⊗Q for an IS distribution ν on Rd . To do that we simulate realizations (Xm , Ym) , m= 1, . . . , M1+M2 , of (X , Y) under the corresponding distribution. The first M1 simulations are used for training and the other M2 for validation. More precisely, we employ a stochastic gradient descent method to minimize the Monte Carlo approximation based on the training samples 1 M1 M1 ∑ m=1lθ(Xm)−Ym2(9) of E(lθ(X)−Y). At the same time we use the validation samples to check whether 1 M2 M1+M2 ∑ m=M1+1lθ(Xm)−Ym2(10) is decreasing as we are updating θ. 3.1.1. Regularization through Tree Structures If the number q of parameters is large, one needs to be careful not to overfit the neural network. For instance, in the extreme case, the network could be so flexible that it can bring (9) down to zero even in cases where the true conditional expectation l(X) = EQ[Y|X] is not equal to Y . To prevent this, one can generate the training samples by first simulating N1 realizations Xi of X and then for every Xi , drawing N2 simulations Yi,j from the conditional distribution of Y given Xi . In the simple example of Section 4.1, we chose N2=1. In Sections 4.2 and 4.3 we used N2=5. 3.1.2. Stochastic Gradient Descent In principle, one can use any numerical method to minimize (9) . But stochastic gradient descent methods have proven to work well for neural networks. We refer to Ruder (2016) for an overview of Risks 2020,8, 16 7 of 17 different (stochastic) gradient descent algorithms. Here, we randomly 2 split the M1 training samples into bmini-batches of size B. Then we update θbased on the θ-gradients of 1 B iB ∑ m=(i−1)B+1lθ(Xm)−Ym2,i=1, . . . , b. We use batch normalization and Adam updating with the default values from TensorFlow. After b gradient steps, all of the training data have been used once and the first epoch is complete. For further epochs, we reshuffle the training data, form new mini-batches and perform b more gradient steps. The procedure is repeated until the training error (9) stops to decrease or the validation error (10) starts to increase. 3.1.3. Initialization We follow standard practice and initialize the parameters of the network randomly. The final operation of the network is x7→ ψ(AJx+bJ). Since the network tries to approximate Y , and in all our examples below we use ψ=exp , we initialize the last bias as b0 J=log 1 M1 M1 ∑ m=1 Ym!. For the other parameters we use Xavier initialization; see Glorot and Bengio (2010). 3.2. Backtesting the Network Approximation After having determined an approximate minimizer θ∈Rq for a given network architecture, one can test the adequacy of the approximation lθ(X) of the true conditional expectation l(X) = EQ[Y|X] . The quality of the approximation depends on different aspects: (i) Generalization error: The true conditional expectation EQ[Y|X] is of the form l(X) for the unique 3 measurable function l:Rd→R minimizing the mean squared distance Eh(l(X)−Y)2i . To approximate l we choose a network architecture and try to find a θ∈Rq that minimizes the empirical squared distance 1 M1 M1 ∑ m=1lθ(Xm)−Ym2. (11) But if the samples (Xm , Ym) , m= 1, . . . , M1 , do not represent the distribution of (X , Y) well, (11) might not be a good approximation of the true expectation Ehlθ(X)−Y2i. (ii) Numerical minimization method: The minimization of (11) usually is a complex problem, and one has to employ a numerical method to find an approximate solution θ . The quality of lθ(X) will depend on the performance of the numerical method being used. (iii) Network architecture: It is well known that feedforward neural networks with one hidden layer have the universal approximation property; see, e.g., Cybenko (1989); Hornik et al. (1989) or Leshno et al. (1993). 2 If the training data is generated according to a tree structure as in Section 3.1.1, one can either group the simulations (Xi , Yi,j) , i= 1, . . . , N1 , j= 1, . . . , N2 , so that pairs with the same Xi -component stay together or not. In our implementations, both methods gave similar results. 3More precisely, uniqueness holds if functions are identified that agree π-almost surely. Risks 2020,8, 16 8 of 17 That is, they can approximate any continuous function uniformly on compacts to any degree of accuracy if the activation function is of a suitable form and the hidden layers contain sufficiently many nodes. As a consequence, E(lθ(X)−l(X))2 can be made arbitrarily small if the hidden layer is large enough and θ is chosen appropriately. However, we do not know in advance how many nodes we need. And moreover, feedforward neural networks with two or more hidden layers have shown to yield better results in different applications. Since we simulate from an underlying model, we are able to choose the size M1 of the training sample large and train extensively. In addition, for any given network architecture, we also evaluate the empirical squared distance (11) on the validation set (Xm , Ym) , m=M1+ 1, . . . , M2 . So we suppose the generalization error is small and our numerical method finds a good approximate minimizer θ of (11) . However, since we do not know whether a given network architecture is flexible enough to provide a good approximation to the true loss function l, we test for each trained network whether it satisfies the defining properties of a conditional expectation. The loss function l:Rd→Ris characterized by E[l(X)ξ(X)] = E[Yξ(X)] (12) for all measurable functions ξ:Rd→R such that Yξ(X) is integrable. Ideally, we would like lθ to satisfy the same condition. But there will be an approximation error, and (12) cannot be checked for all measurable functions ξ:Rd→R satisfying the integrability condition. Therefore, we select finitely many measurable subsets Bi⊆Rd , i= 1, . . . , I . Then we generate M3 more samples (Xm , Ym) of (X , Y) and test whether the differences (a) ∑M1+M2+M3 m=M2+1lθ(Xm)−Ym/M3 (b) ∑M1+M2+M3 m=M2+1lθ(Xm)−Ymlθ(Xm)/M3 (c) ∑M1+M2+M3 m=M2+1lθ(Xm)−Ym1Bi(Xm)/M3 are sufficiently close to zero. If this is not the case, we change the network architecture and train again. 4. Examples As examples, we study three different risk assessment problems from banking and insurance. For comparison we generated realizations (Xm , Ym) of (X , Y) under P⊗Q as well as Pν⊗Q , where ν is the IS distribution on Rd obtained by changing the distribution of X . In all our examples, X has a transformed normal distribution as in Section 2.4. In each case we used 1.5 million simulations with mini-batches of size 10,000 for training, 500,000 simulations for validation and 500,000 for backtesting. After training and testing the network, we simulated another 500,000 realizations of X , once under Pand then under Pν, to estimate VaR99.5%(L)and ES99%(L). We implemented the algorithms in Python. To train the networks we used the TensorFlow package. 4.1. Single Put Option As a first example, we consider a liability consisting of a single put option with strike price K= 100 and maturity T=1/3 on an underlying asset starting from s0=100 and evolving according to dSt=µStdt +σStdWP t=rStdt +σStdWQ t for an interest rate r= 1%, a drift µ= 5%, a volatility σ= 20%, a P -Brownian motion WP and aQ-Brownian motion WQ. As risk horizon we choose τ=1/52. The time τ-value of this liability is L=e−r(T−τ)EQ(K−ST)+|Sτ. Risks 2020,8, 16 15 of 17 0 5 10 15 20 25 30 35 40 25 20 15 10 5 0 5 0 5 10 15 20 25 30 35 40 5 0 5 10 15 20 25 30 35 Figure 11. Empirical evaluation of (c) of Section 3.2 for the sets B1(left) and B2(right) under P⊗QE. We also trained and tested under Pν⊗QE for an IS distribution ν on Rd . Our results are reported in Figure 12. Our estimate of VaR99.5%(L) was 138.64 without and 138.52 with IS. In comparison, the VaR estimate obtained by Ha and Bauer (2019) using 37 monomials and 40 million simulations is 139.74. Our estimates of ES99%(L) came out as 141.12 without and 142.12 with IS. There exist no reference values for this case. 0 100,000 200,000 300,000 400,000 500,000 136 137 138 139 0 100,000 200,000 300,000 400,000 500,000 137 138 139 140 141 142 Figure 12. Convergence of the empirical 99.5%-VaR ( left ) and 99%-ES ( right ) without IS (orange) and with IS (blue) compared to the reference value from Ha and Bauer (2019) (green). 5. Conclusions In this paper we have developed a deep learning method for assessing the risk of an asset-liability portfolio over a given time horizon. It first computes a neural network approximation of the portfolio value at the risk horizon. Then the approximation is used to estimate a risk measure, such as value-at-risk or expected shortfall, from Monte Carlo simulations. We have investigated how to choose the architecture of the network, how to learn the network parameters under a suitable importance sampling distribution and how to test the adequacy of the network approximation. We have illustrated the approach by computing value-at-risk and expected shortfall in three typical risk assessment problems from banking and insurance. In all cases the approach has worked efficiently and produced accurate results. Author Contributions: P.C., J.E., and M.V.W. have contributed equally to this work. All authors have read and agreed to the published version of the manuscript. Funding: As SCOR Fellow, John Ery thanks SCOR for financial support. Conflicts of Interest: The authors declare no conflict of interest. References Bauer, Daniel, Andreas Reuss, and Daniela Singer. 2012. On the calculation of the solvency capital requirement based on nested simulations. ASTIN Bulletin 42: 453–99. Risks 2020,8, 16 16 of 17 Britten-Jones, Mark, and Stephen Schaefer. 1999. Non-linear Value-at-Risk. Review of Finance, European Finance Association 5: 155–80. Broadie, Mark, Yiping Du, and Ciamac Moallemi. 2011. Efficient risk estimation via nested sequential estimation. Management Science 57: 1171–94. [CrossRef] Broadie, Mark, Yiping Du, and Ciamac Moallemi. 2015. Risk estimation via regression. Operations Research 63: 1077–97. [CrossRef] Bru, Bernard, and Henri Heinich. 1985. Meilleures approximations et médianes conditionnelles. Annales de l’I.H.P. Probabilités et Statistiques 21: 197–224. Cambou, Mathieu, and Damir Filipovi´c. 2018. Replicating portfolio approach to capital calculation. Finance and Stochastics 22: 181–203. [CrossRef] Carriere, Jacques F. 1996. Valuation of the early-exercise price for options using simulations and nonparametric regression. Insurance: Mathematics and Economics 19: 19–30. [CrossRef] Castellani, Gilberto, Ugo Fiore, Zelda Marino, Luca Passalacqua, Francesca Perla, Salvatore Scognamiglio, and Paolo Zanetti. 2019. An Investigation of Machine Learning Approaches in the Solvency II Valuation Framework. SSRN Preprint. Cybenko, G. 1989. Approximation by superpositions of a sigmoidal function. Mathematics of Control, Signals and Systems 2: 303–14. [CrossRef] David, Herbert A., and Haikady N. Nagaraja. 2003. Order Statistics, 3rd ed. Wiley Series in Probability and Statistics. Hoboken: Wiley. Duffie, Darrell, and Jun Pan. 2001. Analytical Value-at-Risk with jumps and credit risk. Finance and Stochastics 5: 155–80. [CrossRef] Fiore, Ugo, Zelda Marino, Luca Passalacqua, Francesca Perla, Salvatore Scognamiglio, and Paolo Zanetti. 2018. Tuning a deep learning network for Solvency II: Preliminary results. Mathematical and Statistical Methods for Actuarial Sciences and Finance: MAF 2018. Berlin: Springer. Glasserman, Paul. 2003. Monte Carlo Methods in Financial Engineering. Stochastic Modelling and Applied Probability. Berlin: Springer. Glorot, Xavier, and Yoshua Bengio 2010. Understanding the difficulty of training deep feedforward neural networks. Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics 9: 249–56. Glynn, Peter, and Shing-Hoi. Lee. 2003. Computing the distribution function of a conditional expectation via Monte Carlo: Discrete conditioning spaces. ACM Transactions on Modeling and Computer Simulation 13: 238–58. Goodfellow, Ian, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. Cambridge: MIT Press. Gordy, Michael B., and Sandeep Juneja. 2008. Nested simulation in portfolio risk measurement. Management Science 56: 1833–48. [CrossRef] Ha, Hongjun, and Daniel Bauer. 2019. A least-squares Monte Carlo approach to the estimation of enterprise risk. Working Paper. Forthcoming. Hejazi, Seyed Amir, and Kenneth R. Jackson. 2017. Efficient valuation of SCR via a neural network approach. Journal of Computational and Applied Mathematics 313: 427–39. [CrossRef] Hornik, Kurt, Maxwell Stinchcombe, and Halbert White. 1989. Multilayer feedforward networks are universal approximators. Neural Networks 2: 359–66. [CrossRef] Lee, Shing-Hoi. 1998. Monte Carlo Computation of Conditional Expectation Quantiles. Ph.D. thesis, Stanford University, Stanford, CA, USA. Leshno, Moshe, Vladimir Lin, Allan Pinkus, and Shimon Schocken. 1993. Multilayer feedforward networks with a non-polynomial activation function can approximate any function. Neural Networks 6: 861–67. [CrossRef] Longstaff, Francis, and Eduardo Schwartz. 2001. Valuing American options by simulation: A simple least-squares approach. The Review of Financial Studies 14: 113–47. [CrossRef] McNeil, Alexander J., Rüdiger Frey, and Paul Embrechts. 2015. Quantitative Risk Management: Concepts, Techniques and Tools, rev. ed. Princeton: Princeton University Press. Natolski, Jan, and Ralf Werner. 2017. Mathematical analysis of replication by cashflow matching. Risks 5: 13. [CrossRef] Pelsser, Antoon, and Janina Schweizer. 2016. The difference between LSMC and replicating portfolio in insurance liability modeling. European Actuarial Journal 6: 441–94. [CrossRef] [PubMed] Rouvinez, Christophe. 1997. Going Greek with VaR. Risk Magazine 10: 57–65. Ruder, Sebastian. 2016. An overview of gradient descent optimization algorithms. arXiv arXiv:1609.04747. Risks 2020,8, 16 17 of 17 Tsitsiklis, John, and Benjamin Van Roy. 2001. Regression methods for pricing complex American-style options. IEEE Transactions on Neural Networks 12: 694–703. [CrossRef] [PubMed] Wüthrich, Mario V. 2016. Market-Consistent Actuarial Valuation. EAA Series. Berlin: Springer. Zwingmann, Tobias, and Hajo Holzmann. 2016. Asymptotics for the expected shortfall. arXiv arXiv:1611.07222. c 2020 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).