scieee AI-readable full text Open interactive document viewer

A neural network Monte Carlo approximation for expected utility theory

Zhu, Yichen,Escobar, Marcos

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Zhu, Yichen; Escobar, Marcos Article A neural network Monte Carlo approximation for expected utility theory Journal of Risk and Financial Management Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Zhu, Yichen; Escobar, Marcos (2021) : A neural network Monte Carlo approximation for expected utility theory, Journal of Risk and Financial Management, ISSN 1911-8074, MDPI, Basel, Vol. 14, Iss. 7, pp. 1-18, https://doi.org/10.3390/jrfm14070322 This Version is available at: https://hdl.handle.net/10419/258426 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ Journal of Risk and Financial Management Article A Neural Network Monte Carlo Approximation for Expected Utility Theory Yichen Zhu †and Marcos Escobar-Anel *,†   Citation: Zhu, Yichen, and Marcos Escobar-Anel. 2021. A Neural Network Monte Carlo Approximation for Expected Utility Theory. Journal of Risk and Financial Management 14: 322. https:// doi.org/10.3390/jrfm14070322 Academic Editor: Daniel N. Ostrov Received: 17 May 2021 Accepted: 18 June 2021 Published: 13 July 2021 Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. Copyright: © 2021 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). Department of Statistical and Actuarial Sciences, Western University, London, ON N6A5B7, Canada; [email protected] *Correspondence: mar[email protected] † These authors contributed equally to this work. Abstract: This paper proposes an approximation method to create an optimal continuous-time portfolio strategy based on a combination of neural networks and Monte Carlo, named NNMC. This work is motivated by the increasing complexity of continuous-time models and stylized facts reported in the literature. We work within expected utility theory for portfolio selection with constant relative risk aversion utility. The method extends a recursive polynomial exponential approximation framework by adopting neural networks to fit the portfolio value function. We developed two network architectures and explored several activation functions. The methodology was applied on four settings: a 4/2 stochastic volatility (SV) model with two types of market price of risk, a 4/2 model with jumps, and an Ornstein–Uhlenbeck 4/2 model. In only one case, the closed-form solution was available, which helps for comparisons. We report the accuracy of the various settings in terms of optimal strategy, portfolio performance and computational efficiency, highlighting the potential of NNMC to tackle complex dynamic models. Keywords: neural networks; expected utility theory; CRRA utility; 4/2 stochastic volatility model 1. Introduction Optimally allocating a collection of financial investments such as stocks, bonds and commodities has been a topic of concern to financial institutions and shareholders at least since the pioneering work of Markowitz’s mean-variance portfolio theory in 1952. People then realized the potential of diversification and their work laid the foundations for the development of portfolio analysis in both academia and industry. These initial results were in discrete-time, but it was not long before continuous-time portfolio decisions were produced in the alternative paradigm of expected utility theory, as can be seen in Merton (1969) . The author assumed that the investor is able to continuously adjust their position, and the stock price process is modelled by a geometric Brownian motion (GBM). The optimal trading strategy and consumption policy that maximize the investor’s expected utility were obtained in closed-form by solving a Hamilton–Jacobi–Bellman equation. The beauty and practicality of this continuous-time solution has led many researchers onto this path, producing optimal closed-form strategies for a wide range of models. For example, Kraft (2005) considered the stochastic volatility (SV) Heston model, Heston (1993) . Flor and Larsen (2014) constructed a portfolio of stocks and fixed-income market products to hedge the interest rate risk. Explicit solutions in the presence of regime switching, stochastic interest rate and stochastic volatility was presented in Escobar et al. (2017), whilst the positive performance of their portfolio is confirmed by empirical study. For the commodities asset class, Chiu and Wong (2013) modelled a mean-reverting risky asset by an exponential Ornstein–Uhlenbeck (OU) process and solved the investment problem for an insurer subject to the random payment of insurance claim. These models are particular cases of the quadratic-affine family (see Liu (2006)), one of the broadest models solvable in closed-form. The value function for a model in J. Risk Financial Manag. 2021,14, 322. https://doi.org/10.3390/jrfm14070322 https://www.mdpi.com/journal/jrfm J. Risk Financial Manag. 2021,14, 322 2 of 18 this family is the product of a function of wealth and an exponential quadratic function. Nonetheless, the complexity of financial markets has continued increasing every decade, with researchers detecting new stylized facts and proposing new models outside the quadratic-affine. Needless to say, investors must rely on these advanced models for better financial decisions, however. closed-form solutions are no longer guaranteed. One example of these advanced models is the GBM 4/2 model, introduced in Grasselli (2017). The model improves the Heston model in terms of the better fitting of implied volatility surfaces and historical volatilities patterns. The optimal portfolio problem with the GBM 4/2 model is solvable for certain types of market price of risk (MPR, see Cheng and EscobarAnel (2021)), while the optimal trading strategy has not been found yet with an MPR proportional to the instantaneous volatility. More recently, an OU 4/2 model, which unifies the mean-reverting drift and stochastic volatility in a single model, was presented in Escobar-Anel and Gong (2020) . The model targets two asset classes: commodities and volatility indexes. The optimal portfolio with the OU 4/2 model is not in closed form. This motivates approximation methods for dynamic portfolio choice. Most approximation methods follow the idea from martingale method (see Karatzas et al. (1987)) or dynamic programming technique Brandt et al. (2005). Cvitani´c et al. (2003) proposed a simulation-based method seeking the financial replication of the optimal terminal wealth given in the martingale method. Detemple et al. (2003) developed a comprehensive approach for the same investment problems, and the application of Malliavin calculus enhances its accuracy. The work in Brandt et al. (2005) led to the BGSS method, which was inspired by the popular least-square Monte Carlo method of Longstaff and Schwartz (2001) . BGSS pioneered the recursive approximation method for dynamic portfolio choice. Cong and Oosterlee (2017) enhanced BGSS with the stochastic grid bundling method (SGBM) for conditional expectation estimation introduced in Jain and Oosterlee (2015). More recently, a polynomial affine method for constant relative risk aversion utility (PAMC) was recently developed in Zhu et al. (2020). The method takes advantage of the quadratic-affine structure, leading to superior accuracy and efficiency in the approximation of the optimal strategy and value function. In this paper, we extend the methodology in PAMC using neural networks. The history of artificial neural networks goes back to McCulloch and Pitts (1943), where the author created the so-called “threshold logic” on the basis of the neural networks of the human brain in order to mimic human thoughts. Deep learning has since steadily evolved. Almost three decades later, back propagation, a widely used algorithm in neural network’s parameter fitting for supervised learning, was introduced, see Linnainmaa (1970) . The importance of back propagation was only fully recognized when Rumelhart et al. (1986) showed that it can provide interesting distribution representations. The universal approximation theorem (see Cybenko (1989)) illustrated that every bounded continuous function can be approximated by a network with an arbitrarily small error, which further verifies the effectiveness of the neural network. Neural networks recently attracted a lot of attention of applied scientists, and were successful in fields such as image recognition and natural language processing because they are particular good at function approximation when the form of the target function is unknown. In the realm of dynamic portfolio analyses, Lin et al. (2006) first predicted portfolio covariance matrix with the Elman network and achieved the good estimation of the optimal mean-variance portfolio. More recently, Li and Forsyth (2019) proposed a neural network, representing the portfolio strategy at each rebalancing time, for a constrained defined contribution (DC) allocation problem. Chen and Ge (2021) introduced a differential equation-based method, where the value function with the Heston model is estimated by a deep neural network. In this paper, motivated by the lack of knowledge on the correct expression for the portfolio value function for unsolvable models, we approximated the optimal portfolio strategy for any given stochastic process model with a neural network fitting the value function. Successful fitting relies on a suitable network architecture that captures the connection between input and output variables, as well as reasonable activation functions. J. Risk Financial Manag. 2021,14, 322 3 of 18 We designed two architectures enriching an embedded quadratic-affine structure, and we considered three types of activation functions. Given the lack of closed-form solutions for SV 4/2 models, we used them as our toy examples in the implementations. In particular, we first implemented our methodology in the solvable case (i.e., GBM 4/2 with solvable MPR), so the accuracy and efficiency were demonstrated before it is applied to the unsolvable cases of: GBM 4/2 model with stochastic jumps, GBM 4/2 model with proportional instantaneous volatility MPR, and the OU 4/2 model. Furthermore, we numerically show which network architecture is preferable in each case. The paper is organized as follows. Section 2introduces the dynamic portfolio choice problem, and presents the neural network architectures, activation functions and parameter training details. The step-by-step algorithm of our methodology is provided in Section 3 . Sections 4and 5apply the methodology to the GBM 4/2 and the OU 4/2 models. Section 6concludes. 2. Problem Setting and Architectures of the Deep Learning Model We considered a frictionless market consisting of a money market account (cash, M ) and one stock ( S ). We assume the stock price follows a generalized diffusion process incorporating a one-dimensional state variable X . All the processes are defined on a complete probability space (Ω , F , P) with a right-continuous filtration {Ft}t∈[0,T] , summarized by the stochastic differential equations (SDE):            dMt Mt=r(Xt)dt dSt=Stθ(Xt,St)dt +Stσ(Xt,St)dBt+St−µNdNt dXt=a(Xt)dt +b(Xt)dBX t <dBt,dBX t>=ρdt. (1) Bt and BX t are Brownian motions with correlation ρ . r(Xt) is the interest rate, θ(Xt , St) and σ(Xt , St) are the drift and diffusion coefficients for the stock price. a(Xt) and b(Xt) are measurable functions of state variable Xt . Nt is a pure-jump process independent of Bt and BX t with stochastic intensity λNXt for constant λN> 0, and µN>− 1 denotes the jump size. We consider an investor with risk preference represented by a constant relative risk aversion (CRRA) utility: U(W) = W1−γ 1−γ. (2) Investors can adjust their allocation at a predetermined set of rebalancing times ( 0, ∆t , 2 ∆t , ..., T−∆t) . The investors wish to derive a portfolio strategy π (percentage of wealth allocated to the stock) that maximizes their expected utility of terminal wealth, in other words, E(U(WT)) . The value function, representing the investor’s conditional expected utility, has the following representation: V(t,W,S,X) = max πs≥t E(U(WT)|t,W,S,X)=W1−γ 1−γf(t,S,X). (3) The value function is separated into a wealth factor W1−γ 1−γ and a state variable function f . The NNMC estimates the state variable function f with a neural network model NN and computes the optimal strategy π∗ twith the Bellman principle. 2.1. Architectures of the Deep Learning Model In this section, we present two neural network architectures to fit the value function. According to the separable property of the value function shown in (3) , the only unknown component is the state variable function f , which is therefore the target function for the J. Risk Financial Manag. 2021,14, 322 4 of 18 neural network. The architectures of the networks are built around exponential polynomial functions, which are the most common form of solvable investor’s value functions and used in the PAMC method (see Zhu et al. (2020)). This property of proposed networks ensures that the new method generalizes PAMC. The neural network is expected to achieve a better fit than a polynomial regression if the true state variable function is significantly different from the exponential polynomial function. Furthermore, we designed an initialization method for networks, which is better than a random initialization in terms of portfolio value function fitting. 2.1.1. Sum of Exponential Network We first introduced the sum of the exponential polynomial neural network (SEN), as illustrated in Figure 1. The amount of input depends on the number of state variables. For simplicity, we took two inputs as an example. The first hidden layer computes the monomial of inputs. The second hidden layer obtains the linear combinations of the neuron in the first layer, where the weights are fitted in NNMC. An exponential activation function is applied to the second layer. The final output calculates a linear combination of exponential polynomials, so the exponential polynomial is a specific case of this neural network. Figure 1. Sum of exponential network (SEN). We denote the sum of exponential network by NNSEN ; the proposition next states the estimation of the corresponding optimal allocation. Proposition 1. Given the SEN approximation of the value function at the next rebalancing time t+∆t, (i.e., NNSEN[t+∆t,St,Xt]), the optimal strategy at time t is given by πSEN t=arg max π V(t,Wt,πt,St,Xt)(4) which is the solution of: f2(t,Wt,St,Xt) + f1(t,Wt,St,Xt)πt+NNSEN(t+∆t,St(1+µN),Xt)λNXtµN(1+πtµN)−γ=0, (5) J. Risk Financial Manag. 2021,14, 322 5 of 18 where: f1(t,Wt,St,Xt) = −γNNSEN(t+∆t,Wt,St,Xt)σ2(Xt,St) f2(t,Wt,St,Xt) = NNSEN(t+∆t,Wt,St,Xt)(θ(Xt,St)−r(Xt))) +∂NNSEN(t+∆t,Wt,St,Xt) ∂St Stσ2(Xt,St) +∂NNSEN(t+∆t,Wt,St,Xt) ∂Xt σ(Xt,St)b(Xt)ρ. (6) Notably, πSEN t=−f2(t,Wt,St,Xt) f1(t,Wt,St,Xt) when St follows a diffusion process, i.e., λN= 0. πSEN t= 1 µN((−f2(t,Wt,St,Xt) NNSEN(t+∆t,St,Xt)λNXtµN)−1 γ−1)when Stfollows a jump process, i.e., σ(Xt,St) = 0. Proof. It follows similarly to Theorem 1 in Zhu and Escobar-Anel (2020). According to the Bellman principle: V(t,Wt,St,Xt) = max πt Et(V(t+∆t,Wt+∆t,St+∆t,Xt+∆t)|Wt,St,Xt). (7) We substitute V(t+∆t , Wt+∆t , St+∆t , Xt+∆t) with W1−γ 1−γNNSEN(t+∆t , Wt+∆t , St+∆t , Xt+∆t) and expand the right hand side of the equation with respect to W , S and X , then V(t , Wt , St , Xt) is written as a function of strategy πt . Equation (5) is obtained with the first order condition. 2.1.2. Improving Exponential Network The architecture of an improving exponential network (IEN) is exhibited in Figure 2. Figure 2. Improving exponential polynomial. The target function of IEN is the log of the state variable function f (i.e., ln f ). The neural network consists of three parts. Node 1 is a polynomial with the output denoted by V1 . Node 2 is an artificial neural network with an arbitrary number of hidden layers and neurons; we denoted its output by V2 . Node 3 is a single-layer network with a Sigmoid function which computes a proportion p∈[ 0, 1 ] . The final output is the weighted average of the first two nodes pV1+ ( 1 −p)V2 . The second node is the complement to the exponential polynomial function. Moreover, the similarity between the true value function and the exponential polynomial function is measured by p , which is fitted into the NNMC methodology. Therefore, the network automatically adjusts the weights on the exponential J. Risk Financial Manag. 2021,14, 322 6 of 18 polynomial function and its supplement according to the generated data. Finally, the state variable function fis computed as f=epv1+(1−p)v2= (ev1)p×(ev2)1−p, (8) which is the geometric weighted average of nodes 1 and 2. Letting NNIEN denote the IEN, the estimation of the optimal strategy is given in the next proposition. Proposition 2. Given the IEN approximation of the log value function at time t+∆t (i.e., NNIEN[t+∆t,St,Xt]), the optimal strategy at time t is given by πIEN t=arg max π V(t,Wt,St,Xt)(9) which is the solution of: (f2(t,Wt,St,Xt) + f1(t,Wt,St,Xt)πt) + λNXtexpNNIEN(t+∆t,St(1+µN),Xt)µN(1+πtµN)−γ=0, (10) where: f1(t,Wt,St,Xt) = −γexpNNIEN(t+∆t,St,Xt)σ2(Xt,St) f2(t,Wt,St,Xt) = expNNIEN(t+∆t,St,Xt)(θ(Xt,St)−r(Xt))) +∂NNIEN(t+∆t,Wt,St,Xt) ∂StexpNNIEN(t+∆t,St,Xt)Stσ2(Xt,St) +∂NNIEN(t+∆t,Wt,St,Xt) ∂XtexpNNIEN(t+∆t,St,Xt)σ(Xt,St)b(Xt)ρ. (11) Notably, πIEN t=−f2(t,Wt,St,Xt) f1(t,Wt,St,Xt) when St follows a diffusion process, in other words, λN=0 . πIEN t=1 µN((−f2(t,Wt,St,Xt) exp(NNIEN(t+∆t,St,Xt))λNXtµN)−1 γ− 1 ) when St follows a jump process (i.e., σ[Xt,St] = 0). Proof. The proof follows similarly to Proposition 1. 2.2. Initialization, Stopping Criterion and Activation Function In this section, we disclose more details on training the neural networks. The initialization of weights is the first step of network training, which may significantly impact the goodness of fit. A good initialization prevents the network’s weights from converging to a local minimum and avoids slow convergence. Random initialization is mostly used as the interpretability of the network is usually weak. In contrast, both the SEN and the IEN are extensions of an exponential polynomial function; we suggest taking advantage of the results from the polynomial regression. Hence, the neural network searches the minimum near the exponential polynomial function used in the PAMC ensuring consistency. The polynomial regression initialization achieves superior results to the random initialization. The coefficients of the exponential polynomial were first obtained with a regression model. The output of the SEN is a linear combination of exponential polynomial functions N ∑ i=1 aiexp(Pi n(x , y)) + b , we substitute the coefficients from polynomial regression into P1 n(x , y) and set a1= 1, a2=a3= ... =an=b= 0. For the initialization of the IEN, we substitute the coefficients into the first node and artificially make p=0. The training process minimizes the mean squared error (MSE) between the network’s output and the simulated expected utility, and the sample data are split into a training set and a test set to reduce the overfitting problem. Adam is a back-propagation algorithm that combines the best properties of the AdaGrad and RMSProp algorithms to handle sparse J. Risk Financial Manag. 2021,14, 322 7 of 18 gradients on noisy problems and provides excellent convergence speed. We applied the Adam on the training set for updating the network’s weights, and the test set MSE was computed and subsequently recorded. The test set MSE was expected to be convergent, so the training process was finished when the difference between the moving average of the recent 100 test set MSEs and the most recent test set MSE was less than a predetermined threshold, which was set at 0.00001 in the implementation. The number of exponential polynomials is a hyperparameter in the SEN. We let the SEN be a sum of two exponential polynomial functions for simplicity. Node 2 in the IEN is an artificial neural network, which complements node 1 when the value function significantly deviates from an exponential polynomial function. The number of hidden layers and neurons, as well as the activation function of node 2, are freely determined before fitting the value function. We assume node 2 is a single layer network with 10 neurons and we implement several functions for comparison purposes, such as the logistic (sigmoid): f(x) = 1 1+e−x, (12) the Rectified linear unit (ReLU): f(x) = (0 if x≤0 xif x>0, (13) and the Exponential linear unit (ELU): f(x) = (0 if x≤0 ex−1 if x>0(14) 3. Notation and Algorithm of the Methodology In this section, we clarify the notation and the step-by-step algorithm. Table 1displays a summary of the notation. Table 1. Notation for NNMC is listed here. Notation Meaning Bm tBrownian motion at time tin mth simulated path Sm tStock price at time tin mth simulated path Xm tOther state variable such as interest rate or volatility nrNumber of simulated paths NNumber of simulation to compute expected utility for a given set (W0,Sm t,Xm t) ˆ Wm,n t+∆t(πm)A simulated wealth level at t+∆tgiven the wealth, allocation and other state variables at tare W0,πmand Xm t ˆ Sm,n t+∆tA simulated stock price at t+∆tgiven Sm t ˆ Xm,n t+∆tA simulated state variable at t+∆tgive Xm t V(t,W,S,X)Value function at time tgiven wealth W, stock price Sand state variable X NN(t,X,S)The neural network used to fit f(t,St,Xt)or ln [f(t,St,Xt)] ˆ vmEstimation of f(t,Sm t,Xm t)or ln [f(t,Sm t,Xm t)] πm,n sOptimal strategy at time sgiven wealth, stock price and other state variables are ˆ Wm,n s,ˆ Sm,n sand ˆ Xsm,n ˆ V(0, W0,S0,X0)Estimation of expected utility at time 0. J. Risk Financial Manag. 2021,14, 322 8 of 18 Algorithm We first generated the paths of the stock price Sm t and state variable Xm t . The method starts from t=T−∆t (i.e., the last rebalancing time before the terminal). We computed the optimal strategy πm T−∆t given W0 , Sm T−∆t , Xm T−∆t using the Equation (5) or Equation (10) . Then, ˆ vm is obtained through simulation, which estimates f(T−∆t , Sm T−∆t , Xm T−∆t) when using SEN and ln [f(T−∆t,Sm T−∆t,Xm T−∆t)] when using IEN. The network NN(T−∆t , X , S) , approximating the state variable function, is trained with the input ( Xm T−∆t , Sm T−∆t ) and output ˆ vm . We conduct a similar procedure at each rebalancing point and recursively approximate the value function and optimal strategy until the inception of the portfolio. To evaluate the expected utility, we regenerated the paths of stock price and state variables. The path-wise optimal strategy was computed from NN(t , X , S) , so the optimal terminal wealth is easy to obtain. The average of the utility of optimal terminal wealth approximates the expected utility. Algorithms 1and 2present the pseudo code for NNMC using SEN and IEN, respectively. Simulation variance reduction methods, such as antithetic variates, could be incorporated into both algorithms to reduce the standard error of estimated expected utility. Algorithm 1: NNMC-SEN Input: S0,W0,X0 Output: Optimal trading strategy π∗ 0and expected utility ˆ V(0, W0,S0,X0) 1initialization; 2Generating nrpaths of Bm t,Sm t,Xm tf or m =1...nr; 3while t=T−∆tdo 4Compute optimal allocation πm T−∆twith Equation (5) ; 5Simulate wealth ˆ Wm,n T(πm T−∆t)given W0,Sm T−∆t,πm T−∆tand Xm T−∆tat T−∆t f or n =1...N; 6Compute ˆ vm=1 N N ∑ n=1 U(ˆ Wm,n T(πm T−∆t)) ×1−γ W1−γ 0 f or m =1...nr; 7 Train a network with input ( Xm T−∆t , Sm T−∆t ) and output ˆ vm . Denote the network by NN(T−∆t,X,S) 8for t=T−2∆t to ∆tdo 9 Compute optimal allocation πm t with NN(t+∆t , X , S) and Equation (5) given W0,Sm t, and Xm t; 10 Simulate wealth ˆ Wm,n t+∆t(πm t),ˆ Sm,n t+∆tand ˆ Xm,n t+∆tgiven W0,Sm t,πm tand Xm tat time t f or n =1...N; 11 Compute ˆ vm= [ 1 N N ∑ n=1 (Wm,n t+∆t(πm t))1−γNN(t+∆t,ˆ Xm,n t+∆t,ˆ Sm,n t+∆t)] ×1 W1−γ 0 f or m=1...nr; 12 Train a new network with input ( Xm T−∆t , Sm T−∆t ) and output ˆ vm and denote it by NN(t,X,S); 13 while t=0do 14 Compute π∗ 0with with NN(∆t,X,S)and Equation (5); 15 Generate new paths of Sz t,Xz tf or z =1...N0, use the estimation of value function NN(t,X,S)to compute πz tand Wz T. 16 The expected utility is, ˆ V(0, W0,S0,X0) = 1 N0 N0 ∑ n=1 U(Wz T) 17 return π∗ 0,ˆ V(0, W0,S0,X0) J. Risk Financial Manag. 2021,14, 322 15 of 18 (a) Expected utility (b) CER Figure 5. St follows the 4/2 model with a market price of risk λS(a√Xt+b √Xt) , where ( a ) shows the expected utilities obtained with theoretical results and approximation methods versus investment horizon T; and (b) shows the CERs versus investment horizon Tgiven γ=2. 5. Application to the OU 4/2 Model Motivated by the 4/2 stochastic volatility model and mean-reverting price pattern popular among various asset classes (e.g., commodities, exchange rates, volatility indexes), Escobar-Anel and Gong (2020) defined an Ornstein–Uhlenbeck 4/2 (OU 4/2) stochastic volatility model for volatility index option and commodity option valuation. Equation (22) presents the dynamics involved in the OU 4/2 model, which is a specific case of (1) given θ(Xt , St) = (LS+ (λS−1 2)(aS√Xt+bS √Xt)2−βSln St) , σ(Xt , St) = (aS√Xt+bS √Xt) , a(Xt) = κX(θX−Xt) and b(Xt) = σX√Xt . The parameters used in this section are reported in Table 6, which is estimated from the data of gold Exchange-traded fund (ETF) and the volatility index of gold ETF in Escobar-Anel and Gong (2020). There are two state variables in the OU 4/2 model; hence, the input in both the SEN and the IEN are 2. Furthermore, the degree of polynomial in PAMC and NNMC is 2:            dMt Mt=rdt dSt St= (LS+λS(aS√Xt+bS √Xt)2−βSln St)dt + (aS√Xt+bS √Xt)dBt, dXt=κX(θX−Xt)dt +σX√XtdBX t <dBt,dBX t>=ρdt. (22) Table 6. Parameter value for the OU 4/2 model. Parameter Value Parameter Value T1X00.04 r0.05 λS0.572 ∆re t1 60 ∆si t1 60 S0120.0 M01.0 W01nr100 κX4.7937 θX0.0395 σX0.2873 aS1 bS0.002 ρ−0.08 L3.7672 βS0.78 N2000 N0200,000 J. Risk Financial Manag. 2021,14, 322 16 of 18 SEN performs worse than IEN when fitting the value function with the OU 4/2 model. Sometimes, SEN significantly deviates from the true value function, which results in poor portfolio performances and the occurrence of negative terminal wealth. Therefore, we excluded the results from NNMC-SEN in this section. Table 7compares the optimal allocation, expected utility and CER obtained for the OU 4/2 model. PAMC and NNMCIEN produce similar optimal allocations, both outperforming NNMC-SEN. Furthermore, we also estimated the standard deviation of expected utility and CER, which demonstrates that NNMC leads to a less volatile estimation of expected utility and CER than PAMC in most cases. In contrast to the results for the 4/2 model, IEN is more efficient than SEN. We conclude that IEN is suitable for the model with a complex structure and multiple state variables. The expected utility and CER as a function of the maturity T when γ= 2 is plotted in Figure 6. Both the expected utility and CER increase with T . The expected utility and CER obtained from PAMC and NNMC-IEN visually overlap and are slightly higher than that of NNMC-SEN. Moreover, the selection of activation function in IEN makes little difference. Table 7. Results for the OU 4/2 model. We report the estimation of optimal weights, expected utility and CER obtained via approximations for different levels of risk aversion γ . The standard deviation of estimated expected utility and CER from 100 runs is provided in parentheses. γ=2.0 γ=4.0 γ=6.0 γ=8.0 γ=10.0 PAMC Weights (πPAMC 0)0.068 0.026 0.015 0.010 0.008 Expected utility (VPAMC 0)−0.888 (0.0006) −0.255 (0.0003) −0.136 (0.0002) −0.087 (0.0002) −0.061 (0.0001) CER (%) 12.65 (0.073) 9.28 (0.047) 8.00 (0.035) 7.32 (0.028) 6.90 (0.024) Computational time (seconds) 103.9 104.6 104.4 104.5 104.3 NNMC-SEN Weights (πSEN 0)0.134 0.056 0.042 0.040 0.029 Expected utility (VSEN 0)−0.888 (0.0006) −0.256 (0.0003) −0.136 (0.0002) -0.087 (0.0001) −0.061 (0.0001) CER (%) 12.62 (0.076) 9.26 (0.045) 7.97 (0.032) 7.29 (0.025) 6.87 (0.020) Computational time (seconds) 439.5 477.5 434.3 446.9 449.7 NNMC-IEN (ReLU) Weights (πIEN ReLU 0)0.070 0.028 0.016 0.011 0.007 Expected utility (VIEN ReLU 0)−0.888 (0.0006) −0.255 (0.0003) −0.136 (0.0002) −0.087 (0.0002) −0.061 (0.0001) CER (%) 12.65 (0.072) 9.29 (0.045) 8.00 (0.033) 7.32 (0.026) 6.90 (0.022) Computational time (seconds) 190.3 190.6 190.4 187.8 185.1 NNMC-IEN (sigmoid) Weights (πIEN sigmoid 0)0.067 0.026 0.015 0.010 0.007 Expected utility (VIEN sigmoid 0)−0.888 (0.0006) −0.255 (0.0003) −0.136 (0.0002) −0.087 (0.0001) −0.061 (0.0001) CER (%) 12.65 (0.072) 9.28 (0.044) 8.00 (0.033) 7.32 (0.026) 6.90 (0.022) Computational time (seconds) 185.7 186.0 185.2 181.4 181.9 NNMC-IEN (ELU) Weights (πIEN ELU 0)0.072 0.031 0.015 0.010 0.008 Expected utility (VIEN ELU 0)−0.888 (0.0006) −0.255 (0.0003) −0.136 (0.0002) −0.087 (0.0001) −0.061 (0.0001) CER (%) 12.65 (0.072) 9.28 (0.044) 8.00 (0.033) 7.32 (0.026) 6.90 (0.022) Computational time (seconds) 185.6 184.1 188.1 195.1 193.7 J. Risk Financial Manag. 2021,14, 322 17 of 18 (a) Expected utility (b) CER Figure 6. St follows the OU 4/2 model, where ( a ) shows the expected utilities obtained via approximation methods versus investment horizon T ; and ( b ) shows the CERs versus investment horizon T given γ=2. 6. Conclusions This paper investigated fitting the value function in an expected utility, dynamic portfolio choice using a deep learning model. We proposed two architectures for the neural network, which extends the broadest solvable family of value functions (i.e., the exponential polynomial function). We measured the accuracy and efficiency of various types of NNMC methods on the 4/2 model and the OU 4/2 model. The difference in optimal allocation, expected utility and CER is insignificant when the stock price follows the 4/2 model. The embedded PAMC is superior to NNMC due to the lower parametric space, hence its efficiency. Furthermore, when considering the OU 4/2 model, NNMC-SEN is inferior to a polynomial regression (PAMC) and to the NNMC-IEN in terms of expected utility and CER. In summary, NNMC benefits from the popular exponential polynomial representation (embedded PAMC method) to propose a network architecture flexible enough to reach beyond affine models. Although the best setting, NNMC-IEN (ELU), is not as efficient as PAMC, neural networks demonstrate the way to tackle more advanced models along the lines of Markov switching, Lévy processes and fractional Brownian processes. Author Contributions: The authors contributed equally. All authors have read and agreed to the published version of the manuscript. Funding: This research was funded by NSERC, grant number RGPIN-2020-05068. Conflicts of Interest: The authors declare no conflict of interest. Note 1∆re t is the portfolio rebalancing interval, 1 ∆re t indicates the rebalancing frequency. The Euler method with step size ∆si t is applied in generating the stock price and states variables. References Brandt, Michael W., Amit Goyal, Pedro Santa-Clara, and Jonathan R. Stroud. 2005. A simulation approach to dynamic portfolio choice with an application to learning about return predictability. The Review of Financial Studies 18: 831–73. [CrossRef] Chen, Shun, and Lei Ge. 2021. A learning-based strategy for portfolio selection. International Review of Economics & Finance 71: 936–42. Cheng, Yuyang, and Marcos Escobar-Anel. 2021. Optimal investment strategy in the family of 4/2 stochastic volatility models. Quantitative Finance 1–29. [CrossRef] Chiu, Mei Choi, and Hoi Ying Wong. 2013. Optimal investment for an insurer with cointegrated assets: Crra utility. Insurance: Mathematics and Economics 52: 52–64. [CrossRef] Cong, Fei, and Cornelis W. Oosterlee. 2017. Accurate and robust numerical methods for the dynamic portfolio management problem. Computational Economics 49: 433–58. [CrossRef] Cvitani´c, Jakša, Levon Goukasian, and Fernando Zapatero. 2003. Monte carlo computation of optimal portfolios in complete markets. Journal of Economic Dynamics and Control 27: 971–86. [CrossRef] Cybenko, George. 1989. Approximation by superpositions of a sigmoidal function. Mathematics of Control, Signals and Systems 2: 303–14. [CrossRef] J. Risk Financial Manag. 2021,14, 322 18 of 18 Detemple, Jerome B., Ren Garcia, and Marcel Rindisbacher. 2003. A monte carlo method for optimal portfolios. The Journal of Finance 58: 401–46. [CrossRef] Escobar, Marcos, Daniela Neykova, and Rudi Zagst. 2017. Hara utility maximization in a markov-switching bond–stock market. Quantitative Finance 17: 1715–33. [CrossRef] Escobar-Anel, Marcos, and Zhenxian Gong. 2020. The mean-reverting 4/2 stochastic volatility model: Properties and financial applications. Applied Stochastic Models in Business and Industry 36: 836–56. [CrossRef] Flor, Christian Riis, and Linda Sandris Larsen. 2014. Robust portfolio choice with stochastic interest rates. Annals of Finance 10: 243–65. [CrossRef] Grasselli, Martino. 2017. The 4/2 stochastic volatility model: A unified approach for the heston and the 3/2 model. Mathematical Finance 27: 1013–34. [CrossRef] Heston, Steven L. 1993. A closed-form solution for options with stochastic volatility with applications to bond and currency options. The Review of Financial Studies 6: 327–43. [CrossRef] Jain, Shashi, and Cornelis W. Oosterlee. 2015. The stochastic grid bundling method: Efficient pricing of bermudan options and their greeks. Applied Mathematics and Computation 269: 412–31. [CrossRef] Karatzas, Ioannis, John P. Lehoczky, and Steven E. Shreve. 1987. Optimal portfolio and consumption decisions for a “small investor” on a finite horizon. SIAM Journal on Control and Optimization 25: 1557–86. [CrossRef] Kraft, Holger. 2005. Optimal portfolios and heston’s stochastic volatility model: An explicit solution for power utility. Quantitative Finance 5: 303–13. [CrossRef] Li, Yuying, and Peter A. Forsyth. 2019. A data-driven neural network approach to optimal asset allocation for target based defined contribution pension plans. Insurance: Mathematics and Economics 86: 189–204. [CrossRef] Lin, Chi-Ming, Jih-Jeng Huang, Mitsuo Gen, and Gwo-Hshiung Tzeng. 2006. Recurrent neural network for dynamic portfolio selection. Applied Mathematics and Computation 175: 1139–46. [CrossRef] Linnainmaa, Seppo. 1970. The Representation of the Cumulative Rounding Error of an Algorithm as a Taylor Expansion of the Local Rounding Errors. Master’s thesis, University of Helsinki, Helsinki, Finland, pp. 6–7. (In Finnish) Liu, Jun. 2006. Portfolio selection in stochastic environments. The Review of Financial Studies 20: 1–39. [CrossRef] Liu, Jun, and Jun Pan. 2003. Dynamic derivative strategies. Journal of Financial Economics 69: 401–30. [CrossRef] Longstaff, Francis A., and Eduardo S. Schwartz. 2001. Valuing american options by simulation: a simple least-squares approach. The Review of Financial Studies 14: 113–47. [CrossRef] McCulloch, Warren S., and Walter Pitts. 1943. A logical calculus of the ideas immanent in nervous activity. The Bulletin of Mathematical Biophysics 5: 115–33. [CrossRef] Merton, Robert C. 1969. Lifetime portfolio selection under uncertainty: The continuous-time case. The Review of Economics and Statistics 51: 247–57. [CrossRef] Rumelhart, David E., Geoffrey E. Hinton, and Ronald J. Williams. 1986. Learning representations by back-propagating errors. Nature 323: 533–36. [CrossRef] Zhu, Yichen, and Marcos Escobar-Anel. 2020. Polynomial affine approach to hara utility maximization with applications to ornsteinuhlenbeck 4/2 models. Applied Mathematics and Computation, submitted. Zhu, Yichen, Marcos Escobar-Anel, and Matt Davison. 2020. A polynomial-affine approximation for dynamic portfolio choice. Computational Economics, submitted.