Adaptive Portfolio Optimization using Robbins–Monro Stochastic Approximation 1 Adaptive Portfolio Optimization using Robbins–Monro Stochastic Approximation A Dynamic Framework for Volatility-Aware Financial Decision Support Systems Shashvat Dubey, Ankit Ghosh, Samarth Agarwal Department of Computational Intelligence, School of Artificial Intelligence, SRM Institute of Science and Technology, Chennai, India. November 2025 Abstract This research presents an adaptive stochastic portfolio optimization framework based on the Robbins–Monro stochastic approximation algorithm. The model dynamically adjusts portfolio weights in response to real-time volatility fluctuations derived from Exponentially Weighted Moving Averages (EWMA). Two adaptive parameters—risk penalty (λ t ) and learning rate (α t )—evolve continuously according to market variance, allowing the system to balance profitability and stability. Using ten years of Yahoo Finance data (2014–2024) across major equities, the model demonstrates strong convergence, interpretability, and resilience during periods of high volatility. The proposed system highlights the potential of volatility-aware stochastic learning as a foundation for real-time financial decision-support systems. Keywords: Robbins–Monro Algorithm, Adaptive Portfolio Optimization, EWMA, Volatility Modeling, Financial Decision Systems, Stochastic Learning ∗
[email protected] †
[email protected] ‡
[email protected]
Adaptive Portfolio Optimization using Robbins–Monro Stochastic Approximation 2 1 Introduction 1.1 Background and Motivation Financial markets exhibit high levels of uncertainty and frequent shifts in volatility patterns. Classical portfolio optimization methods, such as Markowitz’s Mean–Variance model, assume a stationary environment where covariances and expected returns remain constant. This assumption rarely holds true in practice, leading to suboptimal portfolio allocations when market regimes shift rapidly. 1.2 Problem Definition The challenge lies in designing a portfolio optimizer that continuously adapts to new market information without retraining from scratch. Traditional batch-learning approaches lack the flexibility to adjust parameters online. The objective of this study is to employ a stochastic optimization approach capable of learning incrementally from data streams. 1.3 Motivation for Stochastic Modeling The Robbins–Monro stochastic approximation method offers a mathematically rigorous way to estimate optimal parameters in noisy, dynamic environments. Its incremental learning nature allows weight updates in response to new market conditions, forming the core of our adaptive framework. 1.4 Objectives of the Study • Integrate EWMA volatility tracking into a stochastic gradient-based learning system. • Develop an adaptive rule for real-time adjustment of risk and learning rate parameters. • Benchmark the model against a static equal-weight portfolio to evaluate adaptability and stability. 2 Literature Survey 2.1 Classical Portfolio Optimization Markowitz (1952) first introduced the Mean–Variance framework [6], which remains the foundation for portfolio theory. However, the assumption of fixed covariances limits its applicability in non-stationary markets.
Adaptive Portfolio Optimization using Robbins–Monro Stochastic Approximation 3 2.2 Volatility Modeling and EWMA Volatility estimation is a critical component of adaptive portfolio design. The Exponentially Weighted Moving Average (EWMA) model [3] provides an efficient realtime estimate of market variance, responding more strongly to recent data while smoothing historical fluctuations. 2.3 Adaptive and Stochastic Methods The Robbins–Monro algorithm [1] forms the basis for numerous adaptive algorithms used in signal processing and online learning. Recent advancements integrate stochastic approximation into reinforcement learning frameworks [5], making it suitable for real-time decision-making in financial contexts. 3 Research Methodology 3.1 Mathematical Formulation The optimization goal is to minimize a time-varying loss function: L t (w t ) = − E[R t (w t )] + λ t V ar(R t (w t )) + γC t (w t ) (1) where: • E[R t ]: Expected portfolio return estimated via EWMA mean. • V ar(R t ): Portfolio variance estimated via EWMA volatility. • C t (w t ): Constraint term enforcing Σ w t = 1. Adaptive parameters evolve dynamically as: Portfolio weights update following the Robbins–Monro rule: w t +1 = w t − α t ∇ L t ( w t ) (3) Λ t= λ0 + kλσt ,
Adaptive Portfolio Optimization using Robbins–Monro Stochastic Approximation 4 4 Experimental Setup 4.1 Data Description The dataset consists of daily adjusted closing prices from Yahoo Finance for ten major equities (AAPL, MSFT, GOOGL, AMZN, META, JPM, TSLA, XOM, NVDA, PG) spanning 2014–2025. The data were processed to compute log returns, rolling statistics, and EWMA volatility using a half-life of 30 days. 4.2 Implementation Pipeline The experiment was structured in three stages: 1. Training (2014–2024): The adaptive model learns online weight adjustments based on historical returns and volatility patterns. 2. Testing (Jan–Feb 2025): The model applies learned dynamics to unseen market conditions. 3. Performance Evaluation: Metrics such as annualized return, volatility, Sharpe ratio, and maximum drawdown are computed. 4.3 Graph 1 – Adaptive Parameters Over Time Description: This figure illustrates the evolution of λ t (risk penalty) and α t (learning rate) over 10 years. Spikes in 2020 correspond to increased market uncertainty, where the model automatically increases λ t and reduces α t to ensure stable learning.
Adaptive Portfolio Optimization using Robbins–Monro Stochastic Approximation 5 5 Results and Discussion 5.1 Adaptive Portfolio Dynamics Graph 2 – Portfolio Weights Over Time: This visualization shows how portfolio allocations change daily as the model responds to volatility and expected returns. Smooth variations indicate stable convergence, while sharper shifts represent adaptive reactions to significant market movements. 5.2 Out-of-Sample Evaluation During the test period (Jan–Feb 2025), the adaptive model demonstrated realistic behavior under mild downturns: • Annualized Return: 22.8% • Volatility: 22.9% • Sharpe Ratio: 0.9956 • Max Drawdown: –43.06% While performance slightly dipped under low market momentum, the model’s adaptability maintained robustness compared to the equal-weight baseline.
Adaptive Portfolio Optimization using Robbins–Monro Stochastic Approximation 6 5.3 Interpreting Parameter Coupling Graph 4 – Volatility–Parameter Relationship: This figure demonstrates how volatility (σ t ) drives adjustments in λ t and α t . Periods of high volatility increase the system’s risk sensitivity (λ t ) and reduce its learning rate (α t ), effectively creating a self-regulating feedback mechanism. 6 Findings and Insights The results validate that the Robbins–Monro algorithm is capable of producing consistent adaptive learning in highly non-stationary market conditions. Key takeaways include: • The model successfully learned to reduce exposure during volatility spikes. • Parameter evolution mirrored real-world investor sentiment patterns. • Despite minor underperformance in calm markets, the adaptive strategy achieved superior stability metrics.
Adaptive Portfolio Optimization using Robbins–Monro Stochastic Approximation 7 These behaviors collectively support the hypothesis that stochastic approximation provides a mathematically grounded yet practically flexible base for continuous financial decision-making. 7 Conclusion and Future Work 7.1 Summary of Contributions This work developed a volatility-aware portfolio optimization algorithm based on Robbins–Monro stochastic approximation. By integrating EWMA volatility tracking and adaptive learning rules, the model achieves self-regulated adjustment of portfolio weights under real-time market uncertainty. 7.2 Limitations and Future Directions While the framework adapts effectively to variance shifts, it currently assumes no transaction costs and ignores cross-asset correlations beyond the EWMA window. Future extensions will explore: • Reinforcement Learning-based reward structures for adaptive trading. • Integration of GARCH volatility models for higher precision. • Real-time market execution with streaming data APIs. Acknowledgments The authors thank the Department of Computational Intelligence, SRM Institute of Science and Technology, for their guidance and computational support. Data was sourced from Yahoo Finance. No conflicts of interest are declared. References [1] Robbins, H. and Monro, S., “A Stochastic Approximation Method,” The Annals of Mathematical Statistics, Vol. 22, No. 3, 1951, pp. 400–407. [2] Bollerslev, T., “Generalized Autoregressive Conditional Heteroskedasticity,” Journal of Econometrics, Vol. 31, No. 3, 1986, pp. 307–327. [3] J.P. Morgan, “RiskMetrics Technical Document,” 4th Edition, New York, 1996. [4] Benveniste, A., M ´ e tivier, M., and Priouret, P., Adaptive Algorithms and Stochastic
Adaptive Portfolio Optimization using Robbins–Monro Stochastic Approximation 8 Approximations, Springer Science & Business Media, 2012. [5] Sutton, R. S. and Barto, A. G., Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018. [6] Markowitz, H., “Portfolio Selection,” The Journal of Finance, Vol. 7, No. 1, 1952, pp. 77–91.