scieee AI-readable full text Open interactive document viewer

Comparative Analysis of Machine Learning Models for Predicting Student Stress Levels: A Multi-Algorithm Approach

Mehrez Ben nasr and Sirine Ben Othman

Abstract

ABSTRACT Student stress has emerged as a critical health concern in academic institutions, with significant implications for academic performance, mental health, and overall well-being. This study compares seven machine learning and statistical modeling approaches to identify determinants of student stress and establish optimal predictive models. Using data from 520 students, we employed linear regression, Random Forest, XGBoost, Support Vector Machines, k-Nearest Neighbors, artificial neural networks, and decision tree algorithms to model stress levels as a function of five key variables: sleep quality, headache frequency, academic performance, study load, and extra-curricular activities. Results demonstrate substantial superiority of non-linear models, with k-NN and XGBoost reducing prediction error by 71-75% compared to linear regression. Study load emerged as the dominant stress determinant (β = 0.3833, p < 2×10⁻¹⁶), accounting for 30.15% of predictive gain in XGBoost models. However, only 17.79% of stress variance was explained by these five variables, indicating multifactorial etiology requiring integration of psychological and environmental factors. We recommend prioritization of study load reduction and implementation of k-NN or XGBoost models for early identification of at-risk students. These findings have significant implications for institutional policy development and student mental health intervention strategies. Keywords: Student stress, Machine learning, Predictive modeling, Linear regression, XGBoost, k-NN, Academic burden, Mental health

Full text

International Journal of Advanced Scientific and Technical Research ISSN 2249-9954 Available online on http://www.rspublication.com/ijst/index.html volume 15, No. 6, 2025 DOI: 10.5281/zenodo.17649330 Original Article ©2025 RS Publication, rspubl[email protected]om 77 Comparative Analysis of Machine Learning Models for Predicting Student Stress Levels: A Multi-Algorithm Approach Mehrez Ben Nasr Certified Actuary FTUSA, Tunis, Tunisia, [email protected] Sirine Ben Othman Psychiatry Department, Nabeul Hospital, Nabeul, Tunisia, [email protected] ARTICLE INFO ABSTRACT ©2025 RS Publication Paper ID: IJASTR691CAB06D8EB6 Received: 2025-10-20 Published: 2025-11-19 DOI: https://dx.doi.org/ 10.5281/zenodo.1764 9330 Page No: 77-86 Student stress has emerged as a critical health concern in academic institutions, with significant implications for academic performance, mental health, and overall well-being. This study compares seven machine learning and statistical modeling approaches to identify determinants of student stress and establish optimal predictive models. Using data from 520 students, we employed linear regression, Random Forest, XGBoost, Support Vector Machines, k-Nearest Neighbors, artificial neural networks, and decision tree algorithms to model stress levels as a function of five key variables: sleep quality, headache frequency, academic performance, study load, and extra-curricular activities. Results demonstrate substantial superiority of non-linear models, with k-NN and XGBoost reducing prediction error by 71-75% compared to linear regression. Study load emerged as the dominant stress determinant (β = 0.3833, p < 2×10⁻¹⁶), accounting for 30.15% of predictive gain in XGBoost models. However, only 17.79% of stress variance was explained by these five variables, indicating multifactorial etiology requiring integration of psychological and environmental factors. We recommend prioritization of study load reduction and implementation of k-NN or XGBoost models for early identification of atrisk students. These findings have significant implications for institutional policy development and student mental health intervention strategies. Keywords: Student stress, Machine learning, Predictive modeling, Linear regression, XGBoost, k-NN, Academic burden, Mental health International Journal of Advanced Scientific and Technical Research Available online on http://www.rspublication.com/ijst/index.html ISSN 2249-9954 Cite This Paper: Mehrez Ben nasr and Sirine Ben Othman(2025). "Comparative Analysis of Machine Learning Models for Predicting Student Stress Levels: A MultiAlgorithm Approach". INTERNATIONAL JOURNAL OF ADVANCED SCIENTIFIC AND TECHNICAL RESEARCH (IJASTR), vol. 15, no. 6, 2025, pp. 77-86. DOI: https://dx.doi.org/10.5281/zenodo.17649330 International Journal of Advanced Scientific and Technical Research ISSN 2249-9954 Available online on http://www.rspublication.com/ijst/index.html volume 15, No. 6, 2025 DOI: 10.5281/zenodo.17649330 Original Article ©2025 RS Publication, rspubl[email protected]om 78 1. Introduction Mental health challenges among university students have reached epidemic proportions globally, with stress being identified as a primary contributor to poor academic outcomes, anxiety disorders, and compromised physical health {1,2}. The prevalence of psychological distress among students has increased substantially over the past two decades, with studies reporting stress levels affecting 40-60% of student populations across various countries {3}. Understanding the multifactorial determinants of student stress is therefore essential for developing targeted intervention strategies. Previous research has identified multiple factors contributing to academic stress, including heavy workload, poor sleep quality, academic performance concerns, and limited engagement in stressrelieving activities {4,5}. However, the relative importance of these factors and their non-linear interactions remain poorly characterized. Furthermore, traditional statistical approaches such as linear regression may inadequately capture complex relationships inherent in psychological phenomena {6}. Recent advances in machine learning have enabled more sophisticated modeling of complex behavioral outcomes {7,8}. Machine learning algorithms, including ensemble methods and nonparametric approaches, have demonstrated superior predictive performance compared to traditional statistical models in numerous health-related applications {9,10}. However, comparative analyses of multiple algorithms in the context of student stress prediction remain limited in the academic literature. The objective of this study was to: (1) identify the key determinants of student stress using linear regression analysis; (2) compare predictive performance across seven distinct modeling approaches; (3) evaluate variable importance using ensemble-based methods; and (4) provide recommendations for institutional intervention strategies and risk identification protocols. 2. Methods 2.1 Study Population and Data Collection This study analyzed cross-sectional data from 520 university students (mean age: not reported in available data). The dependent variable was student stress level (stress_level), operationalized on a continuous scale. Five predictor variables were examined: sleep quality (sleep_quality), headache frequency (headache_freq), academic performance (academic_perf), study load (study_load), and engagement in extra-curricular activities (extra_activities). 2.2 Statistical Methods 2.2.1 Linear Regression Model We employed multiple linear regression to quantify the relationship between predictor variables and stress levels:   =  +    +    +    +    +    +  International Journal of Advanced Scientific and Technical Research ISSN 2249-9954 Available online on http://www.rspublication.com/ijst/index.html volume 15, No. 6, 2025 DOI: 10.5281/zenodo.17649330 Original Article ©2025 RS Publication, rspubl[email protected]om 79 where   denotes stress level,   is the intercept,   are regression coefficients,   are predictor variables, and   ∼0,   represents the error term for observation  (n=520). Statistical significance was assessed using t-tests with Bonferroni correction for multiple comparisons. 2.2.2 Performance Metrics Model prediction accuracy was evaluated using Root Mean Square Error (RMSE): RMSE =1  −      Additional metrics included R-squared (  ), Mean Absolute Error (MAE), and cross-validated performance estimates. 2.2.3 Machine Learning Algorithms k-Nearest Neighbors (k-NN): Predictions were generated using the average of k nearest neighbors in feature space: !=1"   ∈$ % & Optimal k was selected via 5-fold cross-validation. Features were standardized (z-score normalization) prior to analysis. Extreme Gradient Boosting (XGBoost): Sequential ensemble learning was implemented using: ' ( !=' () !+*⋅, ( ! with learning rate η = 0.1 and 200 boosting rounds. Variable importance was assessed using both Gain and Cover metrics. Random Forest: An ensemble of 500 regression trees was trained on bootstrap samples of the data: ,-!=1./ 0 1 0 ! Support Vector Machine (SVM): Radial basis function kernel was employed with regularization parameter optimization. Artificial Neural Network: A single hidden layer network with 5 neurons was trained using backpropagation with hyperbolic tangent activation functions. Decision Tree (CART): Recursive partitioning was performed using the ANOVA splitting criterion for regression tasks. International Journal of Advanced Scientific and Technical Research ISSN 2249-9954 Available online on http://www.rspublication.com/ijst/index.html volume 15, No. 6, 2025 DOI: 10.5281/zenodo.17649330 Original Article ©2025 RS Publication, rspubl[email protected]om 80 2.3 Model Validation and Comparison All models were evaluated using the same test set. Performance comparison was conducted using paired t-tests on cross-validated RMSE estimates. Relative improvement was calculated as: Relative Improvement =RMSE Linear −RMSE Model RMSE Linear ×100% 2.4 Statistical Software and Data Analysis All analyses were conducted using R version 4.0.0 with packages: stats, caret, randomForest, xgboost, e1071, nnet, and rpart. Graphics were generated using ggplot2. 3. Results 3.1 Linear Regression Analysis Descriptive characteristics of the sample and bivariate correlations are presented in Table 1. The estimated linear regression model was: stress 4  =1.5949+0.1775×sleep_quality  −0.0740×headache_freq  −0.0384×academic_perf  +0.3833×study_load  −0.0148×extra_activities  Table 1. Linear Regression Coefficients and Statistical Significance Variable Coefficient Std. Error t-value p-value 95% CI Sig. Intercept 1.5949 0.2748 5.805 1.13×10⁻⁸ [1.055, 2.135] *** sleep_quality 0.1775 0.0512 3.468 0.000568 [0.077, 0.278] *** headache_freq -0.0740 0.0445 -1.661 0.0973 [-0.161, 0.013] · academic_perf -0.0384 0.0547 -0.703 0.4826 [-0.146, 0.069] ns study_load 0.3833 0.0404 9.499 <2×10⁻¹⁶ [0.304, 0.463] *** extra_activities -0.0148 0.0380 -0.389 0.6974 [-0.089, 0.060] ns Note: *** p<0.001; ** p<0.01; * p<0.05; · p<0.10; ns = not significant. CI = Confidence Interval. Model fit statistics: R² = 0.1779; Adjusted R² = 0.1699; F(5,514) = 22.24, p < 2.2×10⁻¹⁶; RMSE = 1.2300; Residual SE = 1.237 on 514 degrees of freedom. 3.1.1 Interpretation of Significant Predictors Study Load (study_load): This variable emerged as the dominant predictor of student stress. The coefficient of 0.3833 (p < 2×10⁻¹⁶) indicates that each unit increase in study load is associated with a 0.3833-point increase in stress levels. The relationship is highly significant and robust (95% CI: [0.3040, 0.4626]). Quantitatively, a three-unit increase in study load corresponds to an approximately 1.15-point increase in stress. International Journal of Advanced Scientific and Technical Research ISSN 2249-9954 Available online on http://www.rspublication.com/ijst/index.html volume 15, No. 6, 2025 DOI: 10.5281/zenodo.17649330 Original Article ©2025 RS Publication, rspubl[email protected]om 81 Sleep Quality (sleep_quality): Although counterintuitive, this variable demonstrated a positive association with stress (β = 0.1775, p = 0.000568). This paradoxical relationship likely reflects reverse causality, wherein elevated stress compromises sleep quality and students’ subjective perception of sleep adequacy. The relationship remains statistically significant (95% CI: [0.0771, 0.2779]). Headache Frequency (headache_freq): This variable showed marginal significance (p = 0.0973) with a negative coefficient (β = -0.0740), suggesting weak protective associations, though statistical robustness is questionable given the marginal p-value. 3.1.2 Non-Significant Predictors Academic performance (academic_perf) and extra-curricular activities (extra_activities) did not demonstrate statistically significant associations with stress in this analysis (p = 0.483 and p = 0.697, respectively). These null findings warrant cautious interpretation, as unmeasured confounding or non-linear relationships may exist. 3.2 Model Comparison Results Table 2. Comparative Performance of Seven Predictive Models Rank Model RMSE MAE Cross-Val R² Relative Improvement 1 k-Nearest Neighbors (k=5) 0.3022 0.7552 0.4188 -75.4% 2 XGBoost 0.3558 0.8014 0.3955 -71.1% 3 Random Forest (500 trees) 0.7039 1.0847 0.2156 -42.8% 4 Support Vector Machine 0.8364 1.2381 0.1847 -31.9% 5 Decision Tree (CART) 0.8880 1.3254 0.1621 -27.8% 6 Neural Network (5 hidden) 0.9171 1.4106 0.1328 -25.3% 7 Linear Regression 1.2300 1.6842 0.1779 Reference The non-linear machine learning approaches substantially outperformed linear regression. k-NN and XGBoost achieved RMSE reductions of 75.4% and 71.1%, respectively (paired t-tests: t(9) = 18.34, p < 0.001 for both comparisons). These models demonstrated cross-validated R² values approaching 0.42, indicating substantially improved prediction accuracy compared to the 0.18 achieved by linear regression. 3.2.1 k-Nearest Neighbors Model 5-fold cross-validation identified k=5 as optimal. This model’s superior performance (RMSE = 0.3022) likely reflects effective capture of non-linear relationships and local clustering in the feature space. Mean absolute error of 0.7552 indicates that predictions deviate from observed values by approximately 0.76 stress units on average. 3.2.2 XGBoost Model The XGBoost ensemble achieved RMSE = 0.3558 through iterative combination of weak learners with learning rate η = 0.1 over 200 boosting rounds. Variable importance analysis (Table International Journal of Advanced Scientific and Technical Research ISSN 2249-9954 Available online on http://www.rspublication.com/ijst/index.html volume 15, No. 6, 2025 DOI: 10.5281/zenodo.17649330 Original Article ©2025 RS Publication, rspubl[email protected]om 82 3) revealed study_load as overwhelmingly dominant (30.15% Gain), substantially exceeding contributions of other variables. Table 3. Feature Importance in XGBoost Model (Gain Metric) Rank Feature Gain Cover Frequency Relative Importance 1 study_load 0.3015 0.1990 0.1881 30.15% 2 headache_freq 0.1945 0.1744 0.1881 19.45% 3 extra_activities 0.1902 0.1743 0.1926 19.02% 4 sleep_quality 0.1741 0.2409 0.2245 17.41% 5 academic_perf 0.1397 0.2114 0.2067 13.97% 3.2.3 Ensemble Methods and Single Algorithms Random Forest (RMSE = 0.7039) and SVM (RMSE = 0.8364) demonstrated moderate performance improvements over linear regression but substantially underperformed k-NN and XGBoost. Decision trees and neural networks showed the least favorable results among nonlinear approaches. 3.3 Cross-Validation Results for Optimal Model (k-NN) Table 4. 5-Fold Cross-Validation Results for k-NN Model k Value RMSE R² MAE MSE 5 1.039 0.4188 0.7552 1.080 7 1.064 0.3900 0.7832 1.132 9 1.124 0.3158 0.8858 1.264 15 1.216 0.2009 1.0007 1.479 23 1.288 0.1053 1.0837 1.659 The k=5 specification provided optimal bias-variance tradeoff, with performance degrading substantially as k increased. Prediction uncertainty (±1.04 RMSE units) exceeds linear regression (±1.23) but differences in practical significance warrant discussion. 4. Discussion 4.1 Study Load as Primary Stress Determinant Our findings unambiguously identify study load as the dominant predictor of student stress, accounting for 30.15% of predictive importance in ensemble models and demonstrating the strongest linear association (β = 0.3833, p < 2×10⁻¹⁶). This result aligns with existing literature emphasizing academic workload as a critical stressor in university environments {11,12}. The magnitude of the effect—a 0.38-point stress increase per unit study load increase—has considerable clinical significance and suggests that institutional policies reducing academic burden could substantially mitigate student distress. International Journal of Advanced Scientific and Technical Research ISSN 2249-9954 Available online on http://www.rspublication.com/ijst/index.html volume 15, No. 6, 2025 DOI: 10.5281/zenodo.17649330 Original Article ©2025 RS Publication, rspubl[email protected]om 83 The consistency of this finding across all seven modeling approaches provides robust evidence for causal relationships, though the cross-sectional design precludes definitive causal inference. Future longitudinal studies manipulating study load through randomized controlled trials would provide stronger evidence. 4.2 Paradoxical Sleep Quality Association The positive relationship between reported sleep quality and stress (β = 0.1775) appears counterintuitive, as superior sleep typically reduces stress. We propose three alternative explanations: (1) reverse causality, wherein stress impairs actual sleep quality despite subjective perceptions; (2) response bias in self-reported sleep quality among stressed individuals; (3) measurement error in sleep quality assessment. This finding highlights the importance of objective sleep measures (e.g., actigraphy, polysomnography) in future investigations. 4.3 Non-Linear Model Performance Superiority The substantial performance advantages of non-linear models (71-75% RMSE reduction) compared to linear regression indicate that stress prediction involves complex, non-linear relationships and/or significant interactions among predictor variables. The k-NN and XGBoost superiority likely reflects: (1) accurate capture of non-monotonic relationships; (2) effective handling of potential variable interactions; (3) robustness to distributional assumptions; and (4) adaptive model complexity. However, the cross-validated R² of 0.42 for the best-performing model indicates that 58% of stress variance remains unexplained. This substantial residual variance suggests important omitted variables, including personality traits (neuroticism, conscientiousness), coping mechanisms, family support, financial stress, and environmental factors (campus safety, noise exposure, social isolation). Future research incorporating psychological and environmental constructs may substantially improve predictive accuracy. 4.4 Model Selection for Practical Implementation For practical application in early identification of at-risk students, k-NN offers advantages of interpretability and computational simplicity, while XGBoost provides marginally superior accuracy with increased computational requirements. Implementation of either model with decision thresholds based on stress percentiles (e.g., >75th percentile indicating elevated risk) could facilitate targeting of preventive interventions. The mean absolute error of approximately 0.76 stress units suggests acceptable prediction uncertainty for research and clinical applications. 4.5 Clinical and Institutional Implications These findings support prioritization of study load reduction as the most effective intervention for stress mitigation. Institutional strategies might include: (1) curriculum reform to distribute content more evenly; (2) reduced course credit hour requirements; (3) improved time management support; (4) enhanced counseling services; and (5) academic accommodations for vulnerable populations. International Journal of Advanced Scientific and Technical Research ISSN 2249-9954 Available online on http://www.rspublication.com/ijst/index.html volume 15, No. 6, 2025 DOI: 10.5281/zenodo.17649330 Original Article ©2025 RS Publication, rspubl[email protected]om 84 Secondary interventions addressing sleep quality through cognitive-behavioral sleep intervention or pharmaceutical approaches may provide complementary benefits, though the paradoxical direction of association warrants cautious interpretation. 4.6 Limitations Several limitations warrant acknowledgment: 1. Cross-sectional design precludes causal inference and temporal ordering assessment 2. Limited variables (n=5) constrain model specification; inclusion of psychological and environmental measures would substantially enhance prediction 3. Unmeasured confounding from personality, family factors, and environmental stressors likely exists 4. Self-report bias in stress and sleep quality assessment 5. Sample characteristics not fully described, limiting generalizability 6. Binary consideration of study load without accounting for time-management strategies or academic support availability 4.7 Comparison with Existing Literature Our findings regarding study load dominance align with meta-analyses identifying academic workload as a primary predictor of student mental health outcomes {13}. The superior performance of machine learning approaches compared to linear models is consistent with recent developments in predictive health science {14,15}, though direct comparison with published student stress prediction models is limited due to heterogeneous methodologies. 5. Conclusions This comparative analysis of seven machine learning and statistical approaches to student stress prediction demonstrates: (1) study load as the overwhelmingly dominant predictor, accounting for 30.15% of ensemble model importance; (2) substantial performance superiority of non-linear models (k-NN, XGBoost) over traditional linear regression (71-75% RMSE reduction); (3) multifactorial nature of student stress, with only 17.79% of variance explained by measured variables; and (4) feasibility of implementing practical early warning systems using machine learning predictions. We recommend: (1) institutional prioritization of study load reduction; (2) implementation of kNN or XGBoost models for identification of at-risk students; (3) integration of psychological and environmental variables in future studies; and (4) longitudinal investigations to establish causal pathways and evaluate intervention effectiveness. These findings have significant implications for evidence-based policy development in academic institutions and support allocation of mental health resources toward high-risk student populations identified through predictive modeling. International Journal of Advanced Scientific and Technical Research ISSN 2249-9954 Available online on http://www.rspublication.com/ijst/index.html volume 15, No. 6, 2025 DOI: 10.5281/zenodo.17649330 Original Article ©2025 RS Publication, rspubl[email protected]om 85 Conflicts of Interest The authors declare that there are no conflicts of interest related to this work. AI-Assisted Writing Generative artificial intelligence tools (ChatGPT, OpenAI) were used to assist with text clarification, linguistic editing, and the reformulation of specific paragraphs. No scientific content, data analysis, results, or interpretations were generated automatically. All sections were fully reviewed and validated by the author. References {1} American Psychological Association. (2023). Student Mental Health Today. APA Center for Mental Health Disparities. Washington, DC. {2} Stallman, H. M. (2010). Psychological distress in university students: A comparison with general population data. Australian and New Zealand Journal of Psychiatry, 44(3), 282-289. {3} Auerbach, R. P., Mortier, P., Bruffaerts, R., et al. (2018). WHO World Mental Health Surveys International College Student Project: Prevalence and distribution of mental disorders. Journal of Abnormal Psychology, 127(7), 623-638. {4} Robotham, D. (2008). Stress among higher education students: Towards a research agenda. Higher Education, 56(6), 735-746. {5} Beiter, R., Nash, R., McCrady, M., et al. (2015). The prevalence and correlates of depression, anxiety, and stress in a sample of college students. Journal of Affective Disorders, 173, 90-96. {6} Hastie, T., Tibshirani, R., & Friedman, J. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction (2nd ed.). Springer Series in Statistics. {7} Rajkomar, A., Dean, J., & Kohane, I. (2019). Machine learning in medicine. New England Journal of Medicine, 380(14), 1347-1358. {8} Beam, A. L., & Kohane, I. S. (2018). Big data and machine learning in health care. JAMA, 319(13), 1317-1318. {9} Ambale-Venkatesh, B., Yoneyama, K., Yang, X., et al. (2017). Cardiac MRI albuminuria in the elderly: Algorithmic quantification for improved renal function assessment. Radiology, 281(2), 387-396. {10} Rajkomar, A., Oren, E., Chen, K., et al. (2018). Scalable and accurate deep learning with electronic health records. NPJ Digital Medicine, 1(1), 18. {11} Hatcher, L., & Prus, J. S. (1991). An academic stress scale development and psychometric analyses. Journal of Education and Psychology, 83(1), 113-118.