Feature Selection and Ensemble Methods for Predicting Parkinson's Disease Outcomes: A Data-Driven Approach
Full text
Gemiten Gnmze Lingual Ortodontik Apareyler 1
1 Feature Selection and Ensemble Methods for Predicting Parkinson’s Disease Outcomes: A Data-Driven Approach SELİM BUYRUKOĞLU 1 GONCA BUYRUKOĞLU 2 1 Assoc. Prof. Dr; Çankırı Karatekin University, Faculty of Engineering, Department of Computer Engineering, [email protected], ORCID No: 0000-0001-7844-3168 2 Assist. Prof. Dr; Abdullah Gül University, Faculty of Life and Natural Sciences, Molecular Biology and Genetics Department, gonca.buyruko[email protected]du.tr, ORCID No: 0000-0002-1202-8778
2 Feature Selection and Ensemble Methods for Predicting Parkinson’s Disease Outcomes: A Data-Driven Approach SELİM BUYRUKOĞLU, GONCA BUYRUKOĞLU Design: All Sciences Academy Design Publication Date: December 2025 Publisher’s Certificate Number: 72273 ISBN: 978-625-8536-19-5 Doi: https://doi.org/10.5281/zenodo.17887611 © All Sciences Academy www.allsciencesacademy.com [email protected]
3 CONTENT ÖNSÖZ 5 PREFACE 6 ÖZET 7 ABSTRACT 8 INTRODUCTION 9 Background and Motivation 9 Machine Learning for Parkinson's Disease Progression Modeling 9 Research Questions and Objectives 11 LITERATURE REVIEW 12 Machine Learning for Parkinson’s Disease Prognosis 12 PPMI Dataset of machine learning research. 12 Feature Selection in Clinical ML Models 13 Ensemble Learning and Meta-Learning In Biomedical Prediction 14 MATERIALS AND METHODOLOGY 15 Datasets 16 Dataset and outcome definition 17 Data Preprocessing 18 Feature selection 18 Machine Learning Models 19 Random Forest (RF) 19 Gradient Boosting Machine (GBM) 20 ET (Extra Trees / Extremely Randomized Trees) 20
iv AB (AdaBoost) 21 Support Vector Machine (SVM) 22 LR (Logistic Regression) 22 Ensemble Learning Models 23 Stacking (Stacked Ensemble) 23 Super Learner Ensemble 24 Performance Metrics 25 RESULTS AND DISCUSSION 25 Model Performances 25 Error Analysis and Confusion Matrices 28 CONCLUSIONS AND FURTHER DIRECTIONS 31 REFERENCES 32
5 ÖNSÖZ Sevgili Okuyucular, Parkinson hastalığının klinik seyrinin öngörülmesi, hem bireyselleştirilmiş tedavi stratejilerinin geliştirilmesi hem de daha etkili klinik deneylerin tasarlanabilmesi açısından günümüz tıbbında giderek daha önemli hâle gelmektedir. Bu kitap, Parkinson hastalığı ilerleyişini makine öğrenimi temelli yöntemlerle modellemeye yönelik kapsamlı bir yaklaşım sunmak amacıyla hazırlanmıştır. Çalışmamız, özellikle PPMI veri seti üzerinden gerçekleştirilen çok yönlü analizlerle, klinik verilerin yapısal özelliklerine uygun modern makine öğrenimi boru hatlarının nasıl tasarlanabileceğine dair bütüncül bir çerçeve ortaya koymaktadır. Bu kitabı oluştururken temel amacımız, araştırmacılara, klinisyenlere ve biyomedikal veri bilimiyle ilgilenen tüm okuyuculara hem metodolojik açıdan rehberlik edecek hem de uygulamaya dönük örneklerle zenginleştirilmiş bir kaynak sunmaktı. Özellikle özellik seçimi, topluluk modelleri ve metaöğrenme gibi modern yöntemlerin Parkinson hastalığı bağlamında nasıl etkili biçimde kullanılabileceğini göstermek, çalışmanın temel motivasyonlarından biri olmuştur. Ele aldığımız yöntemler, yalnızca akademik araştırmalara değil, aynı zamanda klinik karar destek sistemlerinin geliştirilmesine yönelik de önemli ipuçları barındırmaktadır. Bu eserin ortaya çıkmasına katkı sağlayan tüm bilim insanlarına, klinik veri sağlayıcılarına ve Parkinson hastalığının anlaşılması yolunda çaba veren herkese teşekkür ederiz. Kitabın, bu alanda çalışan araştırmacılara ilham vermesi, yeni metodolojik yaklaşımların önünü açması ve Parkinson hastalığıyla ilgili gelecekteki çalışmalara katkıda bulunması en büyük dileğimizdir. Tüm meslektaşlarımıza iyi çalışmalar dileriz. Saygılarımızla, SELİM BUYRUKOĞLU, GONCA BUYRUKOĞLU
6 PREFACE Dear Readers, The prognosis of the development of the disease in Parkinson has gained greater significance in the field of contemporary medicine, not only to create individual approaches to treatment but also to create more effective clinical trials. The presented book has been written to introduce a methodology of modeling the development of the Parkinson disease disease, based on machine learning approaches. In this way, our work, based on a thorough analysis and especially on the PPMI dataset, provides a comprehensive structure in which modern machine learning pipelines are designed in accordance with the structural properties of clinical data. The main aim of writing this book was to give researchers, clinicians, and all the readers interested in biomedical data science a source where they could find methodological advice with practical examples. Illustrating how the modern techniques, including feature selection, ensemble models, and metalearning, can be helpful to be applied under the conditions of the Parkinson disease has been one of the key inspirations of the given work. The methods mentioned in the present paper, in addition to being valuable in terms of academic research, provide useful information in the context of clinical decision-support system development. It is our pleasure to acknowledge all scientists, clinical data suppliers, and other people who contribute to the development of the knowledge about Parkinson disease. We hope that this book will motivate scientists to conduct research in the same area, develop new methodological ideas, and also make some contribution to the future research on the progress of the Parkinson disease. Wishing all our colleague success in their work. Sincerely, SELİM BUYRUKOĞLU, GONCA BUYRUKOĞLU
7 ÖZET Parkinson hastalığı (PH) sürecinin erken tahmini ve ilerleyişinin stratifikasyonu, bireysel tedavi planlaması ve klinik çalışma tasarımları açısından kritik öneme sahiptir. Bu çalışmada, Parkinson Progression Markers Initiative (PPMI) klinik verileri üzerinde, tam takip verisine sahip 1.008 katılımcının farklı hastalık ilerleme örüntülerini üç sınıflı bir olay çıktısı olarak tahmin etmeye yönelik kapsamlı bir makine öğrenimi hattı tasarlanmış ve kapsamlı biçimde değerlendirilmiştir. Analitik süreç; sistematik veri ön işleme adımlarını (>%50 eksik veriye sahip değişkenlerin dışlanması, özellik etiketleme, medyan değerle eksik veri tamamlama, z-skor normalizasyonu), ANOVA F-istatistiği ile tek değişkenli özellik seçimini (SelectKBest, k=20) ve ağaç tabanlı topluluk modelleri (Random Forest, Gradient Boosting, Extra Trees, AdaBoost), destek vektör makineleri, lojistik regresyon, lojistik regresyon meta-öğrenicili yığma (stacking) modeli ve dış örneklem tabanlı meta-öğrenmeyle oluşturulan bir süper öğrenici (super learner) dahil olmak üzere çeşitli sınıflandırma çerçeveleriyle titiz karşılaştırmaları içermektedir. Modeller; doğruluk, makro ortalama kesinlik, duyarlılık, F1-skoru, ROC-AUC ve beş katlı tabakalı çapraz doğrulama ile elde edilen karmaşıklık matrisleri açısından değerlendirilmiştir. Süper öğrenici meta-öğrenme modeli, tahmin başarımı açısından en yüksek performansa ulaşmış (doğruluk: %94,94 ± %2,58; F1-skoru: %93,06) ve tek başına kullanılan Gradient Boosting temel modeline (%93,35 ± %1,73) belirgin ölçüde üstünlük sağlayarak, meta-öğrenici seçiminin orta büyüklükteki klinik veri kümelerinde topluluk modellerinin etkinliği için kritik bir bileşen olduğunu göstermiştir. Özellik seçimi, gradient boosting algoritmalarının performansını artırmış ancak diğer model ailelerinde tutarsız sonuçlar vermiştir. Bu çalışma, uygun şekilde yapılandırılmış meta-öğrenme mimarilerinin yapılandırılmış klinik tablosal verilerde güncel en yüksek tahmin performansına ulaşabildiğini ve özellik önemi ile SHAP analizi üzerinden yorumlanabilir çıktılar sağlayarak, PH ilerleyişinin modellenmesine yönelik gelecekteki araştırmalar için metodolojik bir çerçeve sunduğunu ortaya koymaktadır. Anahtar Kelimeler – Makine Öğrenimi, Parkinson Hastalığı, Sper Öğrenici, PPMI Veri Seti, Hastalık İlerleyii Tahmini, Gradyan Artırma, Yığma Sınıflandırıcı
8 ABSTRACT The main problem is that early prediction and stratification of the progression of the Parkinson disease (PD) process are important to individual treatment planning and the development of clinical trials. We designed and critically assessed in this work a full machine learning pipeline on the Parkinson progression Markers Initiative (PPMI) clinical data to predict a three-class event outcome of the various disease progression patterns of a cohort of 1,008 subjects who have full follow-up data. The analytical procedure involved systematic preprocessing (removal of missing data >50%, labelling features, median measure imputation, z-score normalization), univariate feature choice by ANOVA F-statistics (SelectKBest, k=20), and stringent comparison with several classification frameworks such as treebased ensemble models (Random Forest, Gradient Boosting, Extra Trees, AdaBoost), support machine, logistic regression, stacking ensemble with logistic regression meta-learner, and a super learner based on meta-learning using out-of The models were evaluated in terms of accuracy, macro-averaged precision, recall, F1-score, ROC-AUC and confusion matrices taking a 5 fold stratified cross-validation. The super learner meta-learning model performed the best with regard to predictive performance (accuracy: 94.94% ± 2.58%, F1-score: 93.06%) and exceeded the single Gradient Boosting baseline (93.35% ± 1.73%) by a significant margin, showing that meta-learner selection is a critical factor in the effectiveness of ensembles in medium-sized clinical data. The use of feature selection enhanced the performance of gradient-boosting algorithms but showed inconsistent performance across other family of models. The proposed workflow shows that properly configured meta-learning architectures can get state-of-the-art predictive performance on structured clinical tabular data and give interpretative results by feature importance and SHAP analysis, which can serve as a methodological template of future studies on PD progression modelling. Keywords – Machine Learning, Parkinson's disease, super learner, PPMI dataset, disease progression prediction, gradient boosting, stacking classifier
15 done in an organized and fair manner are still a significant methodological issue in clinical ML. MATERIALS AND METHODOLOGY This chapter describes the overall methodology used in the current research to design and test machine learning systems that predict event outcomes in three classes in the progression of Parkinsonism. The methodology involves systematic processing of the data, univariate feature selection, creation of various classification architectures (base models, stacking ensembles, and meta-learning methods) and intensive evaluation of its performance through stratified cross-validation and various performance measures. The entire pre-processing, modeling, and assessment procedures were developed in such a way that they could guarantee methodological transparency, reproducibility, and reasonable comparison of the competing methods and still provide clinical meaning to the findings. Figure 1 shows the proposed model’s structure. Figure 1: Proposed Model Structure
16 Datasets The present study used Parkinson's Progress Markers Initiative (PPMI) clinical dataset, a large-scale longitudinal observational cohort study aimed at identifying and authenticated biomarkers of the progression of the Parkinson disease. PPMI dataset is a rich clinical, cognitive, motor, non-motor and biomarker evaluation of patients with known diagnosis of PD, prodromal participants and healthy controls in various international sites. Table 1 summarizes the selected predictors and its clinical descriptions. Table 1: Selected Features (SelectKBest, k=20) Feature Name Description ApoE_Genotype Apolipoprotein E genotype Age Patient age at baseline Sex Gender (encoded) NHY Hoehn and Yahr stage LEDD_sum Levodopa equivalent daily dose BL_Lymphocytes Baseline lymphocyte count Lymphocytes Current lymphocyte count HVLT_DELAYED_RECALL_T Hopkins Verbal Learning Test – delayed recall JLO_RAW Judgment of Line Orientation – raw score SDMT_correct Symbol Digit Modalities Test – correct responses RBDSQ_total_score REM Sleep Behavior Disorder Screening Questionnaire ApoE_e4_allele Presence of ApoE ε4 allele ess_total Epworth Sleepiness Scale total EDS Excessive daytime sleepiness indicator ger_tot Geriatric Depression Scale total
17 stai_state State-Trait Anxiety Inventory – state stai_trait State-Trait Anxiety Inventory – trait neuro_total Neuropsychiatric symptoms total eventtime Time from baseline to event (months) numericEDS Numeric encoding of EDS Dataset and outcome definition The final data set used consisted of 1,008 subjects having full followup data and an event outcome that indicated a 3-class event outcome (event ∈ {0, 1, 2}), which is the various stages or progression patterns of the disease. The initial data set of predictors covered around 50-60 candidate predictor variables comprising of demographic data (age, sex), genetic variables (ApoE genotype, ApoE ε4 allele status), motor data (Hoehn and Yahr stage, UPDRS scores), cognitive and neuropsychological data (Hopkins Verbal Learning Test, Symbol Digit Modalities Test, Judgment of Line Orientation), nonmotor symptom data (REM Sleep Behavior Disorder Questionnaire, Epworth Sleepiness Scale, State-Trait An In order to minimize noise and imputation bias, all columns with a greater than 50 percent missing value were removed in a systematic manner and not further analysed. Also, the identifier variables with non-informative data, including patient ID (PATNO), event ID and original entry areas, were excluded to concentrate on clinically meaningful predictors. After preprocessing and univariate feature selection with ANOVA Fstatistics (SelectKBest with k=20) a narrowed down set of 20 predictors was then saved to be used in model training and evaluation. Each model was evaluated with 5-fold stratified cross-validation in order to have balanced representation of classes in each fold and to give good estimates of the performance of the model in generalization. Besides, 20 percent hold-out test set was to be used to perform independent assessment of ROC curves and confusion matrices. This preprocessing and evaluation system provided fair comparison between all modeling methods at the same time being clinically interpretable and having methodological rigor.
18 Data Preprocessing The PPMI clinical data underwent a systematic preprocessing pipeline in order to guarantee the quality of the data and compatibility with the model. To minimize the impact of imputation bias and noise in subsequent analyses, first, columns with a proportion of more than 50% missing values were removed. Fields that do not provide information like patient ID (PATNO), event ID and original entry fields were also eliminated but only clinically significant predictors were kept in their place. Categorical variables, such as sex, ApoE genotype, and excessive daytime sleepiness (EDS) were coded off with label encoding to code the nominal categories as numbers to be used with machine learning algorithms. The rest of the featuresthere were no missing values so the median of each feature was used to impute the missing values which is a very robust method to reduce the effects of outliers whilst preserving distributional characteristics of data. In order to achieve algorithmic stability and smooth convergence, especially of distance-based and gradient-descent models, like support vector machines (SVM) and logistic regression, z-score normalization (StandardScaler) was used to standardize all features. This transformation maps all the features to zero mean and unit variance, which puts all the predictors in roughly equal range and avoids the domination of model training to features with larger numerical ranges. Feature selection The univariate filter feature selection is done with ANOVA F-statistics, which are implemented with the help of the SelectKBest option with k= 20. This technique ranks features based on their univariate statistical association with the target variable and the 20 highest ranked features are used to form a prediction. This dimensionality reduction scheme can achieve several goals: (1) it helps reduce the risk of overfitting due to the dimension being small compared to the sample size (n=1,008), (2) it supports a better understanding of the model because it prevents the need to compare highly dimensionality models with a wide variety of input features (resulting in a fair comparison), and (3) it allows more direct interpretation of the model (because a small number of clinically relevant variables can be considered).
19 The chosen subset of 20 features (named as 𝑿selected ) includes the major non-motor assessment scales and clinical scales, that are known to be significant indicators of the progression of the disease of Parkinson. These consist of sleep-related (Epworth Sleepiness Scale, REM Sleep Behavior Disorder Questionnaire), neuropsychiatric (State-Trait Anxiety Inventory, Geriatric Depression Scale), and cognitive function (Hopkins Verbal Learning Test, Symbol Digit Modalities Test, Judgment of Line Orientation), motor severity (Hoehn and Yahr stage), genetic (ApoE genotype and e4 allele status) and temporal progression variables (event time, levodopa equivalent daily dose). The set of features was kept in all subsequent modeling so that it can be consistent and can be easily interpreted with the aid of post-hoc metrics such as feature importance ranking and SHAP (SHapley Additive exPlanations) analyses. Machine Learning Models The base line ML models that have been used in this research have the following individual benefits in terms of classification functions: Random Forest (RF): This is an ensemble algorithm which trains many decision trees on random data and features, and then combines the predictions by majority voting (classification) or averaging (regression); it is also resistant to noise and overfitting, feature importance scores, and scales to a variety of data (Breiman, 2001). Figure 2 shows the structure of RF. Figure 2: RF Model Structure
20 Gradient Boosting Machine (GBM): This is boosting ensemble method, which trains decision trees in a series, with each new tree specializing in correcting the errors of the earlier trees, using gradient descent to optimize; this framework is commonly found to be highly accurate, but it takes a lot of hyperparameter optimization to ensure underfitting (Friedman, 2002). Figure 3 presents the structure of GBM. Figure 3: GBM Model Structure ET (Extra Trees / Extremely Randomized Trees): Similar to Random Forest except that it is more random in its choice of the split thresholds (not only features), resulting in less variance and less computation time although with high-dimensional data with noise is good (Geurts et al., 2006). Figure 4 depicts the structure of ET.
21 Figure 4: ET Model Structure AB (AdaBoost): This type of adaptive boosting algorithm trains sequential weak learners (typically shallow decision trees), and weights these learners on misclassified cases, boosting weights in this case, to put more emphasis on errors; good with binary classification, but susceptible to outliers and noisy data (Freund & Schapire, 1997). Figure 5 demonstrates the structure of AB. Figure 5: AB Model Structure
22 Support Vector Machine (SVM): This is a supervised algorithm to discover the best hyperplane (or data transformation with a kernel) to maximize the distance between classes with focus on support vectors (data points that are important to the model); it is very effective in high-dimensional spaces, as well as in cases where the margin is clear, but it is computationally expensive with large data sets (Cortes & Vapnik, 1995). Figure 6 shows the structure of SVM. Figure 6: SVM Model Structure LR (Logistic Regression): A linear model which predicts class probabilities using a logistic (sigmoid) function on a linear combination of features; simple, interpretable, and fast, and is a powerful baseline in binary/multiclass classification, in which the relationships are more or less linear (Cramer, 2005). Figure 7 indicates the structure of LR. Figure 7: LR Model Structure
23 Ensemble Learning Models Stacking (Stacked Ensemble) Stacking (also known as stacked generalization) is an ensemble method, which leverages the predictive power of multiple distinct models (base learners) by learning a second-level (meta-learner) model to combine the predictions of the first level models. Stacking Each base learner is initially trained on the training data - possibly with highly disparate learning algorithms and hyperparameters. Their prediction of each instance are then the input features to the meta-learner which learns to effectively integrate the outputs to create a final prediction. In contrast to bagging or boosting which employ homogenous base learners and minimize variance or bias respectively, stacking is particularly effective with heterogeneous models, where the complementary strengths of the models are used. Alternatively, the original input features may be supplied to the meta-learner (so-called passthrough) to provide as much information as possible to the ultimate decision making (Wolpert, 1992) . The base learners that were used in this research included a Gradient Boosting (GB), Random Forest (RF), Extra Trees (ET) and Support Vector Machine (SVM) and the meta-learner was the Logistic Regression. In the stacking meta-model, the predictions of each base learner were out-of-fold as well as the original input features were used (passthrough=True), and the meta-model was able to learn combinations of flexibilities by knowledge of various modeling paradigms. Figure 8 demonstrates the structure of stacked ensemble. Figure 8: Stacked Model Structure
24 Super Learner Ensemble The Super Learner is a widely formalized generalization of stacking which represents the problem of ensemble modeling as an optimization problem with respect to a library of candidate prediction algorithms. Within the Super Learner, the cross-validation is conducted on a variety of base models, and their out-of-fold predictions are gathered into a meta-feature matrix of features of level-one. Such meta-features are then learned using the meta-learner - usually as a minimizer of a loss-function such as log-loss or mean squared error - to learn an optimum, data adaptive combination of the base learners. The Super Learner has a unique property, which is its asymptotic property: the larger this sample, the closer the Super Learner gets to or even equals the performance of the oracle, the best convex combination of its candidate models, and this makes it resistant to weak or poor models (which are given near-zero weight in the ensemble). This method is potent since it is able to combine the models of numerous kinds and enables the information, as opposed to the researcher intuitions, to select the most appropriate combination (Van Der Laan et al., 2007). Figure 9 presents the structure of super ensemble approach. Figure 9: Super Ensemble Model Structure
31 Figure 13. Heatmap of model performance (accuracy, precision, recall, and F1score) for all base and ensemble models. CONCLUSIONS AND FURTHER DIRECTIONS This book illustrates that a well-crafted, yet conceptually straightforward, machine learning pipeline is capable of producing a high predictive accuracy on three-class PD event outcome on the PPMI data. The overall Gradient Boosting of a subset of features yields the best single base model, which has high and consistent accuracy and Macro F1-scores in crossvalidation folds. It is based on this that a meta-learning Super Learner, which integrates several base learners through out-of-fold predictions and a Gradient Boosting meta-learner, performs the best in overall performance, evidently beating the Gradient Boosting base learner and the stacked ensemble. Seen through the methodological prism, those findings represent a number of practical implications of clinical ML applications to medium-sized tabular datasets. To start with, the use of rigorous feature selection and standardized evaluation (5-fold stratified cross-validation, macro-averaged metrics, ROC and confusion-matrix analysis) should at least be rated equally with the model sophistication. Second, ensemble and meta-learning models
32 ought to be fairly compared to the performance of powerful single-model baselines; in our experiment, a naively trained stacked ensemble is not doing well compared to Gradient Boosting, but the better-constructed Super Learner is offering a significant performance improvement. Third, error patterns indicate that the performance of models is mainly limited by the inability to predict the minority class, which highlights the necessity to employ such strategies as explicit class balancing, addition of longitudinal trajectories, or the addition of more comprehensive imaging and biomarker information. Future efforts might build on this pipeline to include temporal modeling of progression (e.g., recurrent or survival model), external validation of independent PD cohorts, and more advanced explainability methods (e.g. SHAP summary and dependence plots) to better understand what clinical and biomarker features are influencing the model decisions. However, the current work offers an example of the design, evaluation, and reporting of machine learning and ensemble models in PD progression prediction based on actual clinical data in a reproducible and end-to-end manner. REFERENCES Alwazy, A. S. H., Buyrukoğlu, G., Buyrukoğlu, S. & Baker, M. R. (2025). Evaluating machine learning and statistical learning techniques for cancer classification and diagnosis. Iran Journal of Computer Science, 1–20. https://doi.org/10.1007/S42044-025-00233-Z Asadi, S., Amiri, S. S. & Mottahedi, M. (2014). On the development of multi-linear regression analysis to assess energy consumption in the early stages of building design. Energy and Buildings, 85, 246–255. Baker, M. R. & Solanki, U. (2024). Artificial Intelligence Models in Pattern Recognition. In Handbook of Artificial Intelligence Applications for Industrial Sustainability: Concepts and Practical Examples. https://doi.org/10.1201/9781003348351-2 Banik, R., Das, P., Ray, S. & Biswas, A. (2021). Prediction of electrical energy consumption based on machine learning technique. Electrical Engineering, 103(2), 909–920. Buyruko\uglu, G. & Philipson, P. (2025). Multilevel joint modelling of hierarchical longitudinal and time-to-event data. Statistics, 1–20. Buyruko\uglu, S. & Y\ilmaz, Y. (2021). An approach for airfare prices analysis with penalized regression methods. Veri Bilimi, 4(2), 57–61. Buyrukoğlu, G. (2024). Survival analysis in breast cancer: evaluating ensemble learning techniques for prediction. PeerJ Computer Science, 10, e2147. https://doi.org/10.7717/peerj-cs.2147 Buyrukoğlu, S., Baker, M. R., Jihad, K. H., Etem, T. & Buyrukoğlu, G. (2025). NBA 2K20 Player Rating Predictions Using Machine Learning and Ensemble
33 Learning Approaches. Lecture Notes in Networks and Systems, 144–153. https://doi.org/10.1007/978-3-031-78940-3_14 Divina, F., Gilson, A., Goméz-Vela, F., Torres, M. G. & Torres, J. F. (2018). Stacking ensemble learning for short-term electricity consumption forecasting. Energies, 11(4), 949. https://doi.org/10.3390/en11040949 Dong, B., Cao, C. & Lee, S. E. (2005). Applying support vector machines to predict building energy consumption in tropical region. Energy and Buildings, 37(5), 545–553. https://doi.org/10.1016/j.enbuild.2004.09.009 Dostmohammadi, M., Pedram, M. Z., Hoseinzadeh, S. & Garcia, D. A. (2024). A GAstacking ensemble approach for forecasting energy consumption in a smart household: A comparative study of ensemble methods. Journal of Environmental Management, 364, 121264. Guo, J., Yun, S., Meng, Y., He, N., Ye, D., Zhao, Z., Jia, L. & Yang, L. (2023). Prediction of heating and cooling loads based on light gradient boosting machine algorithms. Building and Environment, 236, 110252. Hosamo, H. & Mazzetto, S. (2024). Performance Evaluation of Machine Learning Models for Predicting Energy Consumption and Occupant Dissatisfaction in Buildings. Buildings, 15(1), 39. Kavaklioglu, K. (2011). Modeling and prediction of Turkey’s electricity consumption using Support Vector Regression. Applied Energy, 88(1), 368–375. https://doi.org/10.1016/j.apenergy.2010.07.021 Kiprijanovska, I., Stankoski, S., Ilievski, I., Jovanovski, S., Gams, M. & Gjoreski, H. (2020). Houseec: Day-ahead household electrical energy consumption forecasting using deep learning. Energies, 13(10), 2672. Klass, D. L. (2004). Biomass for renewable energy and fuels. Encyclopedia of Energy, 1(1), 193–212. Kumar, R. (2023). Python Machine Learning: A Beginner’s Guide to Scikit-Learn. Jamba Academy. Liao, J.-M., Chang, M.-J. & Chang, L.-M. (2020). Prediction of air-conditioning energy consumption in R\&D building using multiple machine learning techniques. Energies, 13(7), 1847. Liu, Y., Chen, H., Zhang, L., Wu, X. & Wang, X. jia. (2020). Energy consumption prediction and diagnosis of public buildings based on support vector machine learning: A case study in China. Journal of Cleaner Production, 272. https://doi.org/10.1016/j.jclepro.2020.122542 Lü, X., Lu, T., Kibert, C. J. & Viljanen, M. (2015). Modeling and forecasting energy consumption for heterogeneous buildings using a physical-statistical approach. Applied Energy, 144, 261–275. https://doi.org/10.1016/j.apenergy.2014.12.019 Mottahedi, M., Mohammadpour, A., Amiri, S. S., Riley, D. & Asadi, S. (2015). Multilinear regression models to predict the annual energy consumption of an office building with different shapes. Procedia Engineering, 118, 622–629. Ogunsanya, M., Isichei, J. & Desai, S. (2023). Grid search hyperparameter tuning in additive manufacturing processes. Manufacturing Letters, 35, 1031–1042. Oludolapo, O. A., Jimoh, A. A. & Kholopane, P. A. (2012). Comparing performance of MLP and RBF neural network models for predicting South Africa’s energy consumption. Journal of Energy in Southern Africa, 23(3), 40–46. Pham, A.-D., Ngo, N.-T., Truong, T. T. H., Huynh, N.-T. & Truong, N.-S. (2020). Predicting energy consumption in multiple buildings using machine learning for improving energy efficiency and sustainability. Journal of Cleaner
34 Production, 260, 121082. Sharma, V. (2022). Exploring the Predictive Power of Machine Learning for Energy Consumption in Buildings. Journal of Technological Innovations, 3(1). Touzani, S., Granderson, J. & Fernandes, S. (2018). Gradient boosting machine for modeling the energy consumption of commercial buildings. Energy and Buildings, 158(510), 1533–1543. https://doi.org/10.1016/j.enbuild.2017.11.039 Ves, A. V., Ghitescu, N., Pop, C., Antal, M., Cioara, T., Anghel, I. & Salomie, I. (2019). A stacking multi-learning ensemble model for predicting near real time energy consumption demand of residential buildings. 2019 IEEE 15th International Conference on Intelligent Computer Communication and Processing (ICCP), 183–189. Wang, G., Mukhtar, A., Moayedi, H., Khalilpoor, N. & Tt, Q. (2024). Application and evaluation of the evolutionary algorithms combined with conventional neural network to determine the building energy consumption of the residential sector. Energy, 298, 131312. Wang, Z. & Srinivasan, R. S. (2017). A review of artificial intelligence based building energy use prediction: Contrasting the capabilities of single and ensemble prediction models. Renewable and Sustainable Energy Reviews, 75, 796–808. https://doi.org/10.1016/j.rser.2016.10.079 Wei, Y., Zhang, X., Shi, Y., Xia, L., Pan, S., Wu, J., Han, M. & Zhao, X. (2018). A review of data-driven approaches for prediction and classification of building energy consumption. Renewable and Sustainable Energy Reviews, 82, 1027– 1047. World Energy Balances - Data product. (2025). IEA. https://www.iea.org/data-andstatistics/data-product/world-energy-balances Yan, Z. & Wen, H. (2021). Electricity theft detection base on extreme gradient boosting in AMI. IEEE Transactions on Instrumentation and Measurement, 70, 1–9. Zhang, F., Deb, C., Lee, S. E., Yang, J. & Shah, K. W. (2016). Time series forecasting for building energy consumption using weighted Support Vector Regression with differential evolution optimization technique. Energy and Buildings, 126, 94–103. Zhong, H., Wang, J., Jia, H., Mu, Y. & Lv, S. (2019). Vector field-based support vector regression for building energy consumption prediction. Applied Energy, 242, 403–414. https://doi.org/10.1016/j.apenergy.2019.03.078