Comparative Analysis of Meta-Heuristic Optimization–Enhanced Machine Learning Models for Heart Failure Prediction
Full text
Gemiten Gnmze Lingual Ortodontik Apareyler 1
1 Comparative Analysis of Meta-Heuristic Optimization–Enhanced Machine Learning Models for Heart Failure Prediction SELİM BUYRUKOĞLU 1 GONCA BUYRUKOĞLU 2 1 Assoc. Prof. Dr; Çankırı Karatekin University, Faculty of Engineering, Department of Computer Engineering, [email protected], ORCID No: 0000-0001-7844-3168 2 Assist. Prof. Dr; Abdullah Gül University, Faculty of Life and Natural Sciences, Molecular Biology and Genetics Department, gonca.buyruko[email protected]du.tr, ORCID No: 0000-0002-1202-8778
2 Comparative Analysis of Meta-Heuristic Optimization–Enhanced Machine Learning Models for Heart Failure Prediction SELİM BUYRUKOĞLU, GONCA BUYRUKOĞLU Design: All Sciences Academy Design Publication Date: December 2025 Publisher’s Certificate Number: 72273 ISBN: 978-625-8536-18-8 Doi: https://doi.org/10.5281/zenodo.17887588 © All Sciences Academy www.allsciencesacademy.com [email protected]
3 CONTENT ÖNSÖZ 4 PREFACE 5 ÖZET 6 ABSTRACT 7 INTRODUCTION 8 Novelty and Contributions 10 LITERATURE REVIEW 12 Heart Disease Prediction with Machine Learning 12 Use of Meta-Heuristic Optimization Algorithms 13 Generalizability and Comparisons of Datasets 14 MATERIALS AND METHODOLOGY 14 Datasets 15 Data Preprocessing 18 Machine Learning Models 19 Metaheuristic Optimization Algorithms 20 Performance Metrics 21 RESULTS AND DISCUSSION 21 Results on Dataset-1 21 Results on Dataset-2 25 Visual representation 28 Discussion 30 CONCLUSION AND FURTHER DIRECTIONS 31 REFERENCES 33
4 ÖNSÖZ Sevgili Okuyucular, Bilimsel araştırmaların hızla geliştiği günümüzde, makine öğrenmesi ve optimizasyon tekniklerinin sağlık alanında sunduğu olanaklar her zamankinden daha değerlidir. Bu çalışma, kalp yetmezliği gibi küresel ölçekte kritik bir sağlık sorununun erken tanısında kullanılabilecek güçlü, güvenilir ve karşılaştırmalı modeller geliştirmeyi amaçlamaktadır. Meta-sezgisel algoritmalar ve makine öğrenmesi yöntemlerinin bir araya getirilmesiyle oluşturulan hibrit yaklaşımlar hem klinik karar süreçlerine hem de gelecekteki akademik çalışmalara katkı sağlayacak nitelikte bulgular üretmiştir. Bu kitabın ortaya çıkmasında emeği geçen tüm araştırmacılara ve katkılarıyla çalışmayı zenginleştiren bilim insanlarına teşekkür ederiz. Ayrıca, veri bilimi ve sağlık teknolojileri alanında çalışan tüm meslektaşlarımızın, bu eserin sunduğu yöntem ve sonuçlardan yararlanarak daha ileri çalışmalar üretmesini temenni ederiz. Bilimsel ilerlemenin, paylaşılan bilgi ve kolektif çabayla daha da güçleneceğine inanıyoruz. Bu kapsamda hazırlanan bu kitabın, hem akademi hem de sağlık sektöründe çalışan araştırmacılar için faydalı bir kaynak olmasını diliyor; tüm okuyuculara çalışmalarında başarılar ve kolaylıklar diliyoruz. Tüm meslektaşlarımıza iyi çalışmalar dileriz. Saygılarımızla, SELİM BUYRUKOĞLU, GONCA BUYRUKOĞLU
5 PREFACE Dear Readers, The recent years are characterized by the rapid development of machine learning and the methods of optimization that have provided the opportunities to tackle the complex issues in the sphere of healthcare. This book tries to enter into these efforts by providing an all-inclusive comparative study of meta-heuristic optimization algorithms that are combined with machine learning models to predict heart failure. By comparing various hybrid strategies on a variety of datasets, the research aims both to enhance the predictability of the system and to scale up the generalizability and trustworthiness of the diagnostic systems. We would like to thank everyone who has supported and informed us throughout their research works, colleagues, and institutions, whose contributions have had a significant impact on this work. Their commitment to the scientific knowledge growth still motivates the creation of new effective and efficient solutions in the medical information analysis. We hope that the approaches, results, and discussions of this book will provide a good starting point of future studies and implementation of predictive cardiology and health informatics. It is with this purpose that we submit this work to the scientific community with hope that it will aid in the academic exploration as well as practical developments. We extend our wishes to all the readers and other researchers in their work. Wishing all our colleague success in their work. Sincerely, SELİM BUYRUKOĞLU, GONCA BUYRUKOĞLU
6 ÖZET Bu çalışmada, kalp yetmezliğinin erken tespitine yönelik makine öğrenmesi (ML) modellerinin performansını artırmak amacıyla meta-sezgisel optimizasyon algoritmaları (Parçacık Sürüsü Optimizasyonu (PSO), Genetik Algoritma (GA), Yapay Arı Kolonisi (ABC), Benzetimli Tavlama (SA) ve Karınca Kolonisi Optimizasyonu (ACO)) kullanılmıştır. Temel ML modelleri olan Rastgele Orman (RF), Destek Vektör Makineleri (SVM), Gradyan Artırma Makineleri (GBM), Karar Ağaçları (DT) ve Yapay Sinir Ağları (ANN), hem ham halleriyle hem de sezgisel algoritmalarla optimize edilmiş formlarıyla, iki açık veri kümesi üzerinde (UCI Kalp Hastalığı Veri Kümesi ve Kaggle Kalp Hastalığı Veri Kümesi) test edilmiştir. Bulgular, özellikle RFSA (Rastgele Orman – Benzetimli Tavlama) modelinin, DATASET-1 üzerinde 91,30 doğruluk ve 0,9047 F1-skoru ile en başarılı sonuçları verdiğini göstermektedir. Bunun yanında, RF ve SVM-PSO modelleri de Heart veri kümesinde 91,30 doğruluk elde ederek optimize edilmiş hibrit modellerin tahmin performansını belirgin şekilde artırdığını ortaya koymuştur. Optimizasyonun model genellenebilirliği üzerindeki etkisi, farklı veri kümeleri üzerinde gerçekleştirilen karşılaştırmalı analizlerle görülmektedir. Bu çalışma, daha geçerli ve hassas kalp yetmezliği tahmin sistemlerinin geliştirilmesine yönelik önemli katkılar sunmaktadır. Anahtar Kelimeler – Kalp Yetmezliği, Makine Öğrenmesi, Meta-Sezgisel Algoritmalar, Optimizasyon, PSO, GA, ABC, SA, ACO.
7 ABSTRACT In this paper, metaheuristic optimization algorithms (Particle Swarm Optimization (PSO), Genetic Algorithm (GA), Artificial Bee Colony (ABC), Simulated Annealing (SA), and Ant Colony Optimization (ACO)) are used to improve the performance of machine learning (ML) models in terms of early detection of heart failure. Core ML models like the Random Forest (RF), Support Vector Machines (SVM), Gradient Boosting Machines (GBM), Decision Trees (DT) and the Artificial Neural Networks (ANN) are tested in their raw forms and again optimized through heuristic algorithms, on two publicly available datasets (the UCI Heart Disease Dataset and the Kaggle Heart Disease Dataset). The findings indicate that heuristic optimization, especially the RF-SA (Random Forest - Simulated Annealing) model, with 91.30 and F1-score of 0.9047 on DATASET-1, has the best results. Besides, the RF and SVM-PSO models achieve an accuracy of 91.30 on the Heart dataset, which is significantly higher in predictive accuracy when using optimized hybrid models. The impact of optimization on the generalizability of the model is seen upon the comparative analysis of various datasets. This paper provides one of the key contributions to building more valid and precise heart failure prediction systems. Keywords – Heart Failure, Machine Learning, Metaheuristic Algorithms, Optimization, PSO, GA, ABC, SA, ACO.
8 INTRODUCTION Cardiovascular diseases (CVDs) are a wide category of diseases of the heart and blood vessels including coronary artery disease, heart failure, arrhythmias, valvular heart disease and peripheral vascular disease that weaken the effectiveness of cardiovascular system functioning. The conditions are usually as a result of a gradual accumulation of atherosclerotic plaques in the arterial walls resulting in little blood supply, ischemia, and even possible cardiac death, stroke, or myocardial infarction. CVDs are the most common cause of mortality in the world with an estimated 19.8 million deaths per year (or about 32 percent of the total number of deaths), and a tremendous socioeconomic cost, in terms of health care costs and lost productivity and disability-adjusted life years. In spite of the considerable improvement in medical interventions and population health campaigns, people remain increasingly affected by CVDs, especially in the lowand middle-income countries, in which access to early screening and control is still low. The most prominent examples of risk factors that can be modified and influence this epidemic are hypertension, dyslipidemia (elevated levels of low-density lipoprotein cholesterol and lower levels of high-density lipoprotein cholesterol), diabetes mellitus, obesity, smoking, physical inactivity, poor diet (rich in saturated fats, trans fats, sodium, and sugars), and excessive alcohol intake, with the additional factors of vulnerability being advanced age, being male up to menopause (in women), family history Multifactorial etiology of CVDs can also be supported by emerging evidence that shows the contribution of chronic inflammation, chronic infections (e.g., Chlamydia pneumoniae or Helicobacter pylori), autoimmune disorders (e.g., systemic lupus erythematosus), psychosocial stressors such as chronic stress or depression, and environmental exposures in the acceleration of disease progression. In such large proportions, the ability to detect the disease early and exactly evaluate the risks is the most crucial factor, which has a direct relationship with the effectiveness of the therapeutic operations, better patient results, and fewer lives are lost because of the preventable ones. This need leads to the immediate design of high-quality, reliable, and scalable diagnostic instruments to improve the stratification of patients and clinical decisionmaking (World Health Organization, 2025).
15 Figure 1. Proposed Model Structure Datasets In this book, two publicly available datasets were used: Dataset-1: This is a dataset that was acquired at the UCI Machine Learning Repository and that has been extensively applied in literature (Janosi et al., 1989). It has 303 patient records and 13 clinical characteristics (age, sex, type of chest pain, resting blood pressure, etc.). The target variable shows the existence (1) or the absence (0) of heart disease in the patient. This data was available by using ucimlrepo library. The detailed descriptions of all variables used in this study are presented in Table 1.
16 Table 1. Feature definitions and data types for the Dataset-1 Feature Name Description Type age Age of the patient (in years) Numerical sex Sex (1: male, 0: female) Categorical cp Chest pain type (1: typical angina, 2: atypical angina, 3: non-anginal pain, 4: asymptomatic) Categorical trestbps Resting blood pressure (mm Hg on admission to the hospital) Numerical chol Serum cholesterol (mg/dl) Numerical fbs Fasting blood sugar > 120 mg/dl (1: true, 0: false) Categorical restecg Resting electrocardiographic results (0: normal, 1: STT wave abnormality, 2: left ventricular hypertrophy) Categorical thalach Maximum heart rate achieved Numerical exang Exercise-induced angina (1: yes, 0: no) Categorical oldpeak ST depression induced by exercise relative to rest Numerical slope Slope of the peak exercise ST segment (1: upsloping, 2: flat, 3: downsloping) Categorical ca Number of major vessels (0-3) colored by fluoroscopy Numerical/Categorical thal Thalassemia (3: normal, 6: fixed defect, 7: reversible defect) Categorical num Target: Presence of heart disease (0: no, 1-4: yes Categorical
17 with severity; often binarized to 0/1) Dataset-2: It was obtained on the Kaggle open-source database and includes 918 patient records and 12 clinical features in total (fedesoriano, 2021). Like Dataset-1, this dataset is aimed at the diagnosis of heart disease (binary classification). This dataset has a larger sample size, which was used to test the generalizability ability of the models. A summary of the dataset characteristics is given in Table 2. Table 2. Feature definitions and data types for the Dataset-2 Feature name Description (short) Type Age Age of the patient (years) Numerical Sex Sex (M: male, F: female) Categorical ChestPainType Type of chest pain (e.g., ATA, NAP, ASY, TA) Categorical RestingBP Resting blood pressure (mm Hg) Numerical Cholesterol Serum cholesterol level (mg/dl) Numerical FastingBS Fasting blood sugar > 120 mg/dl (1: true, 0: false) Categorical RestingECG Resting electrocardiogram results (Normal, ST, LVH) Categorical MaxHR Maximum heart rate achieved Numerical ExerciseAngina Exercise-induced angina (Y: yes, N: no) Categorical Oldpeak ST depression induced by exercise relative to rest Numerical ST_Slope Slope of the peak exercise ST segment (Up, Flat, Down) Categorical HeartDisease Target variable: presence of heart disease (1: yes, 0: no) Categorical
18 Data Preprocessing To guarantee the quality of data and its compatibility with two different sources as well as the optimal performance of the machine learning models, a common and strict preprocessing pipeline was selected in both datasets. All the variables of the two datasets were a combination of categorical and numerical variables (Zhu et al., 2024). Label Encoding was used to encode the categorical variables into numeric forms. This is essential in preparing non-numeric data to be used with most machine learning programs. Missing Value Imputation: In order to overcome the problem of missing data, imputation was done through the median strategy. It is a superior method to the mean as it is very strong against outliers and is usually vital when dealing with clinical data which normally have skewed distributions (Alam et al., 2023). Feature Scaling and Standardization: After the imputation, all the features got to be standardized with the help of the StandardScaler tool. This is done so that all the features have a mean of 0 and standard deviation of 1 so as to balance out the scales. Standardization of features is also primary in both distance-based algorithms (e.g., K-Nearest Neighbors) and gradient-based algorithms (e.g., Logistic Regression, Neural Networks) since it ensures that large-valued features do not overtake the learning process (Pinheiro et al., 2025). Data Splitting: The processed datasets were broken down into three different subsets i.e., training (70%), validation (15%), and test (15%) sets. This divide and rule approach permits: • Training: Training the model parameters using a maximum amount of the data. • Validation: Hyperparameter tuning of the model and model selection without biasing the model with the last test set. • Testing: Giving a fair analysis of how the final model would perform on an unobserved data.
19 Machine Learning Models The base line ML models that have been used in this research have the following individual benefits in terms of classification functions: Random Forest (RF): This is an ensemble algorithm of learning which uses the forecasts of a large number of decision trees by incorporating them bootstrap aggregating (bagging) to provide high accuracy and generalization abilities. In training, RF builds many decision trees and the mode of their respective predictions is produced, which minimizes overfitting and improves noise-resistance (Breiman, 2001). Its capability to deal with high-dimensional data and feature ranking of importances makes it especially appropriate to use in medical diagnosis. Support Vector Machine (SVM): It is a discriminative classifier, which determines the most optimal separating hyperplane so that it maximizes the separation between distinct classes in the high-dimensional feature spaces. SVM uses the operations of the kernel functions (e.g., radial basis function) to map non-linearly separable data into the higher dimensional space where they become linearly separable (Cortes & Vapnik, 1995). The model can be useful in binary classification where the decision boundaries are complicated. Gradient Boosting Machine (GBM): This is a potent ensemble method because the learners are trained one after another with each successive model attempting to remedy the errors of the last learner. GBM uses the gradient descent to reduce a loss function, which is gradually improved by increasing the predictive accuracy using an additive modeling strategy (Friedman, 2002). The process of this iterative refinement has proven to be more accurate in many medical prediction issues. Artificial Neural Network (ANN): This is a deep learning model based on the biological neural network, which consists of oppositely coupled layers of neurons that can learn a complex and non-linear pattern by backpropagation. Activation functions and adaptive weight modulations used in ANNs are meant to represent complex relationships of data (Aggarwal, 2018). Their ability to be flexible in identifying secrets makes them useful in medical diagnostics, albeit they must be tuned carefully in order to avoid overfitting.
20 Metaheuristic Optimization Algorithms In order to improve the performance of the ML models, the following metaheuristic algorithms were used. The main aim of these algorithms was to optimize hyperparameters i.e. to optimize the model parameters (n estimators, max depth, learning rate, C, gamma) to obtain maximum accuracy on the validation set. In comparison to classical grid search or random search algorithms, metaheuristic algorithms are effective to search high-dimensional hyperparameter space and find near-optimal settings. Particle Swarm Optimization (PSO): PSO is based on flock behaviour of a bird/fish school and it uses a population of particles, each representing a possible solution. Every particle keeps personal best position and the global best position found by swarm which show quick convergence ability using velocity as well as position updates (Kennedy and Eberhart, 1995). PSO is especially useful with continuous optimization, and has been used in large numbers in hyperparameter tuning in ML. Genetic Algorithm (GA): It is an evolutionary algorithm based on natural selection and genetic processes, that is, there are crossover and mutation operators. GA has a population of candidate solutions and evolves them by means of selection, recombination and mutation to allow a comprehensive search and find the best solution (Forrest, 1996). It can be used in optimization complex terrains owing to its capacity to overcome local optima. Artificial Bee Colony (ABC): Following the foraging strategy of honeybees, ABC strikes a balance between exploration and exploitation stages in the quest to find the global optimum. The algorithm uses employed bees, onlooker bees, scout bees, each of which performs different functions in the search of high quality food sources (solutions). ABC has proven to be efficient in optimization of ML hyperparameters because of the efficient search mechanism (Karaboga and Basturk, 2007). Simulated Annealing (SA): Basing this algorithm on the metallurgical annealing process, SA accepts worse solutions with a probabilistic nature, and the probability of this acceptance decreases as the algorithm gets closer to convergence. This form of exploration uses temperature control to make SA search through complicated search space and arrive at the global optimum
21 (Kirkpatrick et al., 1983). SA has been effective in cardiac healthcare ML applications. Ant Colony Optimization (ACO): ACO is a model which imitates the pheromone-based path-finding of ant colonies to build a solution by laying down pheromone trails which is an indication of solution quality and traversing them. The method works especially well on discrete optimization and search spaces that are discrete (Dorigo and Stutzle, 2019). Performance Metrics In order to measure the performance of the models, the following metrics were used: 𝐴𝑐𝑐𝑢𝑟𝑎𝑐𝑦 = TP+TN TP+TN+FP+FN (1) 𝑅𝑒𝑐𝑎𝑙𝑙 = TP TP+FN (2) 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 = TP TP+FP (3) 𝐹1 𝑠𝑐𝑜𝑟𝑒 = 2 ∗ Precision∗Recall Precision+Recall (4) RESULTS AND DISCUSSION Results on Dataset-1 The findings of Dataset-1 are shown in Table 3. This table is a summary of the performance metrics of the baseline ML models and their metaheuristicoptimized version, namely their accuracy, precision, recall and F1-score as evaluated on Dataset-1. Table 3 data offer a direct comparison of the effectiveness of each model in this case under various forms of optimization strategies, and therefore, the most appropriate method of predicting heart disease was possible through this dataset.
22 Table 3: Model Performance on Dataset-1 Heart Disease Dataset Model Accuracy Precision Recall F1-Score RF 0.891304 0.863636 0.904762 0.883721 RF-PSO 0.891304 0.863636 0.904762 0.883721 RF-GA 0.891304 0.863636 0.904762 0.883721 RF-ABC 0.891304 0.863636 0.904762 0.883721 RF-SA 0.913043 0.904762 0.904762 0.904762 RF-ACO 0.869565 0.826087 0.904762 0.863636 SVM 0.804348 0.750000 0.857143 0.800000 SVM-PSO 0.869565 0.857143 0.857143 0.857143 SVM-GA 0.847826 0.791667 0.904762 0.844444 SVM-ABC 0.826087 0.782609 0.857143 0.818182 SVM-SA 0.826087 0.760000 0.904762 0.826087 SVM-ACO 0.804348 0.700000 1.000000 0.823529 GBM 0.804348 0.730769 0.904762 0.808511 GBM -PSO 0.804348 0.730769 0.904762 0.808511 GBM -GA 0.826087 0.760000 0.904762 0.826087 GBM -ABC 0.739130 0.666667 0.857143 0.750000 GBM -SA 0.804348 0.750000 0.857143 0.800000 GBM -ACO 0.739130 0.666667 0.857143 0.750000 ANN 0.847826 0.818182 0.857143 0.837209 ANN -PSO 0.869565 0.826087 0.904762 0.863636 ANN -GA 0.847826 0.791667 0.904762 0.844444 ANN -ABC 0.804348 0.730769 0.904762 0.808511 ANN -SA 0.847826 0.791667 0.904762 0.808511 ANN-ACO 0.847826 0.791667 0.904762 0.808511 RF-SA model had the highest performance on Dataset-1 with an accuracy of 91.30 and F1-score of 90.47%. RF-ABC model has acquired the highest ROC-AUC value of 94.47%. These results establish that metaheuristic
23 optimization largely improves the rank of prediction especially in ensemble learning models like Random Forest. As an example, the optimal RF-ABC model had a ROC-AUC of 0.9447 compared to the baseline RF model which had 0.9333, which is a significant gain in discriminative power. Figure 2 presents a visual comparison of the assessment measures of all the models performing on Dataset-1 that gives a direct picture of the performance variation between the baseline and optimized strategies. Figure 2. Dataset-1 Evaluation Metric Scores Figure 3 shows the comparison of the baseline ML models with their metaheuristic-optimized counterparts in Dataset-1, which shows that optimization is able to contribute to the improvement in performance.
24 Figure 3. Dataset-1 Base Models and Metaheuristic Optimization Figure 4 illustrates the SHAP decision plot of Dataset-1, which is a plot that illustrates the contribution of the individual features in model predictions by mapping the SHAP values of each feature from the base value to the final prediction. Figure 4. Dataset-1 SHAP Decision Plot
31 of local optima can be viewed as the key factor that led to this success. These results are consistent with the published literature that proves the superiority of metaheuristic algorithms in terms of accuracy of the model achieved by the effective selection of features and hyperparameter optimization. CONCLUSION AND FURTHER DIRECTIONS This book has shown clearly that the implementation of meta-heuristic optimization algorithms alongside the traditional machine learning models would widely benefit performance in predicting the achievement of heart failure classification tasks. In both datasets, the optimized hybrid models have better accuracy, recall, F1-score, and overall discriminative ability as compared to their baseline counterparts. The latter improvements are especially apparent in the context of medical data, which is highly complex and heterogeneous, and the nonlinear interactions and delicate feature dependencies of such data can restrict the usefulness of traditional modeling methods. Through a series of comparisons of various algorithms on different data sources, the present study can demonstrate that meta-heuristic optimization does not only fine-tune hyperparameters of different models in a more efficient way, but also leads to more meaningful generalization when the population features vary. The other significant implication of the current research is the variability of the algorithmic performance among datasets with an accent on the fact that not one combination of optimization models can guarantee the highest accuracy. RF-based hybrid models performed better on Dataset-1 whereas SVM-based optimized models performed better on Data set-2. This would imply that decoupling of data structure, distribution of features and classifier architecture is a significant factor in deciding on the best hybrid strategies. These findings, therefore, support the significance of datasetspecific optimization, as opposed to using a generalized methodology in clinical decision-support systems. These insights can inform the practitioners and researchers to choose more suitable optimization methods depending on the complexity of their datasets and clinical setting. In the future, the study of optimized hybrid models should be extended by including wider and more heterogeneous datasets in order to better support
32 the validity of the generalizability of the study. Recent big-data, multiinstitutional studies, particularly those that involve time-series, physiological measurements of cardiovascular diseases, biomarkers in imaging, or genomics, can provide a more comprehensive understanding of the predictive processes of cardiovascular diseases. The ability to incorporate multimodal data into meta-heuristic-optimized models would also help a lot in enabling capability at the early diagnosis, especially in detecting high-risk subgroups that might not be adequately captured by the conventional models. This growth would not just enhance the performance of algorithms, but would also improve the accuracy of the algorithms in the real-world clinical settings where the data are usually multidimensional and it is also erratic. Moreover, future research would be enhanced by considering superior optimization paradigms like multi-objective optimization, reinforcementlearning-based optimization and hybrid meta-heuristic ensembles, which have the potential to trade off between precision, explainability, and computation efficiency. The other avenue worth pursuing is the use of meta-heuristic algorithms in other areas of application than hyperparameter tuning, including dynamic feature selection, neural architecture search, or data augmentation strategies that are automatically selected. Those strategies would support more dynamic and scalable diagnostic models that can be implemented into realtime clinical processes. Also, more focus is to be put on creating elucidable and explicable systems. Explainability tools can help clinicians to gain insight into how their model works, assess its credibility, and incorporate model results into patient-centered decision-making more readily as exemplified with SHAP-based analyses. In conclusion, the research demonstrates the being-transformative power of meta-heuristic optimization in terms of the development of machine learning to be used in predictive cardiology. Optimized hybrid models are a promising future in intelligent healthcare systems by improving the performance of models, improving generalizability, and able to analyze more complex clinical data in more subtle ways. Increased interdisciplinary collaboration, i.e. a merger of talent and expertise in the field of computer science, medicine, and data analytics, will be instrumental in transforming such developments in methods into clinical applications. In the long run, improved and increased applications of such methods will lead to the realization of more effective reliable, interpretable and scalable diagnostic
33 tools that can help deal with the increasing global burden of cardiovascular diseases. REFERENCES Aggarwal, C. C. (2018). An Introduction to Neural Networks. Neural Networks and Deep Learning, 1–52. https://doi.org/10.1007/978-3-319-94463-0_1 Al Jowf, G. I., Kolhar, M., Hofuf, A., Ahsa, A., & Arabia, S. (2025). Key factors in predictive analysis of cardiovascular risks in public health. Scientific Reports 2025 15:1, 15(1), 23455-. https://doi.org/10.1038/s41598-025-07874-x Alam, S., Ayub, M. S., Arora, S., & Khan, M. A. (2023). An investigation of the imputation techniques for missing values in ordinal data enhancing clustering and classification analysis validity. Decision Analytics Journal, 9, 100341. https://doi.org/10.1016/J.DAJOUR.2023.100341 Asadi, S., Roshan, S. E., & Kattan, M. W. (2021). Random forest swarm optimizationbased for heart diseases diagnosis. Journal of Biomedical Informatics, 115, 103690. https://doi.org/10.1016/J.JBI.2021.103690 Ay, Ş., Ekinci, E., & Garip, Z. (2023). A comparative analysis of meta-heuristic optimization algorithms for feature selection on ML-based classification of heart-related diseases. The Journal of Supercomputing, 79(11), 11797. https://pmc.ncbi.nlm.nih.gov/articles/PMC9983547/ Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324 Chen, L., Ji, P., Ma, Y., Rong, Y., & Ren, J. (2023). Custom machine learning algorithm for large-scale disease screening - taking heart disease data as an example. Artificial Intelligence in Medicine, 146, 102688. https://doi.org/10.1016/J.ARTMED.2023.102688 Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20(3), 273–297. https://doi.org/10.1007/bf00994018 DeGoma, E. M., Knowles, J. W., Angeli, F., Budoff, M. J., & Rader, D. J. (2012). The Evolution and Refinement of Traditional Risk Factors for Cardiovascular Disease. Cardiology in Review, 20(3), 118. https://doi.org/10.1097/CRD.0B013E318239B924 fedesoriano. (2021, September). Heart Failure Prediction Dataset. https://www.kaggle.com/datasets/fedesoriano/heart-failure-prediction/data Friedman, J. H. (2002). Stochastic gradient boosting. Computational Statistics & Data Analysis, 38(4), 367–378. https://doi.org/10.1016/S0167-9473(01)00065-2 Hagan, R., Gillan, C. J., & Mallett, F. (2021). Comparison of machine learning methods for the classification of cardiovascular disease. Informatics in Medicine Unlocked, 24, 100606. https://doi.org/10.1016/J.IMU.2021.100606
34 Janosi, A., Steinbrunn, W., Pfisterer, M., & Detrano, R. (1989). Heart Disease [Dataset]. UCI Machine Learning Repository. https://doi.org/https://doi.org/10.24432/C52P4X. Kumar, R., Garg, S., Kaur, R., Johar, M. G. M., Singh, S., Menon, S. V., Kumar, P., Hadi, A. M., Hasson, S. A., & Lozanović, J. (2025). A comprehensive review of machine learning for heart disease prediction: challenges, trends, ethical considerations, and future directions. Frontiers in Artificial Intelligence, 8, 1583459. https://doi.org/10.3389/FRAI.2025.1583459/FULL Mall, S., & Singh, J. (2024). Comparative Study of Heart Failure Using the Approach of Machine Learning and Deep Neural Networks. Lecture Notes in Networks and Systems, 785, 627–641. https://doi.org/10.1007/978-981-99-65441_47/TABLES/2 Narasimhan, G., & Victor, A. (2025). A hybrid approach with metaheuristic optimization and random forest in improving heart disease prediction. Scientific Reports 2025 15:1, 15(1), 10971-. https://doi.org/10.1038/s41598-024-73867x Nouna, Z., Nouna, Z., Bouyghf, H., Nahid, M., & Sabiri, I. (2025). Heart disease prediction optimization using metaheuristic algorithms. IAES International Journal of Artificial Intelligence (IJ-AI), 14(5), 4332–4341. https://doi.org/10.11591/ijai.v14.i5.pp4332-4341 Pinheiro, J. M. H., de Oliveira, S. V. B., Silva, T. H. S., Saraiva, P. A. R., de Souza, E. F., Godoy, R. V., Ambrosio, L. A., & Becker, M. (2025). The Impact of Feature Scaling In Machine Learning: Effects on Regression and Classification Tasks. IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, XX, 1. https://doi.org/10.1109/ACCESS.2025.3635541 Saboor, A., Usman, M., Ali, S., Samad, A., Abrar, M. F., & Ullah, N. (2022). A Method for Improving Prediction of Human Heart Disease Using Machine Learning Algorithms. Mobile Information Systems, 2022(1), 1410169. https://doi.org/10.1155/2022/1410169 Seslier, T., & Karakuş, M. Ö. (2022). IN HEALTHCARE APPLICATIONS of MACHINE LEARNING ALGORITHMS for PREDICTION of HEART ATTACKS. Journal of Scientific Reports-A, 051, 358–370. https://dergipark.org.tr/en/pub/jsr-a/issue/73021/1191927 Sowmiya, M., Banu Rekha, B., & Malar, E. (2025). Optimized heart disease prediction model using a meta-heuristic feature selection with improved binary salp swarm algorithm and stacking classifier. Computers in Biology and Medicine, 191, 110171. https://doi.org/10.1016/J.COMPBIOMED.2025.110171 Teja, M. D., & Rayalu, G. M. (2025). Optimizing heart disease diagnosis with advanced machine learning models: a comparison of predictive performance. BMC Cardiovascular Disorders, 25(1), 212-. https://doi.org/10.1186/S12872025-04627-6/TABLES/6 World Health Organization. (2025, November 26). Cardiovascular diseases (CVDs). https://www.who.int/news-room/fact-sheets/detail/cardiovascular-diseases-
35 (cvds) Yang, L., Wu, H., Jin, X., Zheng, P., Hu, S., Xu, X., Yu, W., & Yan, J. (2020). Study of cardiovascular disease prediction model based on random forest in eastern China. Scientific Reports 2020 10:1, 10(1), 5245-. https://doi.org/10.1038/s41598-020-62133-5 Zhang, H., & Mu, R. (2024). Refining heart disease prediction accuracy using hybrid machine learning techniques with novel metaheuristic algorithms. International Journal of Cardiology, 416, 132506. https://doi.org/10.1016/J.IJCARD.2024.132506 Zhu, W., Qiu, R., & Fu, Y. (2024). Comparative Study on the Performance of Categorical Variable Encoders in Classification and Regression Tasks.