Heart Disease Prediction Using Artificial Intelligence Algorithms: A Comparative Study
DOI:
https://doi.org/10.64149/Keywords:
Heart Disease Prediction, Machine Learning, Comparative Study, Cross-Validation, Feature Importance, Clinical Decision Support, Explainable Ai.Abstract
Cardiovascular disease is the leading contributor to global mortality, and timely risk identification based on existing data can significantly improve clinical outcomes and decrease the cost of invasive diagnostic workups. In this study, synchronous comparative analysis of the five supervised machine learning algorithms-Logistic Regression, Random Forest, Support Vector Machine (SVM), K-Nearest Neighbors (KNN) and Gradient Boosting-to predict binary heart disease prediction is provided. We used a combined clinical dataset, which consists of 920 patient records collected from four sites (Cleveland, Hungary, VA Long Beach and Switzerland), with 13 raw clinical features extended to a total of 27 following one-hot encoding. For a comprehensive pipeline we performed physiologically informed data cleaning, median and mode imputation for continuous and categorial variables respectively, standardization of continuous variables prior to modeling, use of GridSearchCV to tune hyperparameters, plus combined paired significance testing against all held-out testing (as in a typical predictive study), supported by five-fold cross-validation. Random Forest had the highest held-out test accuracy (84.2%) and SVM fit to highest cross-validated accuracy (82.5% ± 2.8%) while paired t-tests demonstrated that neither of the two models could be statistically distinguished, both together forming a best-performing group. The most predictive features were age, maximum heart rate, cholesterol, chest pain type and ST depression. The greatest values of this study highlight the importance of cross-validated, statistically based model selection instead of single-split accuracy.



