Towards Interpretable Alzheimer’s Disease Diagnosis: A Multi-Model Machine Learning Framework with Consensus Feature Selection and SHAP-Driven Explainability

Authors

  • Fatima Tauseef, Ahmad Jamal, Fahad Naseer, Md Shafayet Hossain, Dr. Mahzabin Mahmud Author

DOI:

https://doi.org/10.64149/

Keywords:

Alzheimer’s Disease; Machine Learning; Explainable Artificial Intelligence; Feature Selection; Lightgbm; Shap; Clinical Decision Support; Ensemble Learning.

Abstract

Alzheimer’s disease (AD) is a progressive neurodegenerative disorder representing the leading cause of dementia worldwide. Timely and accurate diagnosis remains a persistent clinical challenge owing to symptom overlap, diagnostic subjectivity, and the inaccessibility of neuro imaging-dependent workflows in resource-limited settings. This study proposes a comprehensive and clinically interpretable machine learning framework for automated AD classification using structured tabular clinical data drawn from a publicly available dataset comprising 2,149 patient records characterised by 34 demographic, lifestyle, clinical, and cognitive features. The proposed pipeline encompasses systematic data cleaning, categorical encoding, StandardScaler-based normalisation, and Synthetic Minority Oversampling Technique (SMOTE) application to address class imbalance. A novel consensus-driven multi-criteria feature selection strategy was developed, integrating Random Forest importance, Recursive Feature Elimination, and SHAP-based ranking into a unified normalised scoring mechanism to identify the 15 most diagnostically discriminative features. Eight classifiers spanning baseline and advanced architectural tiers were subsequently trained, optimised via stratified cross-validation, and evaluated across six performance metrics. LightGBM achieved the highest overall performance, attaining an accuracy of 94.88%, F1-Score of 92.62%, Cohen’s Kappa of 0.8871, and AUCROC of 0.9491. Statistical validation through 10-fold cross-validation and Wilcoxon signed-rank testing confirmed the superiority of gradientboosting ensemble models over baseline classifiers at the p < 0.05 significance level. Based on explainability analysis, further clinically meaningful feature attribution, and enhanced model transparency and interpretability. The proposed framework demonstrates strong diagnostic potential and offers a scalable, cost-effective, and interpretable alternative to imaging-dependent AD diagnostic pipelines.

Downloads

Published

2025-02-15

How to Cite

Towards Interpretable Alzheimer’s Disease Diagnosis: A Multi-Model Machine Learning Framework with Consensus Feature Selection and SHAP-Driven Explainability. (2025). Vascular and Endovascular Review, 8(2s), 345-357. https://doi.org/10.64149/