Article contents
Interpretable Ensemble Learning Approach for Breast Cancer Diagnosis Using SHAP-Based Explainable AI
Abstract
Early and reliable breast cancer diagnosis is clinically important, but machine learning models intended for healthcare use must offer both predictive accuracy and interpretability. This study evaluated a structured machine learning workflow on the Wisconsin Diagnostic Breast Cancer dataset. After preprocessing and label encoding, five machine learning classifiers—Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, and XGBoost—were implemented. Model performance was assessed using accuracy, precision, recall, F1-score, ROC-AUC, Brier score, and stratified cross-validation. SHAP was used to provide both global and local interpretability of the selected tree-based model. On the holdout set, Logistic Regression achieved the best overall discrimination and calibration, with 97.37% accuracy, 95.24% recall, ROC-AUC of 0.9960, and the lowest Brier score of 0.0211. SVM, Random Forest, and XGBoost each achieved 97.37% accuracy with perfect precision, but lower recall (92.86%). Cross-validation results also favored Logistic Regression, which achieved a mean ROC-AUC of 0.9955 ± 0.0038. SHAP analysis of XGBoost identified perimeter_worst, concave points_mean, concave points_worst, and radius_worst as the most influential predictors. The study demonstrates that transparent and comparatively simple machine learning models can achieve highly reliable classification performance on the WDBC benchmark, while SHAP improves interpretability of high-performing ensemble models. Although the results are promising, external clinical validation is required before real-world implementation.

Aims & scope
Call for Papers
Article Processing Charges
Publications Ethics
Google Scholar Citations
Recruitment