Vollständiger Abstract
Worum geht es in dieser Arbeit?
Diabetes prediction plays an important role in re-ducing long-term health risks by enabling early medical interven-tion. Although machine learning models have been widely applied to this task, many existing studies emphasise predictive accuracy while giving comparatively little attention to the reliability, interpretability, and stability of the resulting decisions. This paper develops a reliability-aware and interpretable machine learning framework for diabetes prediction from structured clinical data. Three complementary models—Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost)—are trained on the Pima Indians Diabetes dataset so that both simple linear and complex non-linear relationships are captured. Beyond conventional discrimination metrics, the reliability of the predicted probabilities is quantified using the Brier score and reliability (calibration) diagrams. Interpretability is addressed with SHapley Additive exPlanations (SHAP) at both the global (cohort) and local (individual patient) levels. Because different models frequently emphasise different predictors, we formalise a Feature Consistency Index (FCI) that quantifies the cross-model agreement of SHAP-derived feature importance and combines it with normalised importance into a single ranking score. Finally, a perturbation-based robustness analysis measures the sensitivity of each model’s output to small changes in the input record. Experi-mentally, XGBoost achieves the highest discrimination (accuracy 0.7597, ROC-AUC 0.8374), whereas Random Forest attains the best-calibrated probabilities (Brier score 0.1646), demonstrating that discrimination and reliability are not interchangeable. The FCI identifies Glucose and BMI as simultaneously the most influential and the most consistently attributed predictors, while Blood Pressure and Skin Thickness are both weak and unstable. Under a 5% Gaussian perturbation of a representative patient record, the linear and bagged models shift by less than 0.01 in predicted probability, whereas the boosted model shifts by 0.0386, revealing an accuracy–stability trade-off that a purely accuracy-driven evaluation would not expose
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Ravikumar V, Dr. S. Sasirekha
- Quelle
- International Journal for Research in Applied Science and Engineering Technology
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2321-9653
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Ravikumar V, Dr. S. Sasirekha (2026). A Reliability-Aware and Interpretable Machine Learning Framework for Diabetes Prediction Using Structured Clinical Data. International Journal for Research in Applied Science and Engineering Technology. https://doi.org/10.22214/ijraset.2026.84701