Vollständiger Abstract
Worum geht es in dieser Arbeit?
Chronic illness burden is unevenly distributed across US counties, with some counties experiencing overlapping challenges related to healthcare access, socioeconomic disadvantage, environmental exposure, and chronic disease burden. Several multidimensional frameworks, including the Social Vulnerability Index, Environmental Justice Index, and Pandemic Vulnerability Index, have highlighted the importance of integrating multiple determinants of vulnerability. This study presents an interpretable machine learning-based multi-model framework for constructing and characterizing a county-level chronic illness vulnerability index. The framework integrates five domains: health burden, social conditions, environmental exposure, behavioral factors, and COVID-19-related measures. Four publicly available 2023 datasets were merged into a 12-feature dataset covering 2911 US counties, and a weighted domain score was used to classify counties into low-, medium-, and high-vulnerability categories. Because the class labels were constructed directly from these variables, the resulting categories should be interpreted as an author-defined composite index rather than an independent health outcome. Four machine-learning models, including Random Forest, XGBoost, MLP (Multilayer Perceptron), and TabNet, were evaluated individually and as a majority-vote ensemble. SHAP and LIME were independently applied to interpret model behavior, and the explainability analyses were restricted to model interpretation rather than causal inference. The ensemble reproduced the three-tier classification with an accuracy of 94.68% and an ROC-AUC (Receiver Operating Characteristic–Area Under the Curve) of 0.9974 on 583 held-out counties. Because the class label is a deterministic function of the same 12 features, these results should be interpreted as measures of index reconstruction rather than prediction of an independent outcome; a penalized multinomial logistic regression achieved 98.97% reconstruction accuracy, further indicating the largely linear structure of the index. COPD prevalence ranked first according to both SHAP and LIME across all four models, reflecting the weighting structure of the index rather than an independent causal effect. Six of the twelve features received identical rankings from both methods, and no feature moved more than three positions. The framework relies entirely on publicly available data and provides a reproducible approach for county-level characterization of chronic illness vulnerability.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Ronish Shrestha, Md Masud Rana, Bo Sun, Margot Gage Witvliet, Raouth Kostandy, Timothy Roden, Mohammad Saquib
- Quelle
- International Journal of Environmental Research and Public Health
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 1660-4601
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Ronish Shrestha, Md Masud Rana, Bo Sun, Margot Gage Witvliet, Raouth Kostandy, Timothy Roden, Mohammad Saquib (2026). An Interpretable Machine Learning-Based Multi-Model Framework for Constructing and Characterizing a County-Level Chronic Illness Vulnerability Index. International Journal of Environmental Research and Public Health. https://doi.org/10.3390/ijerph23091104
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1