Vollständiger Abstract
Worum geht es in dieser Arbeit?
Medicaid improper payments in the United States reached an estimated $37.39 billion (6.12%) in fiscal year 2025, up from $31.10 billion (5.09%) the prior year, with more than three-quarters attributable to insufficient documentation rather than confirmed fraud — precisely the gap that Electronic Visit Verification (EVV) systems are federally mandated to close. This study develops and rigorously validates a machine learning framework for provider-level Medicaid and Medicare fraud detection and establishes its relevance to EVV-based program integrity. A 17-dimensional provider-level feature set spanning claim-volume, financial, physician-network, temporal, clinical-complexity, and patient-demographic domains was engineered from real, publicly available claims data covering 5,410 providers, 40,474 inpatient claims, 517,737 outpatient claims, and 138,556 beneficiary records, with fraud labels derived using the Office of Inspector General’s List of Excluded Individuals and Entities methodology. Logistic Regression, Random Forest, XGBoost, and an unsupervised Isolation Forest were trained and compared under stratified five-fold cross-validation, paired statistical significance testing, systematic ablation analysis, and synthetic noise and missing-data robustness testing. Random Forest achieved the strongest held-out F1-score of 0.649 with an area under the ROC curve of 0.950, statistically indistinguishable from XGBoost. Explainability analysis using SHAP identified total reimbursement, average length of stay, and claims-per-beneficiary as the dominant predictive signals, while ablation analysis showed that temporal features contributed the single largest performance drop when removed, directly motivating an extension toward EVV-based visit-timing data. Robustness testing demonstrated graceful, quantified performance degradation under both synthetic feature noise and simulated missing data, and structured error analysis revealed that the framework under-detects lower-volume fraud relative to high-volume fraud, an honestly reported and actionable limitation. Overall, the findings demonstrate that a compact, explainable, and statistically validated machine learning framework can achieve strong, stable, and computationally efficient fraud discrimination, providing a methodologically sound foundation for extending claims-based fraud analytics into real-time, EVV-integrated Medicaid program-integrity systems that protect public healthcare spending at national scale.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Md Sibbir Hossain, A T M Fokrule Hasan
- Quelle
- Frontiers in Computer Science and Artificial Intelligence
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2978-8048
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Md Sibbir Hossain, A T M Fokrule Hasan (2026). Safeguarding Public Healthcare Spending with Explainable AI: A Statistically Validated Machine Learning Framework for Medicaid Fraud Detection and Electronic Visit Verification. Frontiers in Computer Science and Artificial Intelligence. https://doi.org/10.32996/jcsts.2026.5.3.4
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1