Vollständiger Abstract
Worum geht es in dieser Arbeit?
Background Stigmatizing language (SL) in electronic health records (EHRs) can influence clinical decision-making, propagate bias across care encounters, and undermine patient trust. Gender-expansive patients (GEPs) may be particularly vulnerable to documentation-based stigma; however, large-scale quantitative evidence and fairness-aware evaluation of automated SL detection methods remain limited. Objective This study aims to construct a gender-expansive-inclusive EHR corpus, quantify demographic disparities in SL using parallel outcome definitions with and without misgendering, and evaluate fairness-aware natural language processing (NLP) methods for automated detection of SL. Methods We developed an annotated corpus of 754 clinical notes from the MIMIC-IV (Medical Information Mart for Intensive Care IV) database, including 366 GEP notes and 388 matched nongender-expansive patient (NGEP) notes, labeled for SL and its subtypes. Parallel outcome definitions were constructed with and without misgendering. Multivariable logistic regression was used to assess associations between gender-expansive status and SL while adjusting for race, age, and primary language. Multiple NLP models were evaluated for SL detection. Fairness-aware post hoc threshold optimization based on equalized odds principles was applied using training-set predictions to reduce subgroup error disparities. Results SL was identified in 62.3% (228/366) of GEP notes compared with 25.5% (99/388) of NGEP notes. When misgendering was excluded, the prevalence remained higher among gender-expansive notes at 41.8% (153/366). In multivariable models, gender-expansive status was strongly associated with stigmatizing documentation when misgendering was included (adjusted odds ratio [OR] 4.87, 95% CI 3.54-6.70) and remained significant when misgendering was excluded (adjusted OR 2.12, 95% CI 1.54-2.91). Post hoc equalized odds threshold optimization for a state-of-the-art transformer-based detector reduced the difference in false-positive rate (ΔFPR) from 15.76 to 6.65 percentage points (pp) and the difference in true-positive rate (ΔTPR) from 7.22 pp to substantially lower levels while maintaining similar accuracy (82.78%-83.44%). When misgendering was excluded, fairness optimization reduced ΔFPR to 2.98 pp and ΔTPR to 0.65 pp, with an overall accuracy of 88.08%. Conclusions SL is common in EHR documentation and disproportionately affects GEPs, and automated detection models show persistent subgroup performance gaps. Disparities remained significant even when misgendering was excluded, indicating that bias extends beyond identity-specific errors to broader evaluative language. This study introduces the first annotated corpus focused on SL in GEP documentation, quantifies demographic disparities, and demonstrates practical fairness-aware NLP strategies that can reduce error-rate inequities while preserving accuracy. These findings support equity-focused interventions to address SL through fairness-aware models as assistive auditing tools with human oversight.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Liyang Xue, Mary Chayko, Vivek Kumar Singh
- Quelle
- Journal of Medical Internet Research
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 1438-8871
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Liyang Xue, Mary Chayko, Vivek Kumar Singh (2026). Stigmatizing Language in Gender-Expansive Patient Records: Corpus Development, Disparity Analysis, and Natural Language Processing–Based Detection Study. Journal of Medical Internet Research. https://doi.org/10.2196/91089