Frag' FlorenceEvidenz. Klar. Anwendbar.
Uhr 7/8Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Stigmatizing Language in Gender-Expansive Patient Records: Corpus Development, Disparity Analysis, and Natural Language Processing–Based Detection Study

Liyang Xue, Mary Chayko, Vivek Kumar Singh

Journal of Medical Internet Research · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Background Stigmatizing language (SL) in electronic health records (EHRs) can influence clinical decision-making, propagate bias across care encounters, and undermine patient trust. Gender-expansive patients (GEPs) may be particularly vulnerable to documentation-based stigma; however, large-scale quantitative evidence and fairness-aware evaluation of automated SL detection methods remain limited. Objective This study aims to construct a gender-expansive-inclusive EHR corpus, quantify demographic disparities in SL using parallel outcome definitions with and without misgendering, and evaluate fairness-aware natural language processing (NLP) methods for automated detection of SL. Methods We developed an annotated corpus of 754 clinical notes from the MIMIC-IV (Medical Information Mart for Intensive Care IV) database, including 366 GEP notes and 388 matched nongender-expansive patient (NGEP) notes, labeled for SL and its subtypes. Parallel outcome definitions were constructed with and without misgendering. Multivariable logistic regression was used to assess associations between gender-expansive status and SL while adjusting for race, age, and primary language. Multiple NLP models were evaluated for SL detection. Fairness-aware post hoc threshold optimization based on equalized odds principles was applied using training-set predictions to reduce subgroup error disparities. Results SL was identified in 62.3% (228/366) of GEP notes compared with 25.5% (99/388) of NGEP notes. When misgendering was excluded, the prevalence remained higher among gender-expansive notes at 41.8% (153/366). In multivariable models, gender-expansive status was strongly associated with stigmatizing documentation when misgendering was included (adjusted odds ratio [OR] 4.87, 95% CI 3.54-6.70) and remained significant when misgendering was excluded (adjusted OR 2.12, 95% CI 1.54-2.91). Post hoc equalized odds threshold optimization for a state-of-the-art transformer-based detector reduced the difference in false-positive rate (ΔFPR) from 15.76 to 6.65 percentage points (pp) and the difference in true-positive rate (ΔTPR) from 7.22 pp to substantially lower levels while maintaining similar accuracy (82.78%-83.44%). When misgendering was excluded, fairness optimization reduced ΔFPR to 2.98 pp and ΔTPR to 0.65 pp, with an overall accuracy of 88.08%. Conclusions SL is common in EHR documentation and disproportionately affects GEPs, and automated detection models show persistent subgroup performance gaps. Disparities remained significant even when misgendering was excluded, indicating that bias extends beyond identity-specific errors to broader evaluative language. This study introduces the first annotated corpus focused on SL in GEP documentation, quantifies demographic disparities, and demonstrates practical fairness-aware NLP strategies that can reduce error-rate inequities while preserving accuracy. These findings support equity-focused interventions to address SL through fairness-aware models as assistive auditing tools with human oversight.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Liyang Xue, Mary Chayko, Vivek Kumar Singh
Quelle
Journal of Medical Internet Research
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
1438-8871
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Liyang Xue, Mary Chayko, Vivek Kumar Singh (2026). Stigmatizing Language in Gender-Expansive Patient Records: Corpus Development, Disparity Analysis, and Natural Language Processing–Based Detection Study. Journal of Medical Internet Research. https://doi.org/10.2196/91089
RIS BibTeX CSL-JSON