Vollständiger Abstract
Worum geht es in dieser Arbeit?
ABSTRACT Objective To compare the ability of large language models (LLMs) and otolaryngologists to identify prognostic signals from brief clinical vignettes in CRSwNP. Methods In this blinded study, 68 adults initiating biologic therapy for CRSwNP (≥ 36 months follow‐up) were represented by standardized vignettes derived from documentation immediately before biologic initiation. Vignettes summarized symptoms, prior surgery, comorbidities, and medications. Biomarkers, imaging scores, smell testing, and follow‐up data were excluded to isolate text‐based prognostic reasoning under identical constraints. Five attending otolaryngologists, one rhinology fellow, and two residents independently predicted four 5‐year outcomes: subsequent endoscopic sinus surgery, recurrent systemic steroid bursts, biologic switch, and a composite endpoint. Multiple LLMs were queried in identical zero‐shot format across three sessions to assess stability. Predictions were compared with verified outcomes using accuracy, sensitivity, specificity, F1 score, Cohen's κ, and Matthews correlation coefficient. Results Five‐year outcome prevalences were 38.2% (26 of 68) for subsequent sinus surgery, 26.5% (18 of 68) for steroid bursts, 14.7% (10 of 68) for biologic switch, and 58.8% (40 of 68) for composite failure. Macro‐averaged LLM accuracies ranged from 65.0% to 76.1%, compared with 65.7% ± 7.4% for human raters. The best‐performing LLM achieved accuracy comparable to the top attending. Both groups showed higher negative than positive predictive values, indicating better discrimination for patients without adverse events. Agreement between LLMs and clinicians was moderate and within the range of inter‐clinician variability. Conclusion When restricted to short pretreatment text, LLMs produced prognostic judgments within the range of clinician variability for severe, biologic‐treated CRSwNP outcomes. These findings suggest that in information‐limited settings, LLMs may provide an auxiliary reasoning signal rather than a standalone decision tool.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Sholem Hack, Chase Kahn, Ainhoa Garcia‐Lliberos, Cristina Rodriguez‐Prado, Ameen Biadsee, Maxime Fieux, Mark Liu, Eran Glikson, Masayoshi Takashima, Omar G. Ahmed, Raj Sindwani, Dennis M. Tang, Christopher R. Roxbury
- Quelle
- World Journal of Otorhinolaryngology - Head and Neck Surgery
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2095-8811, 2589-1081
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Sholem Hack, Chase Kahn, Ainhoa Garcia‐Lliberos, Cristina Rodriguez‐Prado, Ameen Biadsee, Maxime Fieux, Mark Liu, Eran Glikson, Masayoshi Takashima, Omar G. Ahmed, Raj Sindwani, Dennis M. Tang, Christopher R. Roxbury (2026). Generative AI and Clinicians Show Comparable Prognostic Reasoning From Clinical Narratives in Biologic‐Treated CRSwNP. World Journal of Otorhinolaryngology - Head and Neck Surgery. https://doi.org/10.1002/wjo2.70164
Kontext