Vollständiger Abstract
Worum geht es in dieser Arbeit?
Background Large language models (LLMs) are increasingly used to support digital health communication, yet their reliability in patient-facing cardiovascular imaging education remains uncertain. Cardiovascular imaging involves complex terminology and procedural details that many patients struggle to understand, creating a need for accurate, clear, and reassuring explanations. While prior evaluations of conversational AI have focused primarily on diagnostic reasoning or clinician-oriented tasks, few studies have systematically compared contemporary LLMs in their ability to communicate effectively with patients. Objective This study aimed to compare the accuracy, clarity, completeness, and patient-centered communication quality of responses generated by 3 state-of-the-art conversational agents (DeepSeek, GPT-o1, and GPT-4o) when addressing real-world patient questions about cardiovascular imaging. Methods A prospective methodological evaluation was conducted using 84 unique patient-centered questions curated from authoritative cardiovascular information sources and online patient forums. Each question was independently submitted to DeepSeek, GPT-o1, and GPT-4o in isolated sessions to avoid contextual contamination. Two cardiovascular radiologists scored each response across 4 domains (accuracy, clarity and appropriateness, completeness, and user engagement and reassurance) using a standardized 3-point rubric (total score range 4-12). Discrepancies were resolved through predefined adjudication procedures. Because the scores were ordinal, median domain and composite scores with IQRs were summarized and compared across the 3 models using the Kruskal-Wallis test, with ε2 as an effect size. Statistical significance was defined as an α value of .05. Results Across the 84 patient questions, all 3 models produced largely accurate, clear, and complete responses, with comparably high scores across the accuracy, clarity, and completeness domains (median 3 of 3, IQR 3-3 in each). The only meaningful difference appeared in user engagement and reassurance. A “good” engagement rating was assigned to 96.4% (81/84) of DeepSeek responses and 98.8% (83/84) of GPT-o1 responses but only 53.6% (45/84) of GPT-4o responses (Kruskal-Wallis P
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Ahmed Marey, Basudha Pal, Ayşenur Buz Yaşar, Shree Rath, Giulia Francese, Hossam M Ghorab, Julia Niemierko, Muhammad Shah Wali Jamal, Muhammad Umair
- Quelle
- JMIR Formative Research
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2561-326X
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Ahmed Marey, Basudha Pal, Ayşenur Buz Yaşar, Shree Rath, Giulia Francese, Hossam M Ghorab, Julia Niemierko, Muhammad Shah Wali Jamal, Muhammad Umair (2026). Large Language Models for Patient Education in Cardiovascular Imaging: Prospective Observational Comparative Study. JMIR Formative Research. https://doi.org/10.2196/95883