Frag' FlorenceEvidenz. Klar. Anwendbar.
Uhr 7/8Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Evaluating AI-generated patient education materials for endometrial cancer surgery: a comparative analysis of response quality, reliability, and readability between ChatGPT and DeepSeek models

Ling Tian, Long-yu Tang, Ming-tao Yang, Yan Wang, Hong-ni He, Tian-wen He, Ying Tang, Chuan Lin, Hui-quan Hu, Jun Li

Frontiers in Public Health · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Purpose This study aimed to evaluate and compare the quality, reliability, and readability of patient education materials on endometrial cancer surgery generated by ChatGPT (GPT-5) and DeepSeek (R1). Materials and methods This cross-sectional study analyzed the responses generated by ChatGPT and DeepSeek to totally 41 questions covering four domains: surgical planning, preoperative evaluation, postoperative care, and long-term follow-up. Reliability was assessed through the DISCERN and EQIP instruments, quality was evaluated by the Global Quality Score (GQS), and readability was analyzed by the Flesch Reading Ease Score (FRES), Gunning Fog Index (GFI), and Flesch-Kincaid Grade Level (FKGL). Statistical comparisons were performed by using paired t -tests and Wilcoxon signed-rank tests. Results The two large language models (LLMs) generated education materials of comparable quality, as reflected in GQS scores (median: DeepSeek vs. ChatGPT 5.00 vs. 4.67, p = 0.077). DeepSeek demonstrated statistically significantly higher reliability scores on both DISCERN and EQIP instruments (both p < 0.001). Readability scores (FRES, GFI) were similar between groups, while DeepSeek exhibited a higher FKGL (10.28 vs. 8.84, p < 0.001), indicating the greater text complexity. Subgroup analysis showed that DeepSeek performed better in terms of reliability in the postoperative care and long-term follow-up domains, while ChatGPT exhibited better readability in the surgical planning domain. Conclusion Both DeepSeek and ChatGPT can generate patient education text drafts that are commendable in their structural coherence and linguistic clarity. DeepSeek demonstrates a significant advantage in information reliability, particularly excelling in postoperative and follow-up management content. ChatGPT shows a slight edge in the readability of surgical planning sections. However, the text readability of both models exceeds the general public's health literacy level. This indicates that large language models can only serve as auxiliary tools for generating patient education materials. Their outputs must undergo review by clinical experts and readability optimization to ensure both accuracy and comprehensibility of the information.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Ling Tian, Long-yu Tang, Ming-tao Yang, Yan Wang, Hong-ni He, Tian-wen He, Ying Tang, Chuan Lin, Hui-quan Hu, Jun Li
Quelle
Frontiers in Public Health
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2296-2565
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Ling Tian, Long-yu Tang, Ming-tao Yang, Yan Wang, Hong-ni He, Tian-wen He, Ying Tang, Chuan Lin, Hui-quan Hu, Jun Li (2026). Evaluating AI-generated patient education materials for endometrial cancer surgery: a comparative analysis of response quality, reliability, and readability between ChatGPT and DeepSeek models. Frontiers in Public Health. https://doi.org/10.3389/fpubh.2026.1831988
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1