Vollständiger Abstract
Worum geht es in dieser Arbeit?
Introduction Large language models (LLMs) are increasingly used as sources of information across many fields, including healthcare. As patients turn to these models for health-related queries, evaluating the accuracy and reliability of their responses is essential. Peri-acetabular osteotomy (PAO) is performed on younger patients – a group more likely to use digital tools like LLMs for health information. This study assesses the accuracy and readability performance of two leading LLMs, ChatGPT and Google Gemini, in addressing common patient questions on PAO. Methods A panel of fellowship-trained PAO surgeons created ten commonly asked patient questions based on real-world experience. Responses from each LLM were assessed by the same three surgeons, blinded to response origin, using a 5-point Likert scale to evaluate clarity, accuracy, and completeness. Readability was measured with Flesch–Kincaid Reading scores. Results ChatGPT outperformed Gemini with an average score of 4.17 vs 3.13 (t = −3.08, p = 0.006). ChatGPT’s responses were often rated higher for completeness and clarity, particularly in areas needing detailed explanation, and usually required minimal clarification. Gemini sometimes lacked specificity or included minor inaccuracies that reduced its perceived reliability. Both LLMs produced responses with similar “difficult” Flesch Reading Ease scores. Conclusions There may be significant differences in how effectively LLMs support patients with surgical queries. ChatGPT more consistently met expert standards for clarity and thoroughness. As LLM usage expands, ChatGPT may aid patient education on hip surgery, supporting consultations, informed decisions and postoperative guidance.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- TP Davis, B Guevel, K Logishetty, AG Dick, J Hutt
- Quelle
- The Annals of The Royal College of Surgeons of England
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 0035-8843, 1478-7083
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
TP Davis, B Guevel, K Logishetty, AG Dick, J Hutt (2026). Evaluating large language models in patient education: a comparative analysis addressing frequently asked questions in peri-acetabular osteotomy. The Annals of The Royal College of Surgeons of England. https://doi.org/10.1308/rcsann.2026.0056