Vollständiger Abstract
Worum geht es in dieser Arbeit?
Abstract Introduction Artificial intelligence (AI) tools are increasingly used by patients and caregivers seeking medical information. Developmental dysplasia of the hip (DDH) is a common pediatric condition that frequently prompts parental questions regarding screening, diagnosis, and treatment. The purpose of this study was to evaluate the accuracy, completeness, and clarity of responses generated by a Department of Defense–approved AI chatbot (NIPR-GPT) when answering commonly asked DDH-related questions. Methods Twelve frequently asked patient questions related to DDH were submitted to the NIPR-GPT chatbot. Each question was entered into a new chatbot session to minimize contextual bias. AI-generated responses were independently evaluated by 5 fellowship-trained pediatric orthopedic surgeons. Responses were graded using a 4-point scale assessing accuracy, completeness, and clarity: (1) unsatisfactory (major inaccuracies requiring substantial correction), (2) satisfactory with moderate clarification required, (3) satisfactory with minimal clarification required, and (4) excellent with no clarification required. Inter-rater reliability was assessed using intraclass correlation coefficients (ICC) and weighted kappa statistics, while overall internal consistency among raters was assessed using Krippendorff’s alpha. Results All AI-generated responses (100%) were rated either satisfactory or excellent. Six responses (50%) were graded excellent, requiring no clarification, and 6 responses (50%) were graded satisfactory with minimal clarification required. No responses were graded as unsatisfactory or requiring major correction. Single-rater reliability among individual reviewers was poor (ICC [2, 1] = 0.12), reflecting variability in reviewer thresholds. However, reliability improved when scores were averaged across raters, demonstrating moderate agreement (ICC [2, k] = 0.41). Internal consistency among reviewers was Krippendorff’s α = 0.163, indicating heterogeneity in reviewer grading but not systematic deficiencies in AI responses. Five fellowship-trained pediatric orthopedic surgeons independently graded all 12 responses. No reviewer graded any response as unsatisfactory or requiring substantial correction. Mean question scores ranged from 2.8 to 4.0 across the 12 questions, with an overall mean score of approximately 3.6. Conclusion NIPR-GPT generated generally accurate, clear, and clinically appropriate responses to common DDH-related questions. Half of responses required no clarification, and none required major correction. These findings suggest that AI chatbots may serve as useful adjunct educational tools for patients and caregivers within military healthcare systems. However, variability in expert interpretation and the potential for patient misinterpretation underscore the continued importance of clinician oversight when integrating AI-generated information into patient education.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Michael G Johnston, William Ferris, Harrison Diaz, Andrew Rosso, Mark S Katsma
- Quelle
- Military Medicine
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 0026-4075, 1930-613X
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Michael G Johnston, William Ferris, Harrison Diaz, Andrew Rosso, Mark S Katsma (2026). Artificial Intelligence Answering Patient Questions About Developmental Dysplasia of the Hip: Accuracy, Readability, and Clinical Utility. Military Medicine. https://doi.org/10.1093/milmed/usag394
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1