Vollständiger Abstract
Worum geht es in dieser Arbeit?
<h4>Purpose</h4>The purpose of this study was to evaluate and compare four AI software programs-ChatGPT-5, Microsoft Copilot, Google Gemini (V 2.5), and OpenEvidence-in generating comprehensive dental treatment plans for minimally destructed and severely mutilated teeth using identical clinical inputs.<h4>Material and methods</h4>Ten anonymized clinical cases, each consisting of 1 intraoral photograph and 1 corresponding periapical radiograph, were independently submitted to each AI software program using a standardized prompt. A reference standard was established through consensus among four calibrated specialists (one surgically trained prosthodontist, one restorative dentist, one prosthodontist, and one endodontist). AI-generated responses were evaluated using a structured scoring rubric across five domains: diagnostic accuracy, restorability assessment, multidisciplinary integration, treatment sequencing, and extraction appropriateness (score range: 0-10 per case). Agreement among evaluators was assessed using pairwise Cohen's kappa (κ) analysis based on initial independent scoring before consensus discussions.<h4>Results</h4>Substantial inter-evaluator agreement was observed (κ = 0.74; 95% CI: 0.68-0.80), indicating consistent application of the scoring rubric. Mean total scores were highest for OpenEvidence (7.7 ± 1.4) and ChatGPT-5 (7.6 ± 1.1), followed by Google Gemini (5.1 ± 1.2) and Microsoft Copilot (4.7 ± 3.1). OpenEvidence and ChatGPT-5 more frequently incorporated phased treatment sequencing, ferrule assessment, restorability analysis, and multidisciplinary treatment considerations. Microsoft Copilot failed to generate responses in three of the 10 evaluated cases because of content-filtering restrictions.<h4>Conclusions</h4>AI software programs generated structured and clinically relevant dental treatment proposals; however, meaningful variability existed in diagnostic interpretation, restorability assessment, multidisciplinary integration, and extraction thresholds. OpenEvidence and ChatGPT-5 demonstrated greater agreement with the expert reference standard, although clinically significant diagnostic errors were observed across all evaluated systems. AI-generated treatment plans should therefore be regarded as adjunctive decision-support tools requiring specialist oversight before clinical implementation.
Abstract: PubMed · Datensatz
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Nicht angegeben
- Quelle
- CrossRef Listing of Deleted DOIs
- Publikation
- 2000-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 0849-6757
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
(2000). 10.18808/jopr.2015.1.11. CrossRef Listing of Deleted DOIs. https://doi.org/10.1111/jopr.70226