Frag' FlorenceEvidenz. Klar. Anwendbar.
Uhr 7/8Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Accuracy and Consistency of Artificial Intelligence Chatbots in Dental Anatomy Education: A Comparative Study

Mubashir Baig Mirza, Ali Robaian, Abdullah Saad Alqahtani, Abdullah A. Alqahtani, Khalid K. Alanazi, Mohammed A. S. Abuelqomsan, Shahad AlBader, Hamod Alqahtani, Abdullah AlShehri

European Journal of Education · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

ABSTRACT The potential of Artificial Intelligence (AI), large language models (LLMs) in enhancing dental education emphasises the need for careful selection of AI tools to improve learning outcomes. Therefore, this study evaluates the accuracy and consistency of responses from eight AI chatbots to multiple‐choice questions (MCQs) related to dental anatomy, both in bulk and individually. A total of 85 questions were adapted from Wheeler's Science of Dental Anatomy. A blinded researcher gathered responses to questions posed in bulk as well as individually from eight chatbots: o3 Mini, GPT‐4o, Deepseek, Grok‐2, ChatGPT‐4, Gemini Flash 2, Claude 3.5 Sonnet, and Meta Llama 3.2, with a week's interval between the two rounds of data collection. Chi‐square tests were conducted to compare the responses from the same chatbots across the two weeks and between the two question formats (bulk vs. individual). Cohen's Kappa was utilised to assess agreement among chatbots, while McNemar's test was employed to evaluate response dissimilarities among chatbots for questions posed through both approaches. ChatGPT‐4 demonstrated the highest accuracy (85.53% in round one and two in bulk; 89.41% and 85.9% in round one and two in individual assessment), while Llama 3.2 performed the poorest (55.29% and 44.7% in bulk; 56.47% and 54.71% individually, in rounds one and two respectively). GPT‐4o and Grok‐2 exhibited the highest consistency ( K = 0.887, K = 0.804), whereas o3‐mini and Meta Llama 3.2 were least consistent. Individual question inputs generally yielded higher accuracy. LLMs, especially ChatGPT‐4 and GPT‐4o, show potential as additional tools for MCQ‐based dental anatomy education. However, variability in performance underscores the necessity for thorough evaluation of these tools. Future research should examine various question formats to improve AI integration in dental curricula.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Mubashir Baig Mirza, Ali Robaian, Abdullah Saad Alqahtani, Abdullah A. Alqahtani, Khalid K. Alanazi, Mohammed A. S. Abuelqomsan, Shahad AlBader, Hamod Alqahtani, Abdullah AlShehri
Quelle
European Journal of Education
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
0141-8211, 1465-3435
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Mubashir Baig Mirza, Ali Robaian, Abdullah Saad Alqahtani, Abdullah A. Alqahtani, Khalid K. Alanazi, Mohammed A. S. Abuelqomsan, Shahad AlBader, Hamod Alqahtani, Abdullah AlShehri (2026). Accuracy and Consistency of Artificial Intelligence Chatbots in Dental Anatomy Education: A Comparative Study. European Journal of Education. https://doi.org/10.1111/ejed.70851
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1 · Lizenz 2