Frag' FlorenceEvidenz. Klar. Anwendbar.
Uhr 7/8Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Artificial intelligence in medical education

Kyung Wook Kim, Hong Bae Jeon, Jong Hyuk Choi, Dong Uk Lee, Yoon Sung Jeon, Bum-Jin Shim

Medicine · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Attention toward Chat Generative Pre-trained Transformer (ChatGPT) has greatly increased, including in the medical field, but concerns about its problem-solving ability, accuracy, reliability, and limitations in medical problems always coexist. This study aimed to evaluate ChatGPT’s performance in solving a medical school examination. We used 50-question examinations from the preventive medicine and psychiatry fields applied to 41 fourth-year medical students. Forty-nine questions were separately entered (one was excluded for being image-based) into the text box of ChatGPT models 3.5 and 4.0 with a request to select the single best answer. The initial answer was considered for primary evaluation, but incorrect answers led to model re-prompting to evaluate its answer regeneration function. We assessed the accuracy of both ChatGPT models and compared them with that of students. The student mean score was 29.5 ± 5.9. ChatGPT’s performance was higher in model 4.0 and when re-prompted (models 3.5 and 4.0 initial scores, 28 and 43; re-prompt scores, 40 and 46). Model 4.0 outperformed all students at re-prompting. Both ChatGPT models and all prompts (except for the model 3.5 first prompt) had a higher probability than students of providing correct answers. For easy questions, the accuracy of both ChatGPT models differed significantly from that of students. For difficult questions, the difference between students and the first prompt of model 3.5 was nonsignificant. We observed a nonsignificant tendency in the point estimation when comparing the accuracy of ChatGPT models and prompts by question difficulty. Both models showed excellent performance and learning ability in the medical field, especially model 4.0, and can be considered reliable large language model tools. However, future studies should constantly verify the reliability of the information from these models. Also, educators must be cautious when interpreting/applying ChatGPT output data for clinical and educational purposes.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Kyung Wook Kim, Hong Bae Jeon, Jong Hyuk Choi, Dong Uk Lee, Yoon Sung Jeon, Bum-Jin Shim
Quelle
Medicine
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
0025-7974, 1536-5964
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Kyung Wook Kim, Hong Bae Jeon, Jong Hyuk Choi, Dong Uk Lee, Yoon Sung Jeon, Bum-Jin Shim (2026). Artificial intelligence in medical education. Medicine. https://doi.org/10.1097/md.0000000000050323
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1