Vollständiger Abstract
Worum geht es in dieser Arbeit?
Attention toward Chat Generative Pre-trained Transformer (ChatGPT) has greatly increased, including in the medical field, but concerns about its problem-solving ability, accuracy, reliability, and limitations in medical problems always coexist. This study aimed to evaluate ChatGPT’s performance in solving a medical school examination. We used 50-question examinations from the preventive medicine and psychiatry fields applied to 41 fourth-year medical students. Forty-nine questions were separately entered (one was excluded for being image-based) into the text box of ChatGPT models 3.5 and 4.0 with a request to select the single best answer. The initial answer was considered for primary evaluation, but incorrect answers led to model re-prompting to evaluate its answer regeneration function. We assessed the accuracy of both ChatGPT models and compared them with that of students. The student mean score was 29.5 ± 5.9. ChatGPT’s performance was higher in model 4.0 and when re-prompted (models 3.5 and 4.0 initial scores, 28 and 43; re-prompt scores, 40 and 46). Model 4.0 outperformed all students at re-prompting. Both ChatGPT models and all prompts (except for the model 3.5 first prompt) had a higher probability than students of providing correct answers. For easy questions, the accuracy of both ChatGPT models differed significantly from that of students. For difficult questions, the difference between students and the first prompt of model 3.5 was nonsignificant. We observed a nonsignificant tendency in the point estimation when comparing the accuracy of ChatGPT models and prompts by question difficulty. Both models showed excellent performance and learning ability in the medical field, especially model 4.0, and can be considered reliable large language model tools. However, future studies should constantly verify the reliability of the information from these models. Also, educators must be cautious when interpreting/applying ChatGPT output data for clinical and educational purposes.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Kyung Wook Kim, Hong Bae Jeon, Jong Hyuk Choi, Dong Uk Lee, Yoon Sung Jeon, Bum-Jin Shim
- Quelle
- Medicine
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 0025-7974, 1536-5964
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Kyung Wook Kim, Hong Bae Jeon, Jong Hyuk Choi, Dong Uk Lee, Yoon Sung Jeon, Bum-Jin Shim (2026). Artificial intelligence in medical education. Medicine. https://doi.org/10.1097/md.0000000000050323
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1