Vollständiger Abstract
Worum geht es in dieser Arbeit?
Abstract We evaluated proprietary and open-weight foundation models on 24 German medical licensing examinations (2019–2024), including 7485 items and response data from 119,878 sittings. For fair comparison, the eight vision-capable models were evaluated on the full benchmark and all thirteen on a shared text-only subset. On the full benchmark, Gemini 3.1 Pro achieved the highest overall accuracy, reaching 99.31% on the first (M1) and 98.37% on the second (M2) examination. On the shared text-only subset, proprietary frontier models performed at near-ceiling levels, several open-weight models (including GLM-5 and DeepSeek V3.2-Thinking) were highly competitive, and even compact ones exceeded mean student performance. Image-present items were more difficult for both students and models, but the associated decline was disproportionately larger for models than for students. Human- and model-defined difficulty subsets showed limited overlap, and model-hard subsets revealed residual differences among top systems. These findings underscore the rapid progress of these models, particularly open-weight systems, and the value of official German medical licensing examinations as a restricted-access benchmark with reduced public exposure. They carry implications for high-stakes assessment and AI-assisted medical education, notably multimodal assessment, human-aligned educational tools, and privacy-preserving local deployment.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Lasse Cirkel, Johannes Knitza, Volker Schillings, Alexander Oksche, Jan Carl Becker, Sebastian Kuhn
- Quelle
- npj Digital Medicine
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2398-6352
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Lasse Cirkel, Johannes Knitza, Volker Schillings, Alexander Oksche, Jan Carl Becker, Sebastian Kuhn (2026). Evaluating foundation models on official German medical licensing examinations: Implications for high-stakes assessment and AI-assisted medical education. npj Digital Medicine. https://doi.org/10.1038/s41746-026-03082-7
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1