Vollständiger Abstract
Worum geht es in dieser Arbeit?
Background: Burn extent guides triage, transfer and fluid resuscitation, yet its clinical estimation is imprecise and observer-dependent. Multimodal large language models (MLLMs) process clinical photographs without task-specific training, but their error has rarely been separated into systematic and random components or their performance across skin tones characterized. Methods: Three state-of-the-art MLLMs (Gemini 3.1 Pro, GPT-5.6 Sol, and Fable 5) each assessed 153 burn photographs five times under an identical prompt. The tasks were as follows: burned proportion of the imaged field, against an expert-guided pixel-wise segmentation (tolerance ± 10 percentage points, pp); burned percentage of total body surface area (TBSA), against physician consensus (±2 pp); and binary Fitzpatrick skin tone (FST; light I–III versus dark IV–VI). The first of these was the primary endpoint. Results: The primary endpoint was in the range of 32.5–70.2%, with TBSA at 69.5–77.5%. All models compressed the estimation range (slopes 0.58–0.76, intercepts +10.0 to +22.2 pp); one multiplicative constant per model brought errors differing more than twofold into a 1.4 pp range. Across repeated queries, the median within-image range was 5.0–25.0 pp; averaging the five answers reduced error by only 0.24–2.22 pp. FST accuracy was 83.8–91.9% against a majority-class baseline of 81.0%. Conclusions: Averaging repeated answers removes only the smaller, random component; the larger, systematic one persists and requires calibration against reference data before clinical use can be considered.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Ibrahim Güler, Armin Kraus, Gerrit Grieb, Henrik Stelling
- Quelle
- Bioengineering
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2306-5354
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Ibrahim Güler, Armin Kraus, Gerrit Grieb, Henrik Stelling (2026). Burn Extent and Fitzpatrick Skin Tone Assessment from Clinical Photographs: Systematic and Random Error in Multimodal Large Language Models. Bioengineering. https://doi.org/10.3390/bioengineering13091000
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1