Frag' FlorenceEvidenz. Klar. Anwendbar.
Uhr 7/8Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Bridging semantics and clinical fidelity: a section-based assessment of a vision–language model (RadVLM) for chest x-ray report generation

Saleh Alzughaibi

Frontiers in Digital Health · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Background Artificial intelligence-generated radiology reports may reduce clinician burden, but their factual accuracy remains uncertain. This study evaluated whether surface-level semantic similarity in AI-generated Impression sections reliably reflects structural fidelity in the underlying Findings. Methods RadVLM, a radiology-focused vision–language model, generated Findings and Impression sections for 3,000 public chest radiograph studies. A section-aware evaluation framework assessed Findings-level entity-relation fidelity using RadGraph F1 and Impression-level quality using BERTScore-F1 (semantic similarity), CheXbert cosine similarity (label-space agreement), and readability metrics. Generalized linear models examined associations between Impression-level metrics and Findings-level fidelity. Results Mean RadGraph F1 was 0.251 (95% CI, 0.245–0.257), indicating substantially imperfect structural correctness. Mean BERTScore-F1 was 0.498 and CheXbert cosine similarity was 0.394. AI-generated Impressions were more readable (mean Flesch–Kincaid 11.72 vs. 15.33 for reference). A weak positive association existed between semantic similarity and entity-relation fidelity (coefficient 0.154; 95% CI, 0.124–0.183), and this association attenuated substantially with longer Impressions. In practical terms, the association was too weak to be useful clinically: a large gain in semantic similarity corresponded to only a small change in structural fidelity, so a fluent, readable Impression offered little assurance that the underlying Findings were factually correct. Conclusions In this evaluation of a single model on a single public dataset, surface-level fluency and readability did not reliably indicate factual correctness. These findings apply primarily to RadVLM in the experimental setting studied and should not be generalized to AI report generation as a whole without multi-model, multi-dataset replication. Even so, they reinforce a broader principle: clinical deployment of such systems requires explicit entity-relation validation, structured accuracy assessment, and radiologist oversight to protect patient safety.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Saleh Alzughaibi
Quelle
Frontiers in Digital Health
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2673-253X
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Saleh Alzughaibi (2026). Bridging semantics and clinical fidelity: a section-based assessment of a vision–language model (RadVLM) for chest x-ray report generation. Frontiers in Digital Health. https://doi.org/10.3389/fdgth.2026.1882716
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1