Frag' FlorenceEvidenz. Klar. Anwendbar.
Uhr 7/8Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Systematic evaluation of foundation models for organ-level classification on CT scans

Jan Tagscherer, Sarah de Boer, Fennie van der Graaf, Lena Philipp, Colin Jacobs, Ewoud J. Smit, Alessa Hering

International Journal of Computer Assisted Radiology and Surgery · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Abstract Purpose Foundation models are increasingly used as frozen feature extractors for CT classification tasks, yet the determinants of their downstream performance remain unclear. We assess how foundation model choice, feature aggregation strategy, and abnormality type are associated with performance in organ-level abnormality classification. We further evaluate whether more expressive aggregation strategies outperform simple pooling, and whether performance varies by abnormality type. Methods We evaluate five state-of-the-art medical imaging foundation models on binary organ-level abnormality classification (normal vs. abnormal) across six abdominal organs. Local patch-level embeddings are aggregated using six strategies, including simple pooling methods (e.g., mean pooling) and attention-based multiple instance learning. A linear classifier is trained on feature embeddings, and performance is evaluated on an annotated test set of 200 CT scans. Results Among the evaluated models, 3D CT-native models generally outperformed 2D multi-modal models. None of the tested aggregation strategies significantly outperformed mean pooling, including attention-based multiple instance learning (best-performing aggregation vs. mean: $$\Delta $$ Δ AUC = 0.008, 95% CI $$[-0.002, 0.022]$$ [ - 0.002 , 0.022 ] ). In our organ-level classification setting, performance differed between abnormality type for all well-performing models (SPECTRE, TAP-CT, CT-FM), with lower AUCs observed for focal abnormalities compared to diffuse abnormalities (largest difference: $$\Delta $$ Δ AUC = 0.108, 95% CI [0.080, 0.135]). This gap was not reduced by the evaluated aggregation strategies. Conclusion In our experiments, downstream performance varied more across foundation models than across aggregation strategies. The choice of aggregation strategy, including attention-based multiple instance learning, did not significantly impact performance in this setting. The persistent gap for focal abnormalities suggests that current representations may insufficiently encode localized disease patterns, motivating the development of localization-aware pre-training approaches.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Jan Tagscherer, Sarah de Boer, Fennie van der Graaf, Lena Philipp, Colin Jacobs, Ewoud J. Smit, Alessa Hering
Quelle
International Journal of Computer Assisted Radiology and Surgery
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
1861-6429
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Jan Tagscherer, Sarah de Boer, Fennie van der Graaf, Lena Philipp, Colin Jacobs, Ewoud J. Smit, Alessa Hering (2026). Systematic evaluation of foundation models for organ-level classification on CT scans. International Journal of Computer Assisted Radiology and Surgery. https://doi.org/10.1007/s11548-026-03786-x
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1