Vollständiger Abstract
Worum geht es in dieser Arbeit?
Abstract Purpose Foundation models are increasingly used as frozen feature extractors for CT classification tasks, yet the determinants of their downstream performance remain unclear. We assess how foundation model choice, feature aggregation strategy, and abnormality type are associated with performance in organ-level abnormality classification. We further evaluate whether more expressive aggregation strategies outperform simple pooling, and whether performance varies by abnormality type. Methods We evaluate five state-of-the-art medical imaging foundation models on binary organ-level abnormality classification (normal vs. abnormal) across six abdominal organs. Local patch-level embeddings are aggregated using six strategies, including simple pooling methods (e.g., mean pooling) and attention-based multiple instance learning. A linear classifier is trained on feature embeddings, and performance is evaluated on an annotated test set of 200 CT scans. Results Among the evaluated models, 3D CT-native models generally outperformed 2D multi-modal models. None of the tested aggregation strategies significantly outperformed mean pooling, including attention-based multiple instance learning (best-performing aggregation vs. mean: $$\Delta $$ Δ AUC = 0.008, 95% CI $$[-0.002, 0.022]$$ [ - 0.002 , 0.022 ] ). In our organ-level classification setting, performance differed between abnormality type for all well-performing models (SPECTRE, TAP-CT, CT-FM), with lower AUCs observed for focal abnormalities compared to diffuse abnormalities (largest difference: $$\Delta $$ Δ AUC = 0.108, 95% CI [0.080, 0.135]). This gap was not reduced by the evaluated aggregation strategies. Conclusion In our experiments, downstream performance varied more across foundation models than across aggregation strategies. The choice of aggregation strategy, including attention-based multiple instance learning, did not significantly impact performance in this setting. The persistent gap for focal abnormalities suggests that current representations may insufficiently encode localized disease patterns, motivating the development of localization-aware pre-training approaches.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Jan Tagscherer, Sarah de Boer, Fennie van der Graaf, Lena Philipp, Colin Jacobs, Ewoud J. Smit, Alessa Hering
- Quelle
- International Journal of Computer Assisted Radiology and Surgery
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 1861-6429
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Jan Tagscherer, Sarah de Boer, Fennie van der Graaf, Lena Philipp, Colin Jacobs, Ewoud J. Smit, Alessa Hering (2026). Systematic evaluation of foundation models for organ-level classification on CT scans. International Journal of Computer Assisted Radiology and Surgery. https://doi.org/10.1007/s11548-026-03786-x
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1