We pretrain a 3D CT vision–language model on 159k report–volume pairs with two new supervision signals:
prompt-based disease labels for classification and intra-scan snippet localization for axial depth grounding.
A single unified model reaches state-of-the-art retrieval on CT-RATE, competitive disease
classification, and slice-level localization at 12 mm resolution.
1@inproceedings{ging2026radfinder,
2 author = {Simon Ging and Philipp Arnold and Sebastian Walter and Hani Alnahas and Hannah Bast and Elmar Kotter and Jiancheng Yang and Behzad Bozorgtabar and Thomas Brox},
3 title = {Learning to Read Where to Look: Disease-Aware Vision--Language Pretraining for 3{D} {CT}},
4 booktitle = {Medical Image Computing and Computer Assisted Intervention -- {MICCAI} 2026, Strasbourg, France, September 27 -- October 1, 2026, Proceedings},
5 series = {Lecture Notes in Computer Science},
6 publisher = {Springer},
7 year = {2026},
8 note = {To appear},
9}