An image-grounded measurement VQA benchmark for 3D CT, released with
CheapCT (code).
Every gold answer is computed directly from a segmentation mask (TotalSegmentator) and the Hounsfield units
inside it — never from a radiology report or a language model — so every label is exactly reproducible from
the image. Disease is the one exception: its labels come from the reports. The benchmark scores an encoder one
capability at a time.
Capabilities… See the full description on the dataset page: https://huggingface.co/datasets/LiangRenjie/MeasureVQA.