A collection of 11 pure-inference diagnostic / robustness benchmarks built on top of the
core SpaceNum benchmark. Every subset reuses the SpaceNum MMEval record schema and is
drop-in compatible with the standard evaluation pipeline
(../scripts/run_qwen3vl_spacenum.py, or any MMEval --dataset local@json runner).
All subsets probe off-the-shelf VLMs (no fine-tuning). Training-based studies
(tuning/, reward/… See the full description on the dataset page:
https://huggingface.co/datasets/Sterzhang/spacenum-probing-more.