Stage-3 visual-reasoning data for the staged post-training recipe in
"From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models"
(ICML 2026).
A 16,195-sample mix of visual math and figure-grounded reasoning, drawn from four open-source corpora and packed alongside the source images. Every row also ships with a precomputed pass_rate so the same data can be ordered by sample difficulty for… See the full description on the dataset page:
https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-VisualReasoning-Data.