Filtered subset of RuoliuYang/ULVR_v2_clean: the 101,951 training samples that Qwen2.5-VL-7B-Instruct answered incorrectly given only input_image, but correctly once the intermediate_image_* were also provided (judged by Qwen3-VL-32B-Instruct). Same schema / subsets / train-split structure as the source.