MMR1-SFT is a large-scale, carefully curated vision–language long chain-of-thought (CoT) dataset for cold-start supervised fine-tuning of multimodal reasoning models.It accompanies our work on Variance-Aware Sampling (VAS) for RL post-training and the MMR1 model family.
Scale: ~1.6M multimodal QA examples with verified long CoT rationales and short answers
Quality control: CoTs generated by Gemini-2.5 Pro/Flash, verified by… See the full description on the dataset page:
https://huggingface.co/datasets/MMR1/MMR1-SFT.