π ReVisual-R1 (7B) β Open-Source Multimodal Reasoner
One cold-start, two RL stages, endless reasoning power.
SOTA on 9 tough benchmarks covering visualβmath + text reasoning.
Three-Stage SRO Training
Text Cold-Start β seed deep reflection
Multimodal RL β align vision & logic
Text RL β polish fluency & brevity
PAD (Prioritized Advantage Distillation) keeps gradients alive.
Efficient-Length Reward = concise, self-reflective CoT.
πβ¦ See the full description on the dataset page: https://huggingface.co/datasets/csfufu/mmrl.