vero50k: 45,434 rows · vero600k: 49,654 rows · 54–58 sources · images embedded
Vero is a mixed-modality RL training set covering a broad range of visual reasoning tasks: math, charts, grounding, web navigation, spatial reasoning, VQA, and more. It uses a composite DAPO-style reward (0.7 · accuracy + 0.3 · format) and supports 11 distinct reward types dispatched per-row via extra_info.reward_type.
Two _final versions are provided:… See the full description on the dataset page:
https://huggingface.co/datasets/ngqtrung/vero-rl.