[π Website] β’
[π€ Dataset] β’
[π Paper] β’
[π± GitHub] β’
[π¦ Twitter] β’
[π Rednote]
This dataset consists of 314k variational problems synthesized by the Qwen2.5-32B-Instruct policy during RLVR training on DAPO-17k using the SvS strategy for 600-step training, each accompanied by reference answers.The variational problems undergo a min_hash deduplication with a threshold of 0.85.
from datasets import⦠See the full description on the dataset page:
https://huggingface.co/datasets/RLVR-SvS/Variational-DAPO.