deepseek-r1-qwen-7b generations for deepscaler dataset
The original deepscaler dataset has been filtered:
we removed all synthetic data because their problem-answer may not match.
based on generations from Qwen/Qwen2.5-Math-7B-Instruct (pre-o1), we removed problems that has at least 5/32 correct generations.
We then use deepseek-ai/DeepSeek-R1-Distill-Qwen-7B to generate from this filtered dataset with num_generations=32 and max_tokens=8192