PR1 is a new experimental dataset to get reasoning models to plan before thinking before responding. All responses in the default subset are verified.
It may be beneficial to continue training with GRPO in order to further improve performance.
Statistic
Value
Total Tokens
13,254,651
Avg Tokens (Response)
1,211
Total Questions
36,077
Total Correct
10,942
Correct (%)
30.33%
Model
Llama-3.3-70B-Instruct