Light-R1: Surpassing R1-Distill from Scratch* with $1000 through Curriculum SFT & DPO
*from models without long COT
technical report
GitHub page
Here is the DPO data we used to train Light-R1-32B.
Simply refer to dpo-pairs.json
Model
Trained From
Release Date
AIME24
AIME25
DeepSeek-R1-Distill-Llama-70B
Llama-3.3-70B-Instruct
25.1.20
70.0
54.1
DeepSeek-R1-Distill-Qwen-32B
Qwen2.5-32B
25.1.20
72.6
54.9
LIMO (32B)
Qwen2.5-32B-Instruct25.2.4
56.3
47.1
s1.1-32B… See the full description on the dataset page:
https://huggingface.co/datasets/qihoo360/Light-R1-DPOData.