Views
No views yet
Remix-R1-Distilled-Qwen-1.5B and Remix-R1-Distilled-Qwen-7B, the checkpoints of the ReMix-PPO reported in our paper. We also release a ready-to-use [evaluation scripts] that reproduces all benchmark results.Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model presenting ReMix, a simple yet effective approach that equips on-policy proximal policy gradient methods (eg. PPO and GRPO) with off-policy replay to slash LLM reasoning-finetuning costs while setting new SOTA math performance and, for the first time, exposing how off-policy RL shapes the emergence of reasoning behaviors.| Model | AIME 24 | AMC 23 | MATH 500 | Minerva | OlympiadBench | Avg. | Cost |
|---|---|---|---|---|---|---|---|
| R1-Distill-Qwen-1.5B (Base) | 33.33/18.54 | 43.37/44.58 | 67.40/67.40 | 16.54/17.28 | 27.26/29.19 | 37.58/35.27 | N/A |
| Open-RS1 | 23.33/19.58 | 42.17/45.44 | 64.20/68.00 | 16.18/19.49 | 27.11/29.26 | 34.60/35.91 | 0.058M |
| Open-RS2 | 16.67/20.31 | 45.78/44.92 | 65.00/68.40 | 18.38/17.65 | 26.96/28.15 | 34.56/35.35 | 0.029M |
| Open-RS3 | 16.67/19.69 | 44.58/43.98 | 67.60/67.40 | 15.64/19.12 | 25.48/28.15 | 33.99/35.46 | 0.029M |
| AdaptThink | 13.33/19.06 | 57.83/56.81 | 78.60/76.80 | 23.90/23.90 | 38.07/37.93 | 42.35/42.90 | 0.643M |
| II-Thought | 26.67/28.65 | 56.63/59.72 | 73.00/77.80 | 23.16/23.16 | 40.89/42.37 | 44.07/46.34 | - |
| FASTCuRL-preview | 26.67/25.94 | 60.24/54.27 | 74.20/74.20 | 20.22/21.69 | 32.59/37.04 | 42.78/42.63 | 0.676M |
| FASTCuRL-V3 | 36.67/33.44 | 66.27/63.63 | 84.40/83.40 | 28.67/28.31 | 43.56/44.59 | 51.91/50.67 | 2.478M |
| L1-Exact* | 23.33/25.10 | 71.08/66.57 | 84.00/84.20 | 29.41/26.84 | 44.59/44.3 | 50.48/49.40 | 3.953M |
| L1-Max* | 20.00/23.13 | 69.88/66.79 | 83.00/84.00 | 29.04/27.57 | 46.37/44.44 | 49.66/49.19 | 2.764M |
| DeepScaleR | 40.00/31.96 | 65.06/63.58 | 83.20/81.20 | 29.04/26.1 | 43.41/41.33 | 52.14/48.83 | 2.519M |
| ReMix | 36.67/29.08 | 69.88/64.04 | 82.00/83.80 | 30.15/29.04 | 41.78/42.67 | 52.10/49.73 | 0.079M |
1@article{liang2025squeezesoakedspongeefficient,
2 title={Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model},
3 author={Jing Liang and Hongyao Tang and Yi Ma and Jinyi Liu and Yan Zheng and Shuyue Hu and Lei Bai and Jianye Hao},
4 journal={arXiv preprint arXiv:2507.06892},
5 url={https://arxiv.org/abs/2507.06892},
6 year={2025}
7}