Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Haizhong Zheng1, Yang Zhou1, Brian R. Bartoldson2, Bhavya Kailkhura2,
Fan Lai3, Jiawei Zhao4, Beidi Chen1 1Carnegie Mellon University,
2Lawrence Livermore National Laboratory,
3University of Illinois Urbana-Champaign,
4Meta AI
TL;DR
We propose GRESO, a lightweight pre-rollout filtering method that improves the efficiency of rollout scaling in LLM RL by predicting and skipping low-value prompts.
Figure 1: We train Qwen2.5-Math-1.5B/7B on the DAPO + MATH dataset and evaluate them on five math reasoning benchmarks: MATH500, AMC, Gaokao, Minerva, and Olympiad Bench. Compared to the baseline method (Dynamic Sampling), our approach (GRESO) reduces rollout overhead by up to 2x while achieving comparable training performance, improving the efficiency of rollout scaling.