Views
No views yet

| Method | Memory Usage of Model Parameters |
|---|---|
| ROVER (Ours) | Low (actor model ONLY!😊) |
| GRPO | Medium (actor + reference model) |
| PPO | High (actor + reference + critic model) |
1@article{he2025randompolicyvaluation,
2 title={Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards},
3 author={Haoran He and Yuxiao Ye and Qingpeng Cai and Chen Hu and Binxing Jiao and Daxin Jiang and Ling Pan},
4 journal={arXiv preprint arXiv:2509.24981},
5 year={2025}
6}