Views
No views yet
| Model | Type | Size | Capabilities | Use Cases | Download |
|---|---|---|---|---|---|
| ReasonFlux-PRM | PRM | 7B | • Trajectory-aware scoring • Online/Offline supervision • Dense process rewards | Data selection, RL training, Test-time scaling | 🤗 7B |
| ReasonFlux-PRM | PRM | 1.5B | • Lightweight scoring • Efficient inference • Edge deployment | Resource-constrained applications | 🤗 1.5B |
| ReasonFlux-PRM-Qwen-2.5 | End-to-End Trained Policy Model | 7B | • Long CoT reasoning • Solving complex tasks and problems | Math and Science Reasoning | 🤗 7B |
Note: We obtain ReasonFlux-PRM-Qwen-2.5-7B through an end-to-end training process, first applying SFT on 1k Trajectory–Response pairs selected by ReasonFlux-PRM-7B, followed by RL training with ReasonFlux-PRM-7B integrated GRPO.
1@article{zou2025reasonfluxprm,
2 title={ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs},
3 author={Zou, Jiaru and Yang, Ling and Gu, Jingwen and Qiu, Jiahao and Shen, Ke and He, Jingrui and Wang, Mengdi},
4 journal={arXiv preprint arXiv:2506.18896},
5 year={2025}
6}