Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Qwen-2.5-Math-1.5B-DPO – AI Model by Liang0223 | AlphaNeural AI
You can deploy this model and start earning money today!
Liang0223
/
Qwen-2.5-Math-1.5B-DPO
like
0
safetensors
qwen2
mit
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
This model was presented in the paper
On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
.
Code:
https://github.com/yongliang-wu/DFT?tab=readme-ov-file