Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
distill-1.5b-grpo-minmax – AI Model by hkr04 | AlphaNeural AI
You can deploy this model and start earning money today!
hkr04
/
distill-1.5b-grpo-minmax
like
0
safetensors
qwen2
BytedTsinghua-SIA/DAPO-Math-17k
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
finetune
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Trained on DAPO-Math-17k using GRPO implemented by VeRL.
Batch Size: 32
Group Size: 8
Step: 280
Max Response Length: 8192