Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen3-4b-grpo-dapo17k-invmax-linear – AI Model by hkr04 | AlphaNeural AI
You can deploy this model and start earning money today!
hkr04
/
qwen3-4b-grpo-dapo17k-invmax-linear
like
0
safetensors
qwen3
BytedTsinghua-SIA/DAPO-Math-17k
Qwen/Qwen3-4B
finetune
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Trained on DAPO-Math-17k using GRPO implemented by VeRL.
Batch Size: 32
Group Size: 8
Step: 500
Max Response Length: 8192