Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
RLVR-vs-SFT-Qwen2.5-1.5b-checkpoints – AI Model by jayminbhan | AlphaNeural AI
You can deploy this model and start earning money today!
jayminbhan
/
RLVR-vs-SFT-Qwen2.5-1.5b-checkpoints
like
0
safetensors
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
RLVR vs SFT on Qwen2.5-1.5b Instruct
GRPO and SFT trained checkpoint of
Qwen2.5-1.5B-Instruct
with GSM8K dataset.
Part of a personal project comparing RLVR vs SFT training methods.
Result
: GSM8K 69.7% → 81.6% while also improving MATH (49.2% → 52.3%).
📊
Full results & code (GitHub)
🔍
Explore prompts and model responses during benchmark! (HF Spaces)