Views
No views yet
Qwen/Qwen2-0.5B-Instruct fine-tuned with a from-scratch GRPO (Group Relative Policy Optimization) implementation — no TRL/RLHF library, just a custom training loop — on the GSM8K math reasoning dataset. Produced as part of a take-home engineering exercise; not a production model.beta=0.04) against a frozen reference copy of the same base model.5e-6, linear warmup over the first 18% of steps, AdamW, gradient clipped to norm 0.1.<reasoning>
...step-by-step reasoning...
</reasoning>
<answer>
INTEGER
</answer>