Views
No views yet
| Metric | Value |
|---|---|
| GSM8K Eval Accuracy | 49.70% |
| W&B Run | https://wandb.ai/tripathysagar08/grpo-gsm8k-full/runs/jcfqwy7u |
| Parameter | Value |
|---|---|
| Base model | tripathysagar/Qwen2.5-0.5B-GSM8K-SFT |
| Method | GRPO + LoRA |
| LoRA r / alpha | 16 / 32 |
| LoRA targets | all-linear |
| Steps | 1000 |
| Learning rate | 5e-05 |
| Beta (KL) | 0.02 |
| Num generations | 16 |
| Batch size | 4 x 8 (grad accum) |
| Max completion length | 512 |
| Precision | bf16 |
The answer is: {number}.