Views
No views yet
| Parameter | Value |
|---|---|
| Base model | Qwen/Qwen2.5-0.5B |
| Method | SFT + LoRA |
| LoRA r / alpha | 32 / 16 |
| LoRA targets | all-linear |
| Epochs | 1 |
| Learning rate | 0.0002 |
| Batch size | 8 × 4 (grad accum) |
| Training examples | 1024 |
| Precision | bf16 |
The answer is: {number}.You are a helpful math assistant. Solve the problem step by step, then give your final answer as a single number on the last line in exact format
The answer is: {number}.