Views
No views yet
| Metric | Base | This Model | Paper |
|---|---|---|---|
| Greedy Accuracy | 54.4% | 64.7% | 70.6% |
| Pass@1 | 52.6% | 56.2% | — |
| Pass@5 | 61.5% | 70.1% | — |
| Pass@10 | 64.4% | 74.4% | — |
| Pass@50 | 70.6% | 79.4% | — |
| Parameter | Value |
|---|---|
| Base model | Qwen/Qwen2.5-7B-Instruct |
| Method | On-policy Self-Distillation (SDFT) |
| Dataset | ToolAlpaca (4046 train, 68 test) |
| Learning rate | 1e-5 |
| Batch size | 32 |
| Epochs | 2 |
| EMA alpha | 0.01 |
| Step | 1000 (best of 1011) |
| Hardware | L40S 48GB |
| Step | Greedy Acc |
|---|---|
| 100 | 55.9% |
| 200 | 48.5% |
| 300 | 44.1% |
| 400 | 47.1% |
| 500 | 57.4% |
| 600 | 47.1% |
| 700 | 54.4% |
| 800 | 52.9% |
| 900 | 57.4% |
| 1000 | 64.7% |
| 1011 | 57.4% |