unsloth/Qwen2.5-1.5B-Instruct. Built by
scripts/build_tools_gguf.sh Qwen2.5-1.5B-tools (merge_and_unload of the
latest checkpoint in checkpoints/Qwen2.5-1.5B-tools/).TrainFinetuneRecipeForNextTokenPrediction (NeMo AutoModel 0.5.0)configs/qwen25_1.5b_tools.yaml*.projsft_tools split of r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillationmodels/Qwen2.5-1.5B-tools-GGUF/models/Qwen2.5-1.5B-tools-NVFP4/sft_tools validation split (greedy decoding, 384 max new tokens). Throughput is single-stream greedy decode, not serving throughput.| Model | Quant | n | Tool call emitted | Names match | Exact args match | Δ exact vs BASE | tok/s |
|---|---|---|---|---|---|---|---|
| Qwen2.5-1.5B-tools | BASE (unsloth/Qwen2.5-1.5B-Instruct) | 274 | 272/274 (99.3%) | 59/274 (21.5%) | 24/274 (8.8%) | — | 40.6 |
| Qwen2.5-1.5B-tools | BF16 | 274 | 258/274 (94.2%) | 216/274 (78.8%) | 2/274 (0.7%) | -8pp | 38.0 |
| Qwen2.5-1.5B-tools | GGUF-BF16 | 274 | 272/274 (99.3%) | 167/274 (60.9%) | 2/274 (0.7%) | -8pp | 88.6 |
| Qwen2.5-1.5B-tools | GGUF-Q4_K_M | 274 | 248/274 (90.5%) | 49/274 (17.9%) | 3/274 (1.1%) | -7.7pp | 142.9 |
| Qwen2.5-1.5B-tools | GGUF-Q5_K_M | 274 | 260/274 (94.9%) | 79/274 (28.8%) | 6/274 (2.2%) | -6.6pp | 131.4 |
| Qwen2.5-1.5B-tools | GGUF-Q8_0 | 274 | 271/274 (98.9%) | 187/274 (68.2%) | 2/274 (0.7%) | -8pp | 115.3 |