Views
No views yet
--spec-type draft-mtp) · Q6_K quant| Benchmark | Score | Accuracy | Results | Time | vs No-MTP |
|---|---|---|---|---|---|
| ToolCall-15 🛠️ | 1500/1500 | 100% | 15✅ 0⚠️ 0❌ | 0.65min | 1.4× faster |
| HermesAgent-20 🤖 | 1505/2000 | 75.3% | 12✅ 1⚠️ 7❌ | 5.3min | 1.17× faster |
| BugFind-15 🐛 | 1030/1500 | 68.7% | 9✅ 2⚠️ 4❌ | 1.8min | 4.1× faster |
| Total | 4035/5000 | 80.7% | 36✅ 3⚠️ 11❌ | 7.8min | 1.9× faster |
| Benchmark | Without MTP | With MTP | Δ Score | Δ Speed |
|---|---|---|---|---|
| ToolCall-15 🛠️ | 1400/1500 (93.3%) | 1500/1500 (100%) | +100 pts | 1.4× |
| HermesAgent-20 🤖 | 1545/2000 (77.2%) | 1505/2000 (75.3%) | −40 pts | 1.17× |
| BugFind-15 🐛 | 928/1500 (61.9%) | 1030/1500 (68.7%) | +102 pts | 4.1× |
| Total | 3873/5000 (77.5%) | 4035/5000 (80.7%) | +162 pts | 1.9× |
| Total Time | 14.7 min | 7.8 min | — | 1.9× |
| TC-ID | Result | Scenario |
|---|---|---|
| TC-01–TC-04 | ✅ | Simple / Multi / Nested / Type conversion |
| TC-05 | ✅ | Relative date/time parsing ← fixed by MTP |
| TC-06–TC-15 | ✅ | All remaining scenarios |
| BF-ID | Without MTP | With MTP | Δ |
|---|---|---|---|
| BF-01 | ✅ 100 | ✅ 100 | — |
| BF-02 | ✅ 88 | ✅ 100 | +12 |
| BF-03 | ❌ 0 | ❌ 0 | — |
| BF-04 | ✅ 100 | ✅ 100 | — |
| BF-05 | ❌ 40 | ⚠️ 70 | +30 |
| BF-06 | ❌ 0 | ❌ 0 | — |
| BF-07 | ✅ 100 | ✅ 100 | — |
| BF-08 | ✅ 100 | ✅ 100 | — |
| BF-09 | ✅ 100 | ✅ 100 | — |
| BF-10 | ❌ 0 | ❌ 0 | — |
| BF-11 | ⚠️ 60 | ✅ 100 | +40 |
| BF-12 | ❌ 0 (timeout) | ✅ 100 | +100 |
| BF-13 | ✅ 100 | ✅ 100 | — |
| BF-14 | ⚠️ 70 | ⚠️ 60 | −10 |
| BF-15 | ⚠️ 70 | ⚠️ 60 | −10 |
mtp_num_hidden_layers: 1 remained, but no actual tensors existed in the safetensors).--spec-type draft-mtp to activate.| Model | Total | ToolCall-15 | HermesAgent-20 | BugFind-15 | Total Time |
|---|---|---|---|---|---|
| 🐾 QwenPaw MTP 9B | 4035 🥇 | 100% 🥇 | 75.3% | 68.7% | 7.8min 🥇 |
| 🐾 QwenPaw 9B (no MTP) | 3873 | 93.3% | 77.2% 🥇 | 61.9% | 14.7min |
| 🧠 Qwopus 9B MTP | 3935 | 93.3% | 67.3% ⚠️ | 79.0% 🥇 | 21.3min ⚠️ |
| 🧠 Qwen 35B Thinking ON | 1445 (HA only) | — | 72.3% | — | 7.0min |
| ⚡ Qwen 35B Thinking OFF | 1370 (HA only) | — | 68.5% | — | 5.1min |
| 🔮 Gemma 4 26B | 1405 (HA only) | — | 70.3% | — | 18.6min |
--spec-type draft-mtp with --spec-draft-n-max 2| File | Size | Notes |
|---|---|---|
QwenPaw-Flash-9B-heretic-MTP-Q8_0.gguf | ~9.2GB | High quality, near lossless |
QwenPaw-Flash-9B-heretic-MTP-Q6_K.gguf | ~7.1GB | ✅ Recommended, best value |
QwenPaw-Flash-9B-heretic-MTP-Q4_K_M.gguf | ~5.4GB | Compact |
mmproj-BF16 | ~880MB | Vision encoder (multimodal) — same as non-MTP version |
--spec-type — if not, the model functions as a standard 9B model.