Views
No views yet
⚠️ Highly experimental — for tinkering, not production-grade. Recommend L3H5-DFlash over this one at the current training scale; L5H5 has more capacity but underperforms with only ~226 training shards.
[2, 16, 30, 43, 57]full_attention| k | 1 | 2 | 3 | 4 | 5 | 6 | 8 | 10 | 12 | 15 |
|---|---|---|---|---|---|---|---|---|---|---|
| % | 52.0 | 24.5 | 12.0 | 9.5 | 5.5 | 4.0 | 1.5 | 1.5 | 1.0 | 0.5 |
--speculative-config '{"method":"dflash","model":"MirecX/Qwen3.5-397B-A17B-L5H5-DFlash","num_speculative_tokens":6}'dflash_extract plugin (full residual stream). See L3H5 README for the historical bug context.