This is a block-size-16 Domino draft model for the target model
Qwen/Qwen3.6-27B. It is intended
for speculative decoding with the Domino-enabled SGLang branch.
Methods. AR is target-only decoding. MTP-S3/S7/S15 use the built-in Qwen3.6 MTP heads with 3/7/15 steps, 4/8/16 draft tokens, and top-k 1. DFlash uses the official z-lab/Qwen3.6-27B-DFlash checkpoint at revision 0919688.
For DFlash and Domino, b16 is the regular block-16 run and the target verifies all 16 positions. b8 keeps the same block-16 draft backbone but the target verifies only the first 8 positions. The draft backbone itself is not shortened.
Setting.Qwen/Qwen3.6-27B, TP2/BF16 on 2×A100 80GB, FlashInfer, O4096, thinking enabled, greedy sampling (temperature=0, top_p=1, top_k=1), C1/C8/C32, and three fresh-server repeats per cell. Workloads are GSM8K-128, MATH500-128, HumanEval-164, MBPP-128, MT-Bench-80, and Alpaca-128.
Qwen3.6-27B serving throughput
Each bar is mean output tok/s over three runs. The black outline marks the fastest speculative configuration for each workload.
Throughput and speedup
Each cell is output tok/s (speedup versus AR). Bold marks the fastest speculative configuration in each row.
Concurrency 1
Workload
AR
MTP-S3
MTP-S7
MTP-S15
DFlash b8
DFlash b16
Domino b8
Domino b16
GSM8K
47.2 (1.00x)
126.8 (2.68x)
151.9 (3.22x)
133.9 (2.84x)
179.2 (3.79x)
200.9 (4.25x)
206.0 (4.36x)
248.0 (5.25x)
MATH500
47.3 (1.00x)
132.3 (2.80x)
168.1 (3.55x)
151.6 (3.20x)
203.2 (4.29x)
240.1 (5.07x)
217.7 (4.60x)
270.6 (5.72x)
HumanEval
47.2 (1.00x)
125.3 (2.65x)
149.9 (3.18x)
129.9 (2.75x)
188.0 (3.98x)
211.1 (4.47x)
198.5 (4.20x)
235.1 (4.98x)
MBPP
47.6 (1.00x)
122.6 (2.57x)
141.8 (2.98x)
117.0 (2.46x)
177.5 (3.73x)
186.2 (3.91x)
189.0 (3.97x)
214.0 (4.49x)
MT-Bench
47.1 (1.00x)
115.1 (2.44x)
125.0 (2.65x)
100.7 (2.14x)
141.5 (3.00x)
143.8 (3.05x)
155.7 (3.31x)
161.9 (3.44x)
Alpaca
47.2 (1.00x)
112.1 (2.38x)
119.8 (2.54x)
96.0 (2.03x)
135.7 (2.87x)
133.9 (2.84x)
150.1 (3.18x)
157.7 (3.34x)
Concurrency 8
Workload
AR
MTP-S3
MTP-S7
MTP-S15
DFlash b8
DFlash b16
Domino b8
Domino b16
GSM8K
324.4 (1.00x)
778.7 (2.40x)
913.0 (2.81x)
673.3 (2.08x)
1016.1 (3.13x)
899.0 (2.77x)
1153.2 (3.55x)
1131.2 (3.49x)
MATH500
331.6 (1.00x)
846.3 (2.55x)
1056.6 (3.19x)
805.6 (2.43x)
1191.4 (3.59x)
1107.4 (3.34x)
1255.2 (3.78x)
1252.5 (3.78x)
HumanEval
331.9 (1.00x)
797.5 (2.40x)
939.0 (2.83x)
684.4 (2.06x)
1104.2 (3.33x)
985.8 (2.97x)
1155.9 (3.48x)
1078.5 (3.25x)
MBPP
326.7 (1.00x)
756.2 (2.32x)
872.3 (2.67x)
564.5 (1.73x)
995.4 (3.05x)
850.4 (2.60x)
1053.7 (3.23x)
950.0 (2.91x)
MT-Bench
330.6 (1.00x)
713.7 (2.16x)
776.3 (2.35x)
528.6 (1.60x)
808.0 (2.44x)
655.2 (1.98x)
881.5 (2.67x)
752.8 (2.28x)
Alpaca
332.6 (1.00x)
719.8 (2.16x)
740.9 (2.23x)
510.9 (1.54x)
786.8 (2.37x)
616.7 (1.85x)
862.9 (2.59x)
722.7 (2.17x)
Concurrency 32
Workload
AR
MTP-S3
MTP-S7
MTP-S15
DFlash b8
DFlash b16
Domino b8
Domino b16
GSM8K
862.9 (1.00x)
1555.9 (1.80x)
1495.6 (1.73x)
1034.2 (1.20x)
1601.7 (1.86x)
1214.0 (1.41x)
1817.5 (2.11x)
1574.5 (1.82x)
MATH500
963.3 (1.00x)
1834.1 (1.90x)
1795.0 (1.86x)
1250.0 (1.30x)
1920.0 (1.99x)
1478.7 (1.54x)
2021.3 (2.10x)
1716.2 (1.78x)
HumanEval
953.4 (1.00x)
1701.1 (1.78x)
1607.3 (1.69x)
1099.3 (1.15x)
1795.8 (1.88x)
1351.4 (1.42x)
1879.7 (1.97x)
1541.3 (1.62x)
MBPP
919.5 (1.00x)
1577.2 (1.72x)
1443.2 (1.57x)
948.5 (1.03x)
1595.0 (1.73x)
1218.6 (1.33x)
1547.2 (1.68x)
1342.4 (1.46x)
MT-Bench
866.8 (1.00x)
1352.0 (1.56x)
1144.4 (1.32x)
721.8 (0.83x)
1168.1 (1.35x)
817.3 (0.94x)
1286.9 (1.48x)
920.3 (1.06x)
Alpaca
788.1 (1.00x)
1387.6 (1.76x)
1130.1 (1.43x)
718.4 (0.91x)
1128.9 (1.43x)
790.0 (1.00x)
1315.0 (1.67x)
894.1 (1.13x)
Macro speedup vs AR
Arithmetic mean of the per-workload TPS ratios; Overall is the arithmetic mean of all 18 workload-by-concurrency ratios.
C
MTP-S3
MTP-S7
MTP-S15
DFlash b8
DFlash b16
Domino b8
Domino b16
1
2.59x
3.02x
2.57x
3.61x
3.93x
3.94x
4.54x
8
2.33x
2.68x
1.90x
2.98x
2.59x
3.22x
2.98x
32
1.75x
1.60x
1.07x
1.71x
1.27x
1.84x
1.48x
Overall
2.22x
2.43x
1.85x
2.77x
2.60x
3.00x
3.00x
Accept length
Mean output tokens per target verification step, including the target bonus token. The maxima are 4/8/16 for MTP-S3/S7/S15, 8 for b8, and 16 for b16. Bold marks the highest value within the directly comparable max-8 and max-16 groups.
Concurrency 1
Workload
MTP-S3
MTP-S7
MTP-S15
DFlash b8
DFlash b16
Domino b8
Domino b16
GSM8K
3.556
5.503
6.855
5.449
6.878
6.335
9.163
MATH500
3.608
5.658
7.055
5.700
7.401
6.201
8.760
HumanEval
3.422
5.056
6.048
5.264
6.469
5.655
7.445
MBPP
3.349
4.782
5.439
4.972
5.732
5.405
6.803
MT-Bench
3.206
4.487
5.188
4.278
4.947
4.758
5.851
Alpaca
3.182
4.428
5.032
4.176
4.698
4.707
5.937
Concurrency 8
Workload
MTP-S3
MTP-S7
MTP-S15
DFlash b8
DFlash b16
Domino b8
Domino b16
GSM8K
3.544
5.532
6.819
5.429
6.796
6.309
9.310
MATH500
3.604
5.660
7.001
5.663
7.386
6.194
8.813
HumanEval
3.429
5.055
6.024
5.267
6.440
5.659
7.418
MBPP
3.338
4.787
5.442
4.954
5.770
5.404
6.787
MT-Bench
3.198
4.496
5.201
4.294
4.967
4.749
5.908
Alpaca
3.203
4.415
5.022
4.146
4.669
4.743
5.835
Concurrency 32
Workload
MTP-S3
MTP-S7
MTP-S15
DFlash b8
DFlash b16
Domino b8
Domino b16
GSM8K
3.554
5.510
6.889
5.442
6.873
6.288
9.313
MATH500
3.608
5.646
7.032
5.697
7.321
6.206
8.789
HumanEval
3.423
5.061
6.034
5.293
6.466
5.677
7.449
MBPP
3.346
4.755
5.497
4.942
5.752
5.373
6.842
MT-Bench
3.197
4.505
5.224
4.302
4.983
4.768
5.923
Alpaca
3.185
4.426
5.006
4.171
4.692
4.721
5.812
All 144 displayed method/workload/concurrency cells completed three measured runs.
These are observed serving results. BF16 greedy trajectories can differ across methods, so the TPS differences are not pure kernel-attribution claims.
Citation
bibtex
1@article{huang2026domino,
2 title={Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding},
3 author={Huang, Jianuo and Zhang, Yaojie and Zhang, Qituan and Lin, Hao and Xu, Hanlin and Zhang, Linfeng},
4 journal={arXiv preprint arXiv:2605.29707},
5 year={2026}
6}
Acknowledgements
Domino builds on DFlash and its block-diffusion
drafting formulation.