If you use an environment wrapper, set the same JSON as SPECULATIVE_CONFIG.
The verifier model still needs the serving requirements documented on the NVFP4
model page, including the LLM-jp remote code/plugin setup.
Benchmark (DGX Spark / GB10, SM121)
Representative Japanese streaming decode benchmark on ELYZA-100 prompts,
measured on NVIDIA DGX Spark / GB10 (SM121) with vLLM 0.24.0, stock vLLM
DFlash, temperature 0, serial streaming requests, and exactly 128 generated
tokens per request. Decode speed excludes time-to-first-token.
Setup
Mean decode tok/s
Speedup vs BF16
Notes
BF16 base model
14.42
1.00x
llm-jp/llm-jp-4-8b-instruct, no speculative decoding
NVFP4 verifier
29.73
2.06x
kel-jp/llm-jp-4-8b-instruct-NVFP4, no speculative decoding
NVFP4 + this DFlash drafter
49.04
3.40x
stock vLLM DFlash, num_speculative_tokens=3
The DFlash row is 1.65x faster than the NVFP4 verifier baseline by ratio of
standalone means. In the paired per-prompt run against the same NVFP4 baseline,
the DFlash mean was 48.76 tok/s and the ratio of mean decode throughput was
1.64x.
Paired speedup summary:
Ratio of mean decode throughput: 1.64x
Mean paired request speedup: 1.64x
Median paired request speedup: 1.63x
p05/p95 paired speedup: 1.37x / 1.97x
Requests at least 1.5x faster: 70 / 100
Requests at least 2.0x faster: 5 / 100
Requests slower than baseline: 0 / 100
The raw benchmark artifacts are included under benchmark/:
bf16-elyza100-streaming.json
bf16-elyza100-streaming.csv
baseline-elyza100-streaming.json
baseline-elyza100-streaming.csv
dflash20k-b4-l1-v28k-elyza100-streaming.json
dflash20k-b4-l1-v28k-elyza100-streaming.csv
elyza100-bf16-nvfp4-dflash-speedup-summary.json
elyza100-bf16-nvfp4-dflash-speedup.png
elyza100-baseline-vs-dflash-summary.json
elyza100-baseline-vs-dflash-speedup.png
dflash20k-b4-l1-elyza100-tps-histogram.png
BF16, NVFP4, and DFlash speedup distribution
Validation Metrics
The released checkpoint's held-out validation metrics: