SPEED-Bench (SPEculative Evaluation Dataset) is a unified benchmark designed to evaluate speculative decoding (SD) across diverse semantic domains and realistic serving regimes, using production-grade inference engines.
It measures both acceptance-rate characteristics and end-to-end throughput, enabling fair, reproducible, and robust comparisons between SD strategies.
SPEED-Bench introduces a benchmarking ecosystem for⦠See the full description on the dataset page:
https://huggingface.co/datasets/nvidia/SPEED-Bench.