TL;DR: 4,576 benchmark runs measuring speculative decoding speedup / acceptance rate
across llama.cpp and LM Studio, Qwen3 (8B/14B) and Llama-3.1-8B target models, on a
single consumer RTX 4090 (24GB). Best observed case: the draft-free ngram-mod
self-speculative mode on structured tasks (JSON extraction 2.81x, code 2.76x,
global-median aggregation at temp=0). Open-ended tasks (creative writing, translation)
with a traditional draft… See the full description on the dataset page:
https://huggingface.co/datasets/steven0226/speculative-decoding-bench-rtx4090.