TL;DR — On a real H200 running vLLM, a dynamic controller that reallocates a fixed 96-slot admission budget toward live demand beat a frozen even 48/48 split by +14.0% total throughput (11,274 vs 9,894 tok/s) and +18.0% batch throughput (7,736 vs 6,557 tok/s) at equal-or-better p99 (4.08 vs 4.15 s). Honest caveat: steady-window GPU utilization reached only 67.2% mean (99% peak), short of a sustained 90% target… See the full description on the dataset page:
https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-21-agent-dynamic-batch-tuning-vllm.