TL;DR — Inference platforms ship conservative engine defaults that silently set every endpoint's throughput ceiling. On one controlled three-arm production case (single B200, vLLM 0.24.0, NVFP4 weights) the total tax at the top of the concurrency ladder is at least 17.9x and factors multiplicatively into a 10.0x compile/launch component and a concurrency-ceiling… See the full description on the dataset page:
https://huggingface.co/datasets/thaki-AI/daily-paper-2026-08-24-default-configuration-tax.