Throughput, latency and resource benchmarks for on-premise LLM inference across a wide range
of GPUs — datacenter (B200, L40S), Blackwell workstation (RTX PRO 6000), consumer (RTX 4090/4070),
unified-memory (DGX Spark / GB10) and legacy/AMD — covering the full stack of a production
assistant: chat serving, embeddings, reranking, vision, OCR and speech-to-text.
Engine: vLLM (mostly the NGC 26.03 image, engine v0.17.1; Blackwell… See the full description on the dataset page:
https://huggingface.co/datasets/Scherm-AI/onprem-llm-benchmark.