Agent-generated optimization patches benchmarked on real GPU hardware (NVIDIA H100 80GB).
Scope: Lossfunk/ISO-Bench — 54 tasks (39 vLLM + 15 SGLang)
Patches from: Inferencebench/pass-at-k-samples
Codex CLI (GPT-5)
214
74
288… See the full description on the dataset page:
https://huggingface.co/datasets/Inferencebench/pass-at-k-benchmark-results.