Agent traces from InferenceBench
(GitHub), a benchmark that tests whether
frontier coding agents can optimize LLM serving under a fixed compute budget. The agents
know the techniques; the hard part is running, comparing, and keeping the ones that work.
Each run is one autonomous CLI agent attempting to deploy and optimize an OpenAI-compatible
inference server for a fixed base model (mistralai/Mistral-7B-Instruct-v0.3) under a… See the full description on the dataset page:
https://huggingface.co/datasets/aisa-group/InferenceBench-Trajectories.