A multi-turn agentic coding benchmark that needs no Docker. The model is
given a small but real repository, a deliberately narrow tool set, and a bug
report. It drives its own conversation until it declares itself finished, and is
then graded by running a test suite it was never allowed to see.
Built for llama-eval (--dataset agentic),
but the records are self-contained and usable by any harness.
Most agentic… See the full description on the dataset page:
https://huggingface.co/datasets/ilintar/SACB.