OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding.
Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task?
In real-world agentic coding, agents must… See the full description on the dataset page:
https://huggingface.co/datasets/yuan909815/OctoCodingBench.