ANCHOR is a benchmark for evaluating how well coding agents follow explicit constraints and instructions while solving real-world software engineering tasks. Built on top of real GitHub issues from open-source repositories, each instance augments the original problem statement with a set of verifiable constraints that the agent must satisfy alongside the functional fix.
Existing code generation benchmarks… See the full description on the dataset page:
https://huggingface.co/datasets/Yone-01/ANCHOR.