A benchmark for evaluating language-model agents on transformer circuit explanation. Given an already-localized circuit, the agent must recover what each component does: a functional role tag from a 5-class taxonomy, a task-specific natural-language note, and a derived description of the overall task.
The benchmark has 84 semi-synthetic circuits with 163 annotated components, plus a manually annotated real-model circuit (three-operand addition in Llama-3-8B).… See the full description on the dataset page:
https://huggingface.co/datasets/Antik/AgenticInterpBench.