TAB measures whether a terminal agent does what the user asked, and only what the user asked.
Each task starts from a Terminal-Bench 2.1 task whose instruction has been stripped down to an abstracted version (typically 70–95% of the wording removed) while preserving the goal. The removed part is restored as a helpful cue placed somewhere the agent will naturally read while solving the task. On the same surface we also place an irrelevant distractor: a… See the full description on the dataset page:
https://huggingface.co/datasets/anonymous98570239854/tab.