This dataset contains tasks from Terminal-Bench, a benchmark for evaluating AI agents in real terminal environments. Each task is packaged as a complete, self-contained archive that preserves the exact directory structure, binary files, Docker configurations, and test scripts needed for faithful reproduction.
The archive column contains a gzipped tarball of the entire task directory.
Terminal-Bench evaluates AI agents on… See the full description on the dataset page:
https://huggingface.co/datasets/XUO/terminal-bench.