Welcome to Terminal-Bench-2.0! If you’re reading this you’re a member of the Terminal-Bench community that we’ve selected to get a sneak peek at the latest version of the benchmark.
First, clone Harbor (formerly “Sandboxes”):
git clone
https://github.com/laude-institute/harbor.git
This will install Harbor, our new package for running agent evals.
You should now be able to run TB 2.0!… See the full description on the dataset page:
https://huggingface.co/datasets/penfever/terminal-bench-2.