Compact Harbor eval results for violetxi/qwen3-8b-terminal-action-clean on
Terminal-Bench 2.0 using the terminus-2 agent harness.
Each dataset split is one checkpoint step. Rows are Harbor trial directories and include binary reward,
per-test-case pass/fail data from verifier/ctrf.json, exception text when present, and run metadata.
Raw terminal recordings, panes, and completion logs are not included.
Splits… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb2-eval-qwen3-8b-action-clean.