Compact Harbor eval results for violetxi/qwen3-8b-terminal-wm-nextobs-mixed-clean on
Terminal-Bench 2.0 using the terminus-2 agent harness.
Each dataset split is one checkpoint step. Rows are Harbor trial directories and include binary reward,
per-test-case pass/fail data from verifier/ctrf.json, exception text when present, and run metadata.
Raw terminal recordings, panes, and completion logs are not… See the full description on the dataset page:
https://huggingface.co/datasets/violetxi/tb2-eval-qwen3-8b-wm-nextobs-mixed-clean.