Compact Harbor eval results for violetxi/qwen3-8b-terminal-wm-summary-mixed-clean-vista on
Terminal-Bench 2.0 using the terminus-2 agent harness.
Each dataset split is one checkpoint step. Rows are Harbor trial directories and include binary reward,
per-test-case pass/fail data from verifier/ctrf.json, exception text when present, and run metadata.
Raw terminal recordings, panes, and completion logs are… See the full description on the dataset page:
https://huggingface.co/datasets/violetxi/tb2-eval-qwen3-8b-wm-summary-mixed-clean-vista.