sol-max-v2-record
Complete run record for AgentPTB cell sol-max-v2 — Codex / gpt-5.6-sol @ effort max.
A redo of the sol-max cell from hour 0 on node tb-1, after the original attempt died at
~h16. This one ran the full 100 hours (boot 2026-08-19T19:15:00Z).
| field | value |
|---|
| plot cell | sol-max-v2 |
| driver | Codex / gpt-5.6-sol |
| reasoning effort | max |
| run boot (UTC) | 2026-08-19T19:15:00Z |
| checkpoints published | 36 (agentic-ptb/sol-max-v2.h*) |
| submitted checkpoint | sol-max-v2.h007.pi-agent-sft-v5.step_600 |
The arm submitted an hour-7 checkpoint
Of 36 checkpoints spanning 82 hours, the arm selected one written at h7 — choosing it over
75 further hours of its own training. Its self-reported full-suite numbers:
| harness | suite | episodes | score | 95% CI |
|---|
| submitted | terminal-bench-2 | 89 | 4.49 | [1.76, 10.99] |
| submitted | swe-bench-verified | 500 | 22.40 | [18.96, 26.26] |
| stock-compatible | terminal-bench-2 | 89 | 2.25 | [0.62, 7.83] |
| stock-compatible | swe-bench-verified | 500 | 20.40 | [17.10, 24.15] |
These are the arm's own measurements under its own harness and sample size, and are not
comparable across cells. The controlled cross-cell re-measurement is the sweep in INDEX.
Contents
| path | what |
|---|
driver-session/ | the full driver trajectory — every Codex turn, 576 event files |
evals/ | the arm's own eval runs and logs |
harness/ | harness source, trainer/eval configs, and scripts the arm wrote |
submission/ | final submission record, manifest, and audit |
RUNLOG.md | the arm's own running log of decisions |
supervisor.log | supervisor cycle history |
codex_home/ is intentionally not published — it contains live driver credentials.
Related
agentic-ptb/sol-max-v2.h* — the 36 checkpoints from this run
agentic-ptb/sol-max-v2-data — the training corpus the arm built
agentic-ptb/INDEX — manifest joining every checkpoint to the sweep figures