A benchmark for the moment before an agent acts.
An agent has read the context, chosen a tool action, and is about to change the
world: send a message, update a record, merge code, charge a card. SteerBench-Work
asks one question at that boundary: should the action proceed, or should the system
hold for review?
Both mistakes are scored. Acting on work it should have held, and holding work it
was cleared to do.