110 agent trajectories on ST-WebAgentBench (SuiteCRM), generated by
plan-guided execution and kept only when a database check confirmed the
episode's change actually persisted.
The benchmark's own program_html task evaluators read the agent's current
page, which still contains anything it merely typed. An episode that fills a
form and never saves therefore scores reward 1.0. Traces from an… See the full description on the dataset page:
https://huggingface.co/datasets/icrl-finetuning/2026-08-17-stwebagentbench-expert_synthetic.