Execution trajectories and LLM-as-judge scores for all experiments reported in
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents (Zheng et al., 2026),
evaluated on the AgentIF-OneDay benchmark (104 tasks, 767 instance-level rubric points).
This bundle contains the raw evidence chain for every reported run: the full
agent trajectory, the final deliverable artifacts, and the per-criterion judge
scores. Every headline number in the paper… See the full description on the dataset page:
https://huggingface.co/datasets/zjunlp/onedayagent_traj.