AgentHorizon is a benchmark for evaluating LLM judges of computer-use agents. Each item is a recorded trajectory of a GUI agent attempting a long-horizon desktop task (often 100 to 300+ steps), paired with a ground-truth label that says whether the agent actually completed the task. A judge reads the trajectory (instruction, action sequence, and screenshots) and predicts success or failure.
Path
What it is… See the full description on the dataset page:
https://huggingface.co/datasets/ServiceNow/AgentHorizon.