CUAStepBench is a benchmark of 278 human-annotated GUI-agent trajectories
for evaluating trajectory-level reward judges. Each trajectory records an
agent attempting a task in a real GUI environment (web, desktop, or mobile),
together with human ground-truth annotations of task success and per-step
quality.
It is the companion evaluation dataset of SeekJudge, an agentic reward
judge for GUI-agent trajectories.
Benchmark
Trajectories… See the full description on the dataset page:
https://huggingface.co/datasets/ZJUSCL/CUAStepBench.