A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception.
CAP-Bench evaluates browser agents on Cross-site workflows, complex Actions, and challenging visual Perception. The full benchmark contains 420 tasks across 108 real-world websites in 24 functional domains. Each task requires on average 7 complex execution operations and 4 perception challenges, substantially exceeding the difficulty of prior browser-agent benchmarks.
This… See the full description on the dataset page:
https://huggingface.co/datasets/Warrior0302/CAP-Bench.