Can AI Agents Complete Everyday Online Tasks?
ClawBench evaluates AI agents on 153 everyday tasks (such as booking flights, ordering groceries, submitting job applications) across 144 live websites. We capture 5 layers of behavioral data (session replay, screenshots, HTTP traffic, agent reasoning traces, and browser actions), collect human ground-truth for every task, and score with an agentic evaluator that provides step-level traceable diagnostics.
Paper… See the full description on the dataset page:
https://huggingface.co/datasets/Duke313/ClawBench-test.