Hard, practical, end-to-end evaluation for AI agents — in the wild.
WildClawBench is an agent benchmark that tests what actually matters: can an AI agent do real work, end-to-end, without hand-holding?
We drop agents into a live OpenClaw environment — the same open-source personal AI assistant that real users rely on daily — and throw 60 original tasks at them: clipping goal highlights from a football match, negotiating meeting times over multi-round… See the full description on the dataset page:
https://huggingface.co/datasets/internlm/WildClawBench.