End-to-end transparent benchmark for AI agents acting in the real world.
Paper | Leaderboard | Code
Split
Examples
Description
general
161
Core agent tasks across 24 categories (communication, finance, ops, productivity, etc.)
multimodal
101
Multimodal agentic tasks requiring perception and creation (webpage generation, video QA, document extraction, etc.)
multi_turn
38
Multi-turn conversational tasks where the… See the full description on the dataset page:
https://huggingface.co/datasets/claw-eval/Claw-Eval.