Full execution traces for every model run scored in ClawBench V1.
|π Leaderboard | π Benchmark | π Paper | π» Code | π Website |
This is the companion dataset to NAIL-Group/ClawBench. Where the main dataset publishes the task definitions (instructions, rubrics, eval schemas), this one publishes the raw execution data β one directory per (task Γ model Γ attempt), each with the screen recording, network capture, browser actions, agent reasoning, and the finalβ¦ See the full description on the dataset page:
https://huggingface.co/datasets/ZhangArthurHao/ClawBenchV1Trace.