The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
Dataset Overview
This dataset contains the complete execution trajectories of 17 state-of-the-art language models evaluated on the Toolathlon benchmark. Toolathlon is a comprehensive benchmark for evaluating language agents on diverse, realistic, and long-horizon tasks.
Dataset Statistics: