Comprehensive benchmark for evaluating AI agent capabilities across three core competencies:
Tool Selection - Choosing appropriate tools for tasks
Task Planning - Decomposing complex tasks into step sequences
Timing Judgment - Deciding when to use tools vs. direct answers
Dataset Statistics
Total Samples: ~11,000
Tool Selection: ~6,000 samples
Task Planning: ~3,000 samples
Timing Judgment: ~2,000 samples
Splits: train, dev, test
Usage… See the full description on the dataset page: https://huggingface.co/datasets/intellistream/sage-agent-benchmark.