Website | Paper | GitHub
Agent-Diff is a benchmarking framework for evaluating agentic Large Language Models (LLMs) on real-world tasks that execute code via external APIs. The benchmark provides access to real API interfaces (Slack, Box, Linear, Google Calendar) while sandboxing the environment in which calls are made and evaluated.
The dataset contains 224 tasks utilizing enterprise software workflows, provided with an 80/20… See the full description on the dataset page:
https://huggingface.co/datasets/hubertmarek/agent-diff-bench.