A benchmark for evaluating LLM agents on stateful, multi-step planning and tool use across realistic enterprise workflows.
The dataset is hosted on HuggingFace and can be loaded using the datasets library.
from datasets import load_dataset
calendar_tasks =… See the full description on the dataset page:
https://huggingface.co/datasets/EnterpriseAgents/EnterpriseOpsGym.