This repository contains the ACEBench dataset, formatted for evaluating and training tool-using language models. The dataset has been processed into a unified structure, with problem descriptions merged with their corresponding ground-truth rubrics.
Notebook used to format the dataset: Open in Colab
The dataset is provided under a single configuration, en, which contains three distinct splits:
normal: Standard tool-use scenarios.… See the full description on the dataset page:
https://huggingface.co/datasets/oliveirabruno01/acebench.