Anonymous reviewer-facing release for a NeurIPS 2026 D&B-track submission.
All identifying information has been redacted; full author / institution
attribution will be added with the camera-ready release.
RobustBench-TC is a sim-to-real robustness benchmark for tool-use language
agents. It augments five public single-turn tool-calling benchmarks (BFCL V3,
API-Bank, RoTBench, ToolAlpaca, ToolEyes) with 22 perturbation types
organized along the four components of the… See the full description on the dataset page:
https://huggingface.co/datasets/robustbench-tc/RobustBench-TC.