A synthetic benchmark for evaluating tool-calling (function-calling) behavior
in language models: does the model pick the right tool, fill arguments
correctly, and correctly decline to call a tool when nothing in the catalog
applies. The dataset has three configs: