fc-reward-bench is a benchmark designed to evaluate reward model performance in function-calling tasks. It features 1,500 unique user inputs derived from the single-turn splits of the BFCL-v3 dataset. Each input is paired with both correct and incorrect function calls. Correct calls are sourced directly from BFCL, while incorrect calls are generated by 25 permissively licensed models.
Performance of ToolRM, top reward models from… See the full description on the dataset page:
https://huggingface.co/datasets/ibm-research/fc-reward-bench.