This dataset is a part of ToolRM: Towards Agentic Tool-Use Reward Modeling and serves as a dedicated benchmark for evaluating reward models in tool-use settings. It comprises 2,983 preference annotations buit upon BFCL V3, with assistant responses extracted from archived trajectories available in this github repo.
ToolRM is a family of lightweight generative and… See the full description on the dataset page:
https://huggingface.co/datasets/RioLee/TRBench-BFCL.