T1-Bench is a high-fidelity benchmark for evaluating task-completion and role-playing agents across 25 domains, including 11 single-domain and 14 multi-domain settings. It provides 76 tools and extensive human annotations, enabling systematic evaluation of agents in realistic, policy-grounded multi-domain interactions with natural user–assistant role-playing.
T1-Bench is a fully automated benchmark for… See the full description on the dataset page:
https://huggingface.co/datasets/gentaiscool/test2.