A comprehensive benchmark for evaluating Role-play Agents in Chinese and English scenarios.
Role-play Benchmark is designed to evaluate Role-play Agents' ability to deliver immersive role-play experiences through Situated Reenactment. Unlike traditional benchmarks with verifiable answers, Role-play is fundamentally non-verifiable, e.g., there's no single "correct" response when a tsundere character is asked "Do you like me?".… See the full description on the dataset page:
https://huggingface.co/datasets/MiniMaxAI/role-play-bench.