This dataset is released with the COLM 2025 conference paper: "PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?".
PersonaEval is the first benchmark designed to test whether Large Language Models (LLMs) can reliably identify character roles from natural dialogue. We argue that correctly identifying who is speaking is a fundamental prerequisite for any meaningful evaluation of role-playing quality (how… See the full description on the dataset page:
https://huggingface.co/datasets/lingfengzhou/PersonaEval.