VoiceGenEval: A Benchmark for Controllable Speech Generation in Spoken Language Models
Overview
VoiceGenEval is a bilingual (Chinese & English) benchmark for controllable speech generation. It covers four key tasks:
Acoustic attribute control
Natural language instruction following
Role-playing
Implicit empathy
VoiceGenEval goes beyond checking correctness — it evaluates how well the model speaks. Experiments on various… See the full description on the dataset page: https://huggingface.co/datasets/zhanjun/VStyle.