SDiaReward-Dataset & ESDR-Bench
Preference data and evaluation benchmark for SDiaReward, a reward model for
spoken dialogue that scores multi-turn conversations along two axes:
Modality-awareness — prosody, emotion, acoustic naturalness (real human speech vs. synthetic TTS).
Colloquialness — spontaneous spoken style vs. scripted written style.
The model backbone is Qwen2.5-Omni extended with a pooling layer and a linear
reward head. See the paper and code for details.