-
Modality-awareness — prosody, emotion, and acoustic naturalness
(real human speech vs. synthetic TTS).
-
Colloquialness — spontaneous spoken style vs. scripted written style.
-
📄 Paper:
Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness (ACL 2026 Main Conference) —
arXiv:2603.14889
-
-
📚 Training data (gated):
SDiaReward dataset.
-
🧩 Smaller variant:
SDiaReward-3B.
Trained with TRL's reward trainer on the
SDiaReward preference dataset
(11,630 episode-level preference pairs, ~200 hours of paired speech). Backbone:
Qwen2.5-Omni-7B; pooling variant:
mean_center; 2 epochs.
Released under Apache-2.0 for research use. The reward signal is intended for
evaluating and improving spoken-dialogue systems; it is not a safety classifier.
1@article{lu2026modeling,
2 title={Modeling and benchmarking spoken dialogue rewards with modality and colloquialness},
3 author={Lu, Jingyu and Wang, Yuhan and Zhuo, Fan and Cheng, Xize and Pan, Changhao and Pu, Xueyi and Chen, Yifu and Wen, Chenyuhao and Liang, Tianle and Zhao, Zhou},
4 journal={arXiv preprint arXiv:2603.14889},
5 year={2026}
6}