MultiTalkBench is the first benchmark to jointly evaluate long, multi-party, and bilingual full-duplex dialogue. It tests speech-to-speech systems on:
(a) Long interactions — conversations longer than ten minutes, with explicit probes for long-range entity tracking and topic coherence.
(b) One-model-many-user multi-party interaction — quantitative addressee-selection and turn-taking metrics.
(c) Chinese–English bilingual ability.
To our knowledge, no prior benchmark… See the full description on the dataset page:
https://huggingface.co/datasets/MultiTalk/MultiTalkBench.