RealTalk-CN is the first large-scale, multi-domain, bimodal (speech-text) Chinese Task-Oriented Dialogue (TOD) dataset. All data come from real human-to-human conversations, specifically constructed to advance research on speech-based large language models (Speech LLMs). Existing TOD datasets are mostly text-based, lacking realโฆ See the full description on the dataset page:
https://huggingface.co/datasets/BAAI/RealTalk-CN.