The INTERSPEECH 2025 MLC-SLM Challenge Dataset, curated by Datatang, is derived from fifteen proprietary conversational speech corpora. Distinguished by exceptional annotation accuracy and operational reliability, this dataset is engineered to address critical challenges in multilingual automatic speech recognition (ASR) and long-context comprehension. It meticulously replicates real-world complexities including spontaneous interruptions and speaker overlaps across… See the full description on the dataset page:
https://huggingface.co/datasets/Nexdata-AI/INTERSPEECH-2025-MLC-SLM-Challenge-Data-Sample.