MISP-M3SD is a large-scale multimodal, multi-scenario, and multilingual dataset for robust speaker diarization, constructed from in-the-wild online videos. It contains more than 770 hours of synchronised audio-visual recordings, covering 14 scenarios and 16 languages.
The dataset is designed to support the development of speaker diarization systems with stronger cross-domain generalisation under realistic conditions, including background noise… See the full description on the dataset page: https://huggingface.co/datasets/Igor97/MISP-M3SD.