DAVE-Corpus (open subset) is a training dataset for blind source separation (BSS)
and two-speaker speech separation on Chinese meeting speech. It is the redistributable
portion of the training pool of DAVE
(arXiv:2608.09288), our system for the ISCSLP 2026
Real-World AVSE Challenge, and is generated end-to-end by the released synthesis
pipeline from three permissively licensed corpora — AliMeeting, AISHELL-4
(speech) and MUSAN… See the full description on the dataset page: https://huggingface.co/datasets/TaurenMountain/DAVE-Corpus.