This dataset is a modified version of the MOSEL and VoxPopuli corpus, converted into parquet format to facilitate optimized I/O operations in high-performance and distributed computing environments. The MOSEL corpus is a multilingual dataset collection including up to 950K hours of open-source speech recordings covering the 24 official languages of the European Union.
MOSEL includes the automatic transcripts of 441k hours of unlabeled speech from VoxPopuli and… See the full description on the dataset page:
https://huggingface.co/datasets/meetween/mumospee_mosel.