This datacard provides a comprehensive description of YouTube-Commons and EUVoxCommons (European Parliament proceedings) collected and handled by pleias along with a sample of 1,304 documented audio files.
These datasets represent the largest collection of fully open-source copyright-compliant speech data for the 24 official languages of the European Union and more.
Key… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/audio_samples.