The Common Mozilla Voice dataset consists of a unique MP3 and corresponding text file.
Many of the 26119 recorded hours in the dataset also include demographic metadata like age, sex, and accent that can help improve the accuracy of speech recognition engines.
The dataset currently consists of 17127 validated hours in 104 languages, but more voices and languages are always added.
Before using this dataloader, please accept the acknowledgement at
https://huggingface.co/datasets/mozilla-foundation/common_voice_12_0 and use huggingface-cli login for authentication