This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website
https://commonvoice.mozilla.org/.
NOTE: currently converting to parquet for convenience.. WIP
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese… See the full description on the dataset page:
https://huggingface.co/datasets/mort666/cv_corpus_v22.