MozillaSpeak (Derived from Mozilla Common Voice Corpus)
Overview
This dataset is derived from Mozilla's Common Voice (cv-corpus-22.0-delta-2025-06-20).The original data is collected and maintained by the Mozilla Foundation under the CC BY 4.0 license.
This version has been processed to include:
Cleaned and normalized sentence text
Phoneme transcripts generated from the CMU Pronouncing Dictionary
Audio files aligned with each sentence
Purpose and Use… See the full description on the dataset page: https://huggingface.co/datasets/MinhLe999/MozillaSpeak.