speechocean762 is a speech dataset of native Mandarin speakers (50% adult, 50% children) speaking English.
It contains phonemic annotations using the sounds supported by ARPABet.
It was developed by Junbo Zhang et al. Read more on their official github, Hugging Face dataset, and paper.
This Processed Version
We have processed the dataset into an easily consumable Hugging Face dataset using this data processing script.
This maps the phoneme annotations… See the full description on the dataset page: https://huggingface.co/datasets/KoelLabs/SpeechOcean.