AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co., Ltd.
The dataset is designed for training multi-speaker Text-to-Speech (TTS) systems. It contains roughly 85 hours of emotion-neutral recordings spoken by 218 native Mandarin Chinese speakers, with a total of 88,035 utterances.
Auxiliary speaker attributes, including gender, age group, and native accents, are… See the full description on the dataset page: https://huggingface.co/datasets/SMIIP-lab/AISHELL-3.