This dataset contains 50 Cantonese daily sentences for Text-to-Speech (TTS) model training. The audio totals approximately 4 minutes of clean, segmented recordings. The sentences feature authentic colloquial expressions, longer sentence structures, and are accompanied by accurate Jyutping romanization in the metadata for pronunciation guidance.
Collection Process: The Cantonese sentences were recorded in a quiet environment using a USB condenser microphone. All… See the full description on the dataset page:
https://huggingface.co/datasets/eduhk-compling/11557577_PoonYanKei.