<! -- Provide a longer summary of what this dataset is, -->
About three minutes of English spoken speech, comprising roughly 50 sentences, were captured by one speaker in a calm setting for this dataset. In order to comply with typical text-to-speech dataset standards, the primary recording was first recorded as an M4A file before being converted to 16-bit, 44.1 kHz WAV format. To facilitate simple automated access and use for TTS training, the audio is divided… See the full description on the dataset page:
https://huggingface.co/datasets/eduhk-compling/s1160724_GoQuayAdrieneFernandez.