This dataset is prepared for Text-to-Speech (TTS) model training, containing 60 English sentences (covering daily conversations, common phrases, and simple declarative sentences) and corresponding segmented audio clips. The audio was recorded in a quiet language lab using a standard microphone, with a total duration of 3 minutes and 36 seconds. All audio files are in WAV format (16-bit PCM, 44100 Hz), named sequentially as s001.wav to s060. wav. A metadata.csv… See the full description on the dataset page:
https://huggingface.co/datasets/Iris2/11535335_RENZIMENGIris.