A curated, single-speaker TTS training dataset with 144 segments (~60 minutes total) sourced from YouTube, transcribed using Sarvam AI ASR, and annotated with emotion/style tags.
Sample rate: 16 kHz
Channels: Mono
Format: WAV (16-bit PCM)
Segment length: 20–28 seconds… See the full description on the dataset page:
https://huggingface.co/datasets/Itsharshi/indian-tts-dataset.