This dataset was curated for the Sarvam AI screening assignment. It contains clean,
single-speaker YouTube-sourced sentence audio with manually reviewed transcripts and
emotion, volume, energy, pace, and style tags.
audio: relative path to the audio file in the repository (e.g. audio/filename.wav)
text: manually… See the full description on the dataset page:
https://huggingface.co/datasets/akshat1303/indian-english-hindi-tts-60m.