A small, carefully curated text-to-speech dataset: ~68 minutes of clean,
single-speaker-per-clip audio in Indian English (en-IN) and Hindi (hi-IN), sourced
from YouTube, with accurate transcriptions and per-clip emotion/style tags.
Built as a data-quality exercise: clips were filtered conservatively and a sample was
verified by listening rather than shipped straight from an automated pipeline.