This dataset contains quantized Hindi Text-to-Speech (TTS) samples generated using NVIDIA’s nemo-nano-codec-22khz-0.6kbps-12.5fps neural audio codec.It is designed for training lightweight speech synthesis models, such as token-based TTS models, audio language models, or text-to-codec models.
text
The transcription (Hindi text) corresponding to each audio sample.
speaker
Speaker identity… See the full description on the dataset page:
https://huggingface.co/datasets/ArunKr/tts-quantized-dataset.