Private, single-speaker Urdu-script corpus prepared for the Ghalib TTS project.
It contains 27,326 clips totaling approximately 62.01 hours.
id, audio, text, speaker_id, source, duration_seconds, language, and
is_synthetic. Audio is embedded in Parquet and decoded by datasets.Audio.
The deterministic 90/5/5 split from the training pipeline is preserved. The test
split is locked and must not be used for… See the full description on the dataset page:
https://huggingface.co/datasets/theusamaaslam/ghalib-urdu-synthetic.