This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Contains tokenized speech training shards… See the full description on the dataset page:
https://huggingface.co/datasets/instinct-org/zy_chunked_tokenized.