This is a gated Russian TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
audios/shard-
.tar: tokenized audio shards
txts/shard-.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json: upload-time… See the full description on the dataset page:
https://huggingface.co/datasets/instinct-org/tbp_chunked_speech_restorised_tts_train_clone_pairs.