This public, manually gated repository archives Qwen3-TTS 12 Hz, 16-codebook
training records for the Instinct TTS preparation pipeline. It stores tokenized
JSONL shards, immutable source-repository mappings, per-file hashes, audit
reports, and filterable dataset/split provenance.
The current completed Instinct tranche contains 5,069,906 rows in 635 shards.
Audio is not duplicated here: reference_audio_ref and target_audio_ref
resolve through… See the full description on the dataset page:
https://huggingface.co/datasets/instinct1912/qwen3-tts-12hz-tokenized-v1.