Pre-tokenized speech dataset using DAC at 16kHz with 2 codebooks. Optimized for speech TTS training — 16kHz captures the full speech frequency range without wasting capacity on inaudible frequencies.
Speech lives below 8kHz — 16kHz sample rate is sufficient (Nyquist)
50 tokens/sec per codebook vs 87 at 44kHz — shorter sequences, faster training
2 codebooks at 16kHz produce intelligible speech — verified by listening tests… See the full description on the dataset page:
https://huggingface.co/datasets/treadon/speech-dac-16khz-2cb.