A cleaned, pre-tokenized version of laion/Emolia
prepared for text-to-speech (TTS) training.
The pipeline is two steps:
Quality filtering with the open-source
audio_filter tool — this
removes the dirtiest recordings (noise, clipping, band-limiting, robotic artifacts,
overlapping speakers), which matters a lot for TTS quality.
Discrete audio tokenization with NVIDIA
nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps
(an FSQ neural audio… See the full description on the dataset page:
https://huggingface.co/datasets/chukypedro/english_data.