A cleaned and tokenized version of the English data from Mozilla Common Voice 11 dataset.
Cleaning steps:
Filtered on samples with >2 upvotes and <1 downvotes]
Removed non voice audio at start and end through pytorch VAD
Tokenization:
Audio tokenized through EnCodec by Meta
Using 24khz pre-trained model, and target bandwidth of 1.5
Represented in text as audio_token_0 - audio_token_1023