Discrete speech tokens for Moroccan Darija. 101 hours of transcribed speech,
encoded to a single-codebook neural audio codec and paired with text, ready to
train a text-to-speech model that predicts tokens directly.
Same speech, ~78× smaller. At the codec level that is 256 kbps of PCM
reduced to 0.8 kbps — a 320× reduction in bits — since in this case, one second of… See the full description on the dataset page:
https://huggingface.co/datasets/Tilas/MoulSot-Tokens-v1.