Pre-extracted HiggsAudioV2 audio tokens for training OmniVoice Amharic TTS.
Downloading these tokens (~648 MB, ~2 min) skips the 30+ minute token extraction step (steps 1-3 of the training pipeline).
Samples per shard
500 (last shard: 231)… See the full description on the dataset page:
https://huggingface.co/datasets/demeleww/Amharic_tokens.