This release contains 90,253 Shona speech clips, totalling 441.585 hours. Each clip keeps its original FLAC audio and a Sunbird Whisper pseudo-transcript. These are model outputs, not human reference transcriptions.
The source is the unlabeled Shona ASR split from WAXAL NLP, preserved in the operational checkpoint manassehzw/sna-waxal-annotated-unlabeled. The source checkpoint has no transcripts. This… See the full description on the dataset page:
https://huggingface.co/datasets/manassehzw/shona-waxal-pseudo-labeled.