This dataset consists of approximately 56 hours of Afrikaans speech extracted from church sermons, paired with cleaned and aligned transcripts. It is specifically prepared for fine-tuning multilingual ASR models like OpenAI's Whisper (particularly large-v3) on low-resource Afrikaans speech
The audio is segmented into fixed 30-second chunks (with 3-second overlaps for context… See the full description on the dataset page:
https://huggingface.co/datasets/andreoosthuizen/afrikaans-30s.