This dataset was cultivated from Beijuka/xhosa_parakeet_50hr. This dataset orginally came from NCHLT isiXhosa Speech Corpus (see below).
The original corpus contained audio and transcription in 3-5 word segments. This meant that the majority of the dataset was ~5 seconds long. Whisper can receive an input of 30 seconds. This meant that the dataset required substantial padding. To reduce the amount of padding, the audio segments were merged together sequentially… See the full description on the dataset page:
https://huggingface.co/datasets/ilyes25/wjbmattingly_xhosa_merged_audio.