This dataset contains paired audio and text data for training and evaluating speech-to-text models in Swahili. The audio files have been processed to remove silence, converted to 44.1kHz mono FLAC format, and are paired with corresponding transcriptions.
audio_*.flac: Audio files in FLAC format, named by their corresponding text corpus ID.
metadata.jsonl: JSON Lines file with metadata for each audio-text pair. Each line is a JSON… See the full description on the dataset page:
https://huggingface.co/datasets/stem-content-ai-project/swahili-speech.