Odyssey SA Voice Corpus (V0.4) — Evaluation Preview
Overview
The Odyssey SA Voice Corpus (V0.4) is a 1000-hour multilingual South African speech dataset spanning 8 languages, designed for automatic speech recognition (ASR), code-switching research, and large audio model (LAM) evaluation.
This repository provides a 69-minute evaluation preview (Episode 001 — Nobantu Vilakazi). The preview mirrors the segmentation, metadata schema, validation pipeline, and structural standards used in the full… See the full description on the dataset page:
https://huggingface.co/datasets/ODYSSEYAILABS/odyssey-sa-voice-v0.4-sample.