*The IWSLT shared task submission details and the test set are now available at IWSLT 2026 *
This dataset is a large-scale multilingual speech corpus curated for speech-to-speech translation, speech-to-text, and multilingual speech processing research.
The data is organized by language, speaker (user_id), and dataset split (train, dev), and includes rich acoustic and metadata… See the full description on the dataset page:
https://huggingface.co/datasets/McGill-NLP/NaijaS2ST.