Verse-aligned text for 693 African languages, plus several world languages,
for building monolingual and parallel corpora. Every language is aligned
on a shared verse key, so any single language can be pulled on its own or any
two joined into a parallel corpus:
Monolingual corpus for any single language
African ↔ English (English is the default pair)
African ↔ African (e.g. Twi ↔ Yoruba, Hausa ↔ Amharic)
African ↔ other language (French, Arabic, Chinese… See the full description on the dataset page:
https://huggingface.co/datasets/AfriSpeech/africa-corpus.