Bambara speech paired with the French source line it renders, a written Bambara translation of
that line, and a machine transcription of the audio. 58,447 rows, 36.66 hours, 48.77 GB of
Parquet.
Load
The config is default and the splits are not_combined and combined — there is no train
split, so a bare load_dataset returns a DatasetDict keyed by those two names.
from datasets import load_dataset