Read speech corpus for Amharic (አማርኛ) automatic speech recognition. Converted from the original ALFFA project and restructured to match the google/waxalnlp schema for interoperability.
This is a restructured version of hadamard-2/alffa-amharic.
Changes from v1
utterance_id renamed to id
transcript renamed to transcription
speaker_id set to "unknown" — the original ALFFA Kaldi files shipped with utt2spk mapping each utterance to itself… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/alffa-amharic-v2.