This dataset was collected, cleaned and adjusted for huggingface hub and ready to be used for whisper finetunning/training.
From MGB-3 website:
The MGB-3 is using 16 hours multi-genre data collected from different YouTube channels. The 16 hours have been manually transcribed.
The chosen Arabic dialect for this year is Egyptian.
Given that dialectal Arabic has no orthographic rules, each program has… See the full description on the dataset page: https://huggingface.co/datasets/MightyStudent/Egyptian-ASR-MGB-3.