QASR Dataset: QASR is the largest transcribed Arabic speech corpus, collected from the broadcast domain. It contains
2,000 hours of multi-dialect speech sampled at
16kHz from
Al Jazeera News Channel, with lightly supervised transcriptions aligned with the audio segments. Unlike previous datasets, QASR includes
linguistically motivated segmentation, punctuation, speaker information, and more. The dataset is suitable for
ASR, Arabic dialect identification, punctuation restoration, speaker identification, and NLP applications. Additionally, a
130M-word language model dataset is available to aid language modeling. Speech recognition models trained on QASR achieve competitive
WER compared to the MGB-2 corpus, and it has been used for downstream tasks like
Named Entity Recognition (NER) and
punctuation restoration.