An open-source collection of 1,093 fully diacritized Arabic speech recordings, crowd-sourced from native speakers via Nahw.ai.
Dataset summary
Stat
Value
Total recordings
1,093
Speakers
10
Language
Arabic (ar)
Sampling rate
16 kHz
License
CC-BY-4.0
Features
audio: The speech recording, resampled to 16 kHz.
transcription: The fully diacritized Arabic sentence that was read aloud.
sentence: The same… See the full description on the dataset page: https://huggingface.co/datasets/NahwAI/arabic-tashkeel-speech.