Beta
Explore
Marketplace
Neural Labs
Playground
Wallet
Docs
everyayah-phonemes – Dataset by hetchyy | AlphaNeural AI
You can deploy this model and start earning money today!
hetchyy
/
everyayah-phonemes
like
0
automatic-speech-recognition
text-to-speech
ar
cc-by-4.0
100K<n<1M
parquet
audio
text
datasets
dask
mlcroissant
polars
us
Views
No views yet
Model card
Files and Versions
Community
API
Phoneme-labelled Quran Datatset
This dataset contains recitations from 45 professional Quran reciters, sourced from EveryAyah and QUL. The audio has been automatically phoneme-labelled using a custom phonemizer that encodes Tajweed rules.
Dataset Structure
audio: 16 kHz resampled mono audio
duration: length of the audio in seconds
verse: reference in {surah_num}_{ayah_num} format
reciter: name of the Qari'
text: diacritised text in Uthmani script
phonemes:… See the full description on the dataset page:
https://huggingface.co/datasets/hetchyy/everyayah-phonemes
.