PersianAudiobook is a collection of short Persian audiobook speech clips paired with conservatively refined pseudo-transcriptions. The initial release contains 39,454 accepted examples representing 218.564 hours of mono, 16 kHz speech. It is intended for speech-recognition research, audiobook-domain language modeling, speech representation learning, and carefully reviewed text-to-speech research.
Raw data was collected from IranSeda audiobooks.
The… See the full description on the dataset page:
https://huggingface.co/datasets/Pooya-Fallah/PersianAudiobook.