Neyshekar is an open, community-driven Persian speech dataset collected via a web-based crowdsourcing platform at ney.shekar.io. The recordings are provided by a combination of volunteer contributors and paid voice actors, all of whom are native Persian speakers. It supports automatic speech recognition, text-to-speech, speech representation learning, and other downstream Persian speech applications.
Audio is mono, 16-bit PCM, sampled at 16 kHz. Each release represents… See the full description on the dataset page:
https://huggingface.co/datasets/shekar-ai/neyshekar-v5-persian-asr-fa.