Neyshekar is an open Persian speech dataset collected from native Persian speakers through a community-driven crowdsourcing platform. It supports automatic speech recognition, text-to-speech, speech representation learning, and related Persian speech research.
This release contains 40,008 transcribed WAV clips totaling approximately 63 hours. Audio is mono, 16-bit PCM, and sampled at 16 kHz.