A large-scale, multi-source Persian speech dataset curated for ASR and spoken language research
Qoqnus (ققنوس — the Persian Phoenix) is a consolidated, production-grade Persian speech corpus assembled and released by GinkgoQ. It unifies 16 independent datasets spanning read speech, conversational audio, podcast recordings, TTS synthesis, and crowd-sourced contributions — forming one of the largest open Persian ASR corpora… See the full description on the dataset page:
https://huggingface.co/datasets/GinkgoQ/Qoqnus.