[!IMPORTANT]
Official dataset for our INTERSPEECH 2026 paper
"A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563).
Part of the Balalaika Russian speech data-processing pipeline — code:
https://github.com/lab260ru/balalaika.
If you use this resource, please cite it.
A curated Russian speech dataset for advanced speech generative tasks.
OpenSTT… See the full description on the dataset page:
https://huggingface.co/datasets/lab260/openstt_balalaika.