Cyrillic dataset of 8 Turkic languages spoken in Russia and former USSR
Dataset Description
The dataset is a part of the [Leipzig Corpora (Wiki) Collection]: https://corpora.uni-leipzig.de/
For the text-classification comparison, Russian has been included to the dataset.
Paper:
Dirk Goldhahn, Thomas Eckart and Uwe Quasthoff (2012): Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages. In: Proceedings of the Eighth… See the full description on the dataset page: https://huggingface.co/datasets/tatiana-merz/cyrillic_turkic_langs.