This corpus contains texts in 26 languages using the Cyrillic alphabet, designed for the task of automatic language identification.
Languages: 26 individual Cyrillic languages, plus 17 mixed language combinations created through augmentation
Source: texts collected from Wikipedia and augmented using various methods
Size: 20,232 examples in total, including 18,334 original articles collected from Wikipedia and 1,898 texts… See the full description on the dataset page:
https://huggingface.co/datasets/AlmaznayaGroza/cyrillic-language-classification.