A curated corpus of Ukrainian-language text collected from public social media platforms, designed for language model pretraining and fine-tuning.
This dataset is part of the UkrLM initiative — an open effort to build foundational NLP resources for the Ukrainian language.
Telegram… See the full description on the dataset page:
https://huggingface.co/datasets/ruvimx/UkrLM-social.