The corpus was constructed from multiple sources to ensure diversity and representation of real-world Ukrainian social discourse. We systematically scraped comments and posts from Ukrainian Telegram channels, collecting content dated between February 2022 and September 2024.
The volume of the scraped documents amounts to 8,064 texts. Also, we integrated two publicly available datasets: TG samples from D. Baida… See the full description on the dataset page:
https://huggingface.co/datasets/YShynkarov/COSMUS.