UberText-GEC is a large corpus of social media texts scraped from Ukrainian Telegram
(UberText 2.0. Dataset; Paper, Chaplynskyi, 2023)
and automatically corrected using the approach TBU.
Structure
uber_text_gec.csv - main data.
language - language of text;
text - original text;
correction - corrected text;
uber_uk_annotations.csv - contains human annotations for 1500 samples.
text - original text;
correction - corrected… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/UberText-GEC.