CORE task as adapted for "Tokenization is Sensitive to Language Variation" paper, see arxiv.
Originally downloaded from
https://github.com/TurkuNLP/CORE-corpus, see also Register identification from the unrestricted open Web using the Corpus of Online Registers of English. If you want to use the original CORE dataset refer to these original sources.
Warning: the column confusingly named "genre" was used as the ground truth label for experiments in the "Tokenization is Sensitive to Language… See the full description on the dataset page:
https://huggingface.co/datasets/AnnaWegmann/CORE.