This dataset contains 621,357 English, Finnish, and Swedish social media documents extracted from the HPLT 2.0 web corpus, enriched with web register labels and thematic subregister cluster labels. It was produced as part of Fin-CLARIAH deliverable D3.3.3 ("Machine Learning-Based Enrichment of Social Media") and is intended to support corpus linguistics research and downstream NLP on social media language varieties.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/hplt-social-media-registers.