The Icelandic Gigaword Corpus (IGC-2022)[1] is a diverse collection of Icelandic texts from 9 individual corpora. Each is available as raw text placed in
tags and as processed text with tokenization, POS tags, and lemmatization.
Dataset Description
This dataset was created for Icelandic text classification from the IGC-2022 corpus. XML files were parsed. The books category, which contains the fewest total tokens, was selected as a benchmark. The news category, having the lowest average… See the full description on the dataset page:
https://huggingface.co/datasets/elenaovv/igc-labeled.