The Slovene Web genre identification corpus GINCO 1.0 contains web texts, manually annotated with genre,
from two Slovene web corpora, the slWaC 2.0 corpus, crawled in 2014, and a web corpus, crawled in 2021 in the scope of the MaCoCu project.
The corpus allows for automated genre identification and genre analyses as well as other web corpora research.
This is a subcorpus of suitable texts, containing 1002 texts (478,969 words), manually annotated with 24 genre categories (News/Reporting, Announcement,
Research Article, Instruction, Recipe, Call (such as a Call for Papers), Legal/Regulation, Information/Explanation, Opinionated News, Review,
Opinion/Argumentation, Promotion of a Product, Promotion of Services, Invitation, Promotion, Interview, Forum, Correspondence, Script/Drama,
Prose, Lyrical, FAQ (Frequently Asked Questions), List of Summaries/Excerpts, and Other).
The texts in the suitable subset are annotated with up to three genre categories, where the primary label is the most prevalent,
and secondary and tertiary labels denote presence of additional genre(s). They are encoded in three levels of detail, allowing experiments
with the full set (24 labels), set of 21 labels (labels with less than 5 instances are merged with label Other) and set of 12 labels (similar labels
are merged). Additionally, the corpus contains some metadata about the text (e.g. url, domain, year) and its paragraphs (e.g. near-duplicates and their usefulness for the genre identification).