This is a text corpus of Meadow Mari
and Hill Mari languages.
The dataset is plain text without any additional annotations except for basic meta data.
Please note, that the texts have been extracted from PDFs, so they are quite dirty.
Source
The corpus is based on texts from Mari Lab PDFs
of magazines and school books.
Composition
The first part of the corpus includes the following magazines: