Dataset Card for language_relatives_2025
Dataset Summary
The dataset is a collection of mono- and multilingual text corpora of many language relatives of standard Estonian, i.e. Finno-Ugric languages and dialects, excluding Finnish and Hungarian. Multilingual corpora include other languages as translation equivalents, among them also Estonian, Finnish and Hungarian.
The aim is to provide data for language technology, first and foremost for machine translation.… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/smugri4-data.