Estonian National Corpus 2021: Morphologically Tagged and Clean Text
Dataset Summary
This repository contains two versions of the Estonian National Corpus 2021, representing around 43GB of data and millions of sentences annotated with morphological features. This corpus is designed to be useful for:
Morphological analysis.
Natural language understanding for Estonian.