pioNER corpus provides gold-standard and automatically generated named-entity datasets for the Armenian language.
Alongside the datasets, we release 50-, 100-, 200-, and 300-dimensional GloVe word embeddings trained on a collection of Armenian texts from Wikipedia, news, blogs, and encyclopedia.
The generated corpus is automatically extracted and annotated using Armenian Wikipedia. We used a modification of… See the full description on the dataset page:
https://huggingface.co/datasets/Karavet/pioNER-Armenian-Named-Entity.