This repository contains the Armenian Heritage Dataset, a high-quality, curated dataset designed for Armenian Natural Language Processing (NLP) tasks. It serves as a benchmark and training resource for various downstream applications, including text generation, masked language modeling, and token classification.
The Armenian Heritage Dataset provides a clean, well-structured, and verified collection of Armenian text. It is… See the full description on the dataset page:
https://huggingface.co/datasets/andovirab/armenian_heritage_small_dataset.