A clean, unlabeled corpus of Ossetian news articles designed for unsupervised learning tasks: clustering, embedding training, language modeling, and semantic search.Contains approximately 12,000 curated records (filtered from a raw pool of 59,375 articles) in the Ossetian language, with metadata including source URL, publication date, media outlet, and author.
🔗 This is the source corpus for the following supervised datasets:• Ossetian-News-Binary — binary… See the full description on the dataset page:
https://huggingface.co/datasets/OssetianNLPWorld/ossetian-news-corpus.