Dataset Card for Ossetian Web Corpus
Dataset Details
Dataset Description
The Ossetian Web Corpus is a comprehensive collection of Ossetian-language texts gathered from three main sources: online news portals, Wikipedia articles, and digitized books. The corpus is designed to support natural language processing (NLP) research and development for the Ossetian language, a low-resource language spoken in the Caucasus region.