A high-quality, denoised corpus derived from the Kyrgyz Wikipedia. This dataset is specifically prepared for pre-training and fine-tuning Large Language Models (LLMs) in the Kyrgyz language.
📊 Dataset Benchmark
General Statistics
Metric
Value
Total Articles
76,519
Total Characters
92,945,532
Total Words
11,340,634
Avg. Words per Article
148.21
Language Purity Index
91.71% (Cyrillic)
Character Length… See the full description on the dataset page: https://huggingface.co/datasets/Zhantas/Cleaned-Kyrgyz_Wikipedia.