A cleaned and structured collection of 228,810 articles from the Kazakh Wikipedia, curated for Large Language Model (LLM) pre-training, fine-tuning, and Natural Language Processing (NLP) tasks.
📊 Dataset Summary
Property
Value
Total Articles
228,810
Source
Wikimedia Foundation (kkwiki)
Language
Kazakh (Cyrillic script)
Format
JSON Lines (.jsonl)
Article Scope
Namespace 0 (Main Articles only)
🔧 Preprocessing… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Cleaned-Kazakh-Wikipedia.