A dataset of Kazakh-language texts derived from CulturaY — a large multilingual corpus built from Internet Archive data. The kk subset went through an additional multi-stage cleaning and quality-filtering pipeline.
Please note that this dataset is the first version and may contain mixed Russian-Kazakh text, minor spam, and adult content. This dataset is already ready for SLM / LLM training; however, we are currently preparing a large Kazakh corpus that will be… See the full description on the dataset page:
https://huggingface.co/datasets/salyamq/culturay-kk-v1.