SauatAI is a grammar-focused dataset built from 170 children’s stories scraped from ertegiler.kz on July 5, 2025. The dataset was designed to support Kazakh language grammar correction, error detection, and text augmentation research.
s170 — 170 unique stories were scraped and sentence-tokenized.
len60 — Only sentences with ≤60 characters were retained.
n6 — Each correct sentence has 5… See the full description on the dataset page:
https://huggingface.co/datasets/alphazhan/sauatai-ertegiler-kz-misspellings-kk-s170-len60-n6-m1-2-v1.