Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
sozkz-corpus-clean-v3 – Dataset by stukenov | AlphaNeural AI
You can deploy this model and start earning money today!
stukenov
/
sozkz-corpus-clean-v3
like
0
text-generation
kk
apache-2.0
10M<n<100M
parquet
text
datasets
dask
polars
mlcroissant
us
kazakh
language-modeling
cleaned
deduplicated
corpus
Views
No views yet
Model card
Files and Versions
Community
API
SozKZ Corpus Clean v3 — Cleaned Kazakh Text Corpus
A large-scale cleaned and deduplicated Kazakh text corpus assembled from 18 public sources. Designed for pre-training causal language models on Kazakh text.
Overview
Total texts 13,700,018
Train split ~13,563,018 (99%)
Validation split ~137,000 (1%)
Raw input 28,431,116 texts
Pass rate 48.2%
Dedup removed 3,170,330
License Apache 2.0
Sources
Source Raw Clean Pass… See the full description on the dataset page:
https://huggingface.co/datasets/stukenov/sozkz-corpus-clean-v3
.