TurMix (
https://arxiv.org/abs/2512.18834) is a Turkish pretraining corpus containing 168 billion tokens across 219 million documents (in the minhash subset). Rather than scraping the web again, TurMix combines five publicly available Turkish datasets, applies Turkish-specific quality filtering, and performs cross-dataset deduplication.
We train a 1.4B parameter language model through nanotron on 30 billion tokens to show that the matched subset of TurMix outperforms the… See the full description on the dataset page:
https://huggingface.co/datasets/Alptekinege/TurMix.