IndMix (
https://arxiv.org/abs/2512.18834) is an Indonesian pretraining corpus built by combining six publicly available Indonesian datasets, applying Indonesian-specific quality filtering, and performing cross-dataset deduplication.
matched
Documents appearing in 2+ source datasets
The matched subset uses… See the full description on the dataset page:
https://huggingface.co/datasets/AdaMLLab/IndMix.