VieMix (
https://arxiv.org/abs/2512.18834) is a Vietnamese pretraining corpus built by combining six publicly available Vietnamese datasets, applying Vietnamese-specific quality filtering, and performing cross-dataset deduplication.
The matched subset uses… See the full description on the dataset page:
https://huggingface.co/datasets/AdaMLLab/VieMix.