This repository contains a cleaned Indonesian-language corpus for experiments in continued pretraining, instruction tuning, and evaluation-pipeline development. The active version was rebuilt from the previous JSONL shards and an Indonesian-only cleansed Wikipedia subset from sabilmakbar/indo_wiki.
The final corpus contains 100,001,389 SentencePiece tokens according to the repository's tokenizer.model. It… See the full description on the dataset page:
https://huggingface.co/datasets/sansaks/dataset_indonesia.