A large-scale cleaned Indonesian text corpus, built from multiple open sources and processed through a multi-stage quality filtering and deduplication pipeline.
The raw merged corpus of 19,814,163… See the full description on the dataset page:
https://huggingface.co/datasets/AiRukua/IndoCleanCorpus.