70,587 documents · 36,468,751 words
Built by Deflated — clean Indonesian text corpus for training language models.
Most Indonesian datasets are full of navigation noise, cookie banners, and boilerplate. Cleanesia is filtered and cleaned.
domain
encyclopedia 8000
legal 261
news 732
web 61594
Wikipedia ID — encyclopedia… See the full description on the dataset page:
https://huggingface.co/datasets/ripkiiiii/cleanesia.