KoPI-CC (Korpus Perayapan Indonesia)-CC is Indonesian only extract from Common Crawl snapshots using ungoliant, each snapshot also filtered using some some deduplicate technique such as exact hash(md5) dedup technique and minhash LSH neardup
Each folder name inside snapshots folder denoted preprocessing technique that has been applied .
Raw
this processed directly from cc snapshot using ungoliant without any addition filter ,you can read it… See the full description on the dataset page:
https://huggingface.co/datasets/acul3/KoPI-CC.