License-audited Korean/English pretraining corpus for the Sage Korean LLM project.
Records: id, text, source, license, meta (JSON of original fields).
⚠️ Mixed licenses — license: other. Comply with each source's license individually. CC-BY / ODC-BY require attribution; CC-BY-SA carries ShareAlike.
fineweb2_ko
46,470,574
ODC-BY-1.0
FineWeb-2 Korean (HuggingFaceFW/fineweb-2 kor_Hang)… See the full description on the dataset page:
https://huggingface.co/datasets/seongchaeae/sage-pretrain-corpus.