Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
normalized-datasets-for-koreanLLM – Dataset by mkd-chanwoo | AlphaNeural AI
You can deploy this model and start earning money today!
mkd-chanwoo
/
normalized-datasets-for-koreanLLM
like
0
1M<n<10M
json
text
datasets
dask
polars
mlcroissant
us
Views
No views yet
Model card
Files and Versions
Community
API
Normalized Datasets for Korean LLM (Stage 0.5)
This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains.
Quick Stats
Metric Value
Total Documents ~944,545,628
Source Datasets 42
Format JSONL (one JSON object per line)
Languages English, Korean
Domains English, Korean, Code, Science
Pipeline Stage Stage 0.5 (after download, before… See the full description on the dataset page:
https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM
.