Beta
Explore
Marketplace
Neural Labs
Playground
Wallet
Docs
ko_llm_annotations – Dataset by devngho | AlphaNeural AI
You can deploy this model and start earning money today!
devngho
/
ko_llm_annotations
like
0
text-classification
machine-generated
machine-generated
HAERAE-HUB/KOREAN-WEBTEXT
blueapple8259/c4-ko-cleaned-2
ko
mit
1M<n<10M
parquet
text
datasets
dask
mlcroissant
polars
us
synthetic
Views
No views yet
Model card
Files and Versions
Community
API
Dataset
이 데이터셋은 fineweb-edu의 방법을 한국어에 적용하기 위해 만들어진 합성 데이터셋입니다. v1과 v2는 퀄리티가 낮습니다. v3을 사용하는 것을 권장합니다. This synthetic dataset was created to apply the methods of fineweb-edu to Korean datasets. v1 and v2 are of low quality. I recommend using v3.
v1
source: HAERAE-HUB/KOREAN-WEBTEXT, sampled 500k analysis model: mistralai/Mistral-Nemo-Instruct-2407
temperature: 0.5 min_p: 0.1 max_model_len: 8192
generation time: ~42 hrs
v2
source:… See the full description on the dataset page:
https://huggingface.co/datasets/devngho/ko_llm_annotations
.