한국어 중심 500M급 decoder-only 모델을 처음부터 학습하기 위한 사전 토큰화
데이터입니다. 각 행의 input_ids는 uint16[2049]이며 앞의
2,048개 토큰을 causal-LM 입력·레이블로 사용합니다.
train: 5,075,616 rows / 10,399,937,184 stored tokens /
10,394,861,568 effective 2048-token training positions
validation: 9,761 rows / 20,000,289 stored tokens
tokenizer: 한국어 SentencePiece Unigram 32K, NFC + identity + byte fallback
원천 데이터: HuggingFaceFW/fineweb-2, kor_Hang
고정 revision:… See the full description on the dataset page:
https://huggingface.co/datasets/Infinity08/KAWK500M-Korean-Pretraining-10B.