Korean supervised fine-tuning (SFT) dataset in fully-rendered ChatML format, produced by the
Keural SFT data pipeline (collect → structure → clean → quality-filter → safety-filter → dedup →
format → tokenize → package → audit). Intended for instruction-tuning of the Keural model family.
Language: Korean (ko)
Total samples: 2,018,417 (~498M tokens; see manifest.json for per-shard counts and sha256 checksums)
Format: one JSON object per line (JSONL)… See the full description on the dataset page:
https://huggingface.co/datasets/Mkd-Yonas/keural-SFT-chatml-ko-v2.