Dataset Card for nphearum/khmer-raw-text-3M
Dataset Summary
nphearum/khmer-raw-text-3M is a large-scale raw text corpus containing approximately 50000 completed records with 3 million text segments in Khmer, curated for large language model (LLM) pre-training, continued pre-training, and domain adaptation.
The dataset emphasizes Khmer-language coverage, a historically underrepresented low-resource language, while retaining bilingual context for cross-lingual learning.… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/khmer-raw-text-3M.