KapCode-1B is a high-quality, 1-billion-token curated dataset designed for Continued Pre-Training (CPT) and domain adaptation of compact Large Language Models. Engineered specifically to empower models under 1 billion parameters with robust code generation, technical comprehension, mathematical reasoning, and Fill-in-the-Middle (FIM) infilling capabilities, KapCode-1B combines multi-lingual code… See the full description on the dataset page:
https://huggingface.co/datasets/kaptaan45/KapCode-1B.