Llama-3.2-1B Korean Continued Pretraining (KcBERT Corpus)
This model is a continuation of Llama-3.2-1B, pretrained on the KcBERT Pretraining Corpus (Korean News Comments).
- Original Corpus: KcBERT Kaggle Dataset
- Target Language: Korean (news comments, informal, colloquial)
- Perplexity Improvement: 18% (PPL reduced from 7.59 ➔ 6.20)
- Model Type: Llama-3 architecture
- Context Window: 4096 tokens
Fine-tuned for better performance on real-world Korean language generation tasks.