Context Pattern Inference
The Copain dataset is intended to quickly benchmark the in-context capabilities of a large languaeg model (LLM) in a language-agnostic manner.
Paper: Emergent Abilities of Large Language Models under Continued Pretraining for Language Adaptation
Code: base code for pretraining
Paper Abstract
Continued pretraining (CPT) is a popular approach to adapt existing large language models (LLMs) to new languages. When doing so, it is common practice to include a portion of… See the full description on the dataset page:
https://huggingface.co/datasets/ahmedselhady/CoPain.