Views
No views yet
mistral-12b-cpt is a continual-pretrained version of the Mistral-12B Nemo Instruct model.| Dataset | Description |
|---|---|
arxiv.jsonl | Scientific and technical papers |
gov.jsonl | Government and policy documents |
news.jsonl | News articles |
wiki.jsonl | Wikipedia text |
| Parameter | Value |
|---|---|
| Sequence length | 2048 |
| Micro batch size | 2 |
| Gradient accumulation | 2 |
| Epochs | 10 |
| Max steps | 10000 |
| Learning rate | 0.0002 |
| LR scheduler | cosine |
| Optimizer | AdamW (8-bit) |
| Warmup steps | 10 |
| Weight decay | 0.0 |
| LoRA rank (r) | 16 |
| LoRA alpha | 32 |
| LoRA dropout | 0.05 |
| LoRA targets | q_proj, k_proj, v_proj, o_proj |
| Gradient checkpointing | ✅ |
| Flash attention | ✅ |
| Loss watchdog (threshold/patience) | 5.0 / 3 |
AutoTokenizer<|end_of_text|>