Views
No views yet
commandr-35b-sft is a supervised fine-tuned variant of Cohere’s Command-R 35B model.axolotl_deduplicated_synthetic_qa.jsonlalpaca_chat.load_qa schema.| Parameter | Value |
|---|---|
| Sequence length | 2048 |
| Micro batch size | 1 |
| Gradient accumulation | 2 |
| Epochs | 1 |
| Learning rate | 0.0001 |
| LR scheduler | cosine |
| Optimizer | AdamW (8-bit) |
| Warmup steps | 20 |
| Weight decay | 0.0 |
| LoRA rank (r) | 16 |
| LoRA alpha | 32 |
| LoRA dropout | 0.05 |
| LoRA target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Gradient checkpointing | ✅ |
| Flash attention | ✅ |
| Auto resume | ✅ |
| Loss watchdog threshold | 8.0 |
| Loss watchdog patience | 20 |
AutoTokenizer<|end_of_text|> as pad_token