-
Method Continued pretraining
-
Epochs 3
-
Batch size 4
-
Grad Accum 2
-
Learning rate 0.00005
-
Embedding Learning Rate 0.000005
-
Optimizer AdamW 8-bit
-
Context length 2048
-
Warmup steps 4
-
Packing True
-
weight decay 0.001
-
LR scheduler Cosine
-
seed 3407
-
LoRA Rank 32
-
Alpha 64
-
Dropout 0.0
-
Variant RS-lora
-
Dataset 2-Pass CPT Corpus
This qwen3_5 model was trained 2x faster with
Unsloth and Huggingface's TRL library.