Method Lora(16 bit)
Epochs 2
Batch size 8
Grad Accum 1
Learning rate 0.0001
Optimizer AdamW 8-bit
Context length 2048
Warmup steps 2
Packing False
weight decay 0.001
seed 3407
This qwen3_5 model was trained 2x faster with
Unsloth and Huggingface's TRL library.