This qwen3_5 model was trained 2x faster with
Unsloth and Huggingface's TRL library.
Method Lora(16 bit)
Epochs 0
Batch size 8
Grad Accum 2
Learning rate 0.0002
Optimizer AdamW 8-bit
Max steps 100
Context length 2048
Warmup steps 4
Packing False
weight decay 0.001
seed 3407