Views
No views yet
| Param | Value |
|---|---|
| LoRA rank | 64 |
| LoRA alpha | 128 (ratio 2.0) |
| LoRA dropout | 0.0 (implicit regularization from rank constraint) |
| Learning rate | 5e-5 |
| Scheduler | Cosine with 1000-step warmup |
| Steps | 15000 (~3.1 epochs) |
| Batch size | 4 × 4 grad_accum = 16 effective |
| EMA decay | 0.999 |
| Validation | Every 3000 steps, best checkpoint selected |