Views
No views yet
unsloth/Qwen2.5-3B-Instruct-bnb-4bit, then this distinct adapter was trained on a deterministic
2,000-pair slice of argilla/ultrafeedback-binarized-preferences-cleaned.| Setting | Value |
|---|---|
| LoRA rank / alpha | 16 / 16 |
| DPO beta | 0.1 |
| Learning rate | 5e-7 |
| Epochs / optimizer steps | 1 / 250 |
| Max sequence / prompt length | 512 / 256 |
| Micro-batch / accumulation | 1 / 8 |
| Train loss | 0.6761 |
| Runtime including reference pass | 2,777.4 s |
| Peak allocated VRAM | 4.932 GB |
training_metrics.json for all 51 logged points.