-
Training steps: 500
-
Learning rate: Automagic (adaptive)
- Learning rate schedule: Automagic optimizer
- Warmup steps: 0 (commented out: 100)
-
Max grad value: 1.0
-
Effective batch size: 4
- Micro-batch size: 4
- Gradient accumulation steps: 1
- Number of GPUs: 1
-
Gradient checkpointing: True (unsloth)
-
Prediction type: logit_normal
-
Optimizer: automagic
-
Trainable parameter precision: Pure BF16
-
Base model precision: Pure BF16
-
Caption dropout probability: 0.0% (not specified)
-
LoRA Rank: 64
-
LoRA Alpha: 64 (auto-set to rank)
-
LoRA Dropout: 0.0
-
LoRA initialisation style: default
-
LoRA mode: Standard