epochs: 3
learning rate: 3e-4
lora rank: 8
lora alpha: 16
lora dropout: 0.05 for cutoff 1024 13B, otherwise no dropout due to gradient checkpointing
masking: none
mbatch size: 4 (1 for 30B)
batch size: 8 (2 for 30B)
val set size: 0.2
sdp implementation: xformers
optimizer: AdamW
eval strategy: none
Code snippets… See the full description on the dataset page:
https://huggingface.co/datasets/kaiokendev/SuperCOT-dataset.