Views
No views yet
| Hyperparameter | Value |
|---|---|
| Training steps | 19080 |
| Batch size | 256 |
| Weight decay | 0.1 |
| Context length | 1024 |
| Learning rate | 2e-5 linear decay to 2e-6 |
| Precision | bf16 |
| Optimizer | AdamW |
@article{luhtaru2024err,
title={To Err Is Human, but Llamas Can Learn It Too},
author={Luhtaru, Agnes and Purason, Taido and Vainikko, Martin and Del, Maksym and Fishel, Mark},
journal={arXiv preprint arXiv:2403.05493},
year={2024}
}