We encountered two major loss spikes while
training K2.
We are releasing these checkpoints so others can study this interesting phenomena in large model training.
Loss spikes are still a relatively unknown phenomena. By making these spikes and associated training details available, we hope others use these artifacts to further the worlds knowledge on this topic.
View all the evaluations on our
Weights & Biases here
The LLM360 Research Suite is a comprehensive set of large language model (LLM) artifacts from Amber, CrystalCoder, and K2 for academic and industry researchers to explore LLM training dynamics. Additional resources can be found at llm360.ai.
1@misc{
2 title={LLM360-K2-65B: Scaling Up Open and Transparent Language Models},
3 author={The LLM360 Team},
4 year={2024},
5}