The following data mix was used to train K2 and achieve results in line with Llama 2 70B.
K2 was trained on 1.4T tokens across two stages. The data sources and data mix for each stage are listed below.
Dataset Description: Stage 1
uspto (from the Pile)
4.77B
3x… See the full description on the dataset page:
https://huggingface.co/datasets/IFM/K2Datasets.