This release integrates the entire data sequence utilized in the CrystalCoder training. It encompasses data sequences from the three pre-training stages, combining information from two prior works: the SlimPajama dataset and StarCoder, totaling approximately 1300 billion tokens. These tokens are distributed across three stages, each with distinct weights.
During this initial stage, half of the SlimPajama data is utilized… See the full description on the dataset page:
https://huggingface.co/datasets/IFM/CrystalCoderDatasets.