This repository contains approximately 259 billion tokens of pretrain data generated using Qwen2.5-14B-Instruct, Qwen 2.5-32B-Instruct and Qwen 2.5-Coder-32B.
The dataset utilizes a MGA-style methodology and a persona-driven data synthesis methodology
to create diverse and comprehensive training data mainly from the Yulan, the Stack-V2 and the Pile datasets.
The dataset is available under the Apache 2.0 license.
This dataset is mainly in English.
The… See the full description on the dataset page:
https://huggingface.co/datasets/Tiiny/PowerCoding.