This repository contains approximately 10 billion tokens of pretrain data generated using Qwen2.5-14B-Instruct.
The dataset utilizes a MGA-style methodology
to create diverse and comprehensive training data from the MegaMath dataset. The dataset is available under the Apache 2.0 license.
This dataset is mainly in English.
The dataset inherits the biases, errors, and omissions known to exist in data used for seed sources and models used for data… See the full description on the dataset page:
https://huggingface.co/datasets/Tiiny/PowerMath.