This subset is derived from the meta-math/MetaMathQA dataset, which contains 395,000 samples. The MetaMathQA dataset augments samples from the training sets of GSM8K and MATH. For this subset, we selected only the 240,000 samples that were augmented from GSM8K.
@article{yu2023metamath,
title={MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models},
author={Yu, Longhui and Jiang, Weisen and Shi, Han and Yu, Jincheng and Liu… See the full description on the dataset page:
https://huggingface.co/datasets/fxmeng/MetaMath-GSM240K.