This dataset is provided to facilitate access to GSM8k-Aug, originally from
https://github.com/da03/Internalize_CoT_Step_by_Step and
https://arxiv.org/pdf/2405.14838.
This dataset is used to train CODI (
https://arxiv.org/abs/2502.21074)
Description:
We utilize two datasets to train our models--GSM8k-Aug and GSM8k-Aug-NL. (1) We use the GSM8k-Aug dataset, which has proven effective for training implicit CoT methods. This dataset extends the original GSM8k training set to 385k samples by… See the full description on the dataset page:
https://huggingface.co/datasets/zen-E/GSM8k-Aug.