This dataset is a merged and cleaned version of openai/gsm8k and Thomas-X-Yang/gsm8k-prolog, introduced in the paper Training Language Models to Use Prolog as a Tool.
Code: aisilab/Prolog-as-a-Tool
During preprocessing, 15 errors were identified and manually corrected across the original datasets: 14 in openai/gsm8k and 1 in Thomas-X-Yang/gsm8k-prolog. This ensured accurate alignment between the natural language questions, numeric answers, and symbolic Prolog… See the full description on the dataset page:
https://huggingface.co/datasets/niklasm222/gsm8k-prolog-prover.