The LEMMA is collected from MATH and GSM8K. The training set of MATH and GSM8K is used to generate error-corrective reasoning trajectories. For each question in these datasets, the student model (LLaMA3-8B) generates self-generated errors, and the teacher model (GPT-4o) deliberately introduces errors based on the error type distribution of the student model. Then, both "Fix & Continue" and "Fresh & Restart" correction strategies are applied to these… See the full description on the dataset page:
https://huggingface.co/datasets/panzs19/LEMMA.