This dataset is derived from the GSM8K training set questions. The process to create this dataset involved the following steps:
Initial Prompting: Each question from the GSM8K train set was initially answered by the Fine-Tuned Mistral model.
Filtering Incorrect Answers: Incorrect responses were filtered out.
Refinement: The model was prompted to refine its answers based on the incorrect responses.
Final Filtering: The refined responses were… See the full description on the dataset page: https://huggingface.co/datasets/August4293/gsm8k_preference_dataset_it_2.