This dataset is focused on improving LLM logical reasoning skills and was used to train the Platypus2 models. It is comprised of the following datasets, which were filtered using keyword search and then Sentence Transformers to remove questions with a similarity above 80%:
ScienceQA
Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International
TheoremQA
MIT… See the full description on the dataset page:
https://huggingface.co/datasets/botp/Open-Platypus.