This dataset aims to be a compilation of simple preference pairs for reasoning tasks generated synthetically using gemini-2.5-flash It has been generated using this colab notebook.
For RL and reasoning tasks.
Use it as a preference dataset for DPO tuning.
Though a detailed schema isn’t provided in its card, it follows a standard Hugging Face DatasetDict. Once… See the full description on the dataset page:
https://huggingface.co/datasets/ritwikraha/reasoning.