Use for whatever you want.
Made to replicate the thought traces of OpenAI's o1, I'll release RL datasets including DPO soon enough.
For fine-tuning smaller models such as Google's google/gemma-2-2b-it with this dataset, I recommend fine-tuning for 2-3 epochs, the loss will be at around 1.6 at the beginning, and 1.3 by the end of the training job with learning rate of 2e-6.
Suggested system prompt:
Always respond in strict JSON format with a reasoning_steps array and a response field. Each… See the full description on the dataset page:
https://huggingface.co/datasets/minchyeom/Thinker-JSON.