This is grimulkan/theory-of-mind with "rejected" responses generated using mistralai/Mistral-7B-Instruct-v0.2, and the file formatted for use in DPO training.
The code used to generate the dataset can be found in this repository:
https://github.com/DocShotgun/LLM-datagen