DPO dataset, using a subset of the data from the other magpie datasets, filtered by quality.
5 answers generated per input and ranked using a reward model, with the highest scoring output in the chosen column and the lowest socring output in the rejected column.
Reward model used: https://huggingface.co/RLHFlow/ArmoRM-Llama3-8B-v0.1