[Paper] [Model]
This is a multilingual preference dataset generated using human written prompts and responses from 7 LLMs. We evaluate each set of responses 5 times using GPT4.
Note that this model has a non-commerical license as we used the Command R and Command R+ models to create this data.
We are currently working on a developing a commerically usable model, so stay tuned for that!
This is the ORPO training dataset derived from the… See the full description on the dataset page:
https://huggingface.co/datasets/lightblue/mitsu_top25_borda.