This is a roleplaying dataset of 4000 chosen/rejected pairs. It may be regenerated in the future with a better teacher model, which would fix some of the issues this dataset currently has.
To distill pairs, we go through the following process:
Pick a character card and a question.
Give the LLM the character card and ask it to write an analysis on how that character would reply to the given question.