This is the training dataset for the DPO Turn strategy in PLAYPEN: An Environment for Exploring Learning From Dialogue Game Feedback.
These preference dataset have been obtained from these games' instances using this script, with --preference_depth turn.
Given the huge number of chosen vs rejected pairs in the first turn of the conversation, we limit the numbers of chosen and rejected pairs for the first turn to 10k samples (--first_turn_limit True).… See the full description on the dataset page:
https://huggingface.co/datasets/clembench-playpen/DPO_turn_bug.