This is the training dataset for the DPO Dialogue strategy in PLAYPEN: An Environment for Exploring Learning From Dialogue Game Feedback.
This preference dataset has been obtained from these games' instances using this script, with --preference_depth dialogue.
Dataset Description
Preference Dataset where chosen vs rejected continuations are the successful vs unsuccessful interactions stored.
Language(s) (NLP): English
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/clembench-playpen/DPO_dialogue.