Synthetic therapy conversations using user questions from seed user questions and assistant personas
The dataset is constructed turn by turn from using just 10 (seed user question, assistant persona) pairs
At each turn, the user or the assistant chooses from 5 conversational tactics to respond.
The goal of this dataset is to act as a source dataset for human preference annotations. The preference annotations can be used to create a reward model for empathetic responses in hard conversations.