This is the dataset for Per-Training GRAM.
Each item of the dataset includes following keys:
instruction: any prompt in following template:[User Question]
{your prompt here}
input: the input for above prompt, can be empty if there is not.
output: two responses in following template:[The Start of Assistant A's Answer]
{answer of assistant A}
[The End of Assistant A's Answer]
[The Start of Assistant B's Answer]
{answer of assistant B}
[The End of Assistant B's Answer]