These datasets are exactly like the Evaluation datasets except the model_responses array are budget forcing rounds.
So the first response is at a maximum total context length of 4k, the second response (2nd index in the array) is a continuation of that last response up to a total of 8,192 tokens.
Column
Description
question
The question we want the model to answer
answer
The string answer
task
The name of the task the row belongs to