[!NOTE]
For full information, go check out the Tmax paper here.
This is the dataset we used to train Tmax 9b (and our other tmax models), formatted for use with our open-instruct fork here.
In general, this is a collection of roughly 15k RL environment instances.
For details on how we generated this dataset and its makeup, please see our paper!
You can find a more generic version of this⦠See the full description on the dataset page:
https://huggingface.co/datasets/allenai/tmax-15k-open-instruct.