This dataset includes the training mix used to train the SciTulu models described in our paper SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific Literature.
It contains 35K instances from the SciRIFF dataset (1,000 instances per train task), together with a matching number of instances randomly sampled from the Tulu V2 mix. See the dataset cards of the component datasets for more information.
You can load the dataset like:… See the full description on the dataset page:
https://huggingface.co/datasets/allenai/SciRIFF-train-mix.