This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples for multiple tasks at once.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This was used to train this model.
This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources.
We are releasing this dataset under the terms of ODC-BY. By using this, you… See the full description on the dataset page:
https://huggingface.co/datasets/hamishivi/rds-sels-multitask-rrmax-top326k.