Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
The OLMo v2 SFT mixture was used to train the OLMo models.
It contains 939,344 samples from the following sets:
CoCoNot (ODC-BY-1.0), 10,983 prompts (Brahman et al., 2024)
FLAN v2 via ai2-adapt-dev/flan_v2_converted, 89,982 prompts (Longpre et al., 2023)
No Robots (CC-BY-NC-4.0), 9,500… See the full description on the dataset page:
https://huggingface.co/datasets/Uzbekswe/tulu-3-sft-olmo-2-mixture.