This dataset is licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
The Dolci Think 7B DPO mixture was used to preference tune Olmo 3 Think 7B. It contains 150,000 preference pairs created with the preference heuristic described in Delta Learning (Geng et al. 2025).
@misc{olmo2025olmo3,
title={Olmo 3},
author={Team Olmo and Allyson Ettinger and Amanda Bertsch… See the full description on the dataset page:
https://huggingface.co/datasets/leideng/Dolci-Think-DPO-7B-4K-Plus.