This dataset is licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
The Dolci Instruct DPO mixture was used to preference tune Olmo 3 Instruct 7B. It contains 260,000 preference pairs in total, including:
125,000 pairs created with the preference heuristic described in Delta Learning (Geng et al. 2025)
125,000 pairs created with a delta-aware Ultrafeedback-esque GPT-judge pipeline… See the full description on the dataset page:
https://huggingface.co/datasets/leideng/Dolci-Instruct-DPO-4K-Plus.