This dataset contains the final DPO training pairs used to train
ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO.
The pairs are derived from 7B-native rollouts on OmniGAIA train questions:
Roll out the SFT model on answer-hidden train inputs.
Audit each rollout with Gemini using the private reference answer and annotated solution.
Locate the first erroneous assistant sub-step.
Convert the corrected prefix (tau_win) and the original erroneous prefix… See the full description on the dataset page: https://huggingface.co/datasets/ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO-dataset_for_DPO.