A pairwise preference dataset for DPO fine-tuning of paper-recommendation rerankers, generated from team git merge history via the IPD (Implicit Preference Distillation) framework. First instantiation of the broader IPD method (paper: Code Is Context: Tuning Coding Agents via Implicit Preference Distillation).
4,418 pairwise preference records across two Remyx projects:
VQASynth: 1,418 pairs… See the full description on the dataset page:
https://huggingface.co/datasets/remyxai/mhpd-dpo-v0.