An rationale-enhanced version of the paired preference learning dataset Intel-ORCA-DPO.
These rationales are more specific, detailed explaning why the chosen response is preferred over the rejected response.
The dataset was generated according to this paper: Data-Centric Human Preference Optimization with
Rationales.