Quick preference dataset I put together with MS3.2 as the response generator, Polar-7B as the RM, and V3.2 (nonthinking) for reference responses.
Contains duplicate responses and responses cut off due the 2048 token response limit.
For a ready to train version, use ConicCat/Wildchat-IF-Preference-MS3.2
All responses were generated with .8 temp and all references with 1.0 temp + .95 top_p.