Direct Preference Optimization (DPO) dataset pairing original DeepSeek-R1
distillation responses against synth-style reasoning rewrites produced by
DeepSeek V4 Flash.
The rejected side originates from
Shekswess/trlm-dpo-stage-3-final-2.
Each original assistant response followed the DeepSeek-R1-Distill style
...\n\n
layout. For each record the
block was stripped of its tags to… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/trlm-dpo-stage-3-synth.