Preference (chosen / rejected) dataset generated with an Empirical-MCTS (Em-Mcts) rollout pipeline
and scored by a reward model. Every sample contains a higher-quality chosen response and a
lower-quality rejected response for the same prompt, making it suitable for DPO / preference
optimization and reward-model training.
Records: 4,959
Format: JSON Lines (one JSON object per line)
Language: English
Generation model:… See the full description on the dataset page:
https://huggingface.co/datasets/Minami-su/Step-3.5-Flash-Instruct-EmMcts.