This dataset is a preprocessed version of qihoo360/Light-R1-DPOData, adapted for use with the verl training pipeline. It is designed for DPO (Direct Preference Optimization) training, containing pairs of chosen and rejected responses for mathematical reasoning problems.
The original Light-R1-DPOData is part of the "Light-R1: Surpassing R1-Distill from Scratch with $1000 through… See the full description on the dataset page:
https://huggingface.co/datasets/LLMcompe-Team-Watanabe/math_Light-R1-DPOData_preprocess.