This dataset is a preprocessed version of SURESHBEEKHANI/medical-reasoning-orpo, formatted for preference tuning tasks like DPO or ORPO.
Data Structure
The dataset contains three columns:
question: A combination of the original instruction and Input fields.
accepted: The preferred response, formatted with thinking process and final answer tags.
rejected: The dispreferred response, also formatted with tags.
Answer… See the full description on the dataset page: https://huggingface.co/datasets/LLMcompe-Team-Watanabe/medical-reasoning-orpo_preprocess.