This dataset is constructed using Qwen2.5-72B-Instruct and QwQ-32B-Preview, and it serves as the foundation for fine-tuning FineMedLM-o1.
This repository contains the DPO data used during the training process. To access the dataset, you can either use the load_dataset function or directly download the parquet files from the designated folder.
For details, see our paper and GitHub repository.
If you find our data useful, please consider citing our… See the full description on the dataset page:
https://huggingface.co/datasets/hongzhouyu/FineMed-DPO.