This dataset run on dataset after filtering in Stage 1.
This is a synthetic dataset generated using state-of-the-art (SOTA) models (<32B) on our medical QA dataset. Each question has three generated answers.
We used Sglang to generate the data, leveraging code from the open-r1 project by Hugging Face.
⚠ Important Note: This dataset is unfiltered and may contain problematic content. Please apply filtering before use.
Usage… See the full description on the dataset page: https://huggingface.co/datasets/OpenMedical/medical-data-stage1.