II-Medical-32B-Preview is the latest advanced large language model developed by Intelligent Internet, specifically designed to enhance AI-driven medical reasoning. As our first 32B-scale model version, it significantly advances the capabilities of medical question answering.
II. Training Methodology
We collected and generated a comprehensive set of reasoning datasets for the medical domain and performed SFT fine-tuning on the Qwen3-32B model.
For the hyperparameter:
Max Length: 16378.
Batch Size: 128.
Learning-Rate: 2e-5.
Number Of Epoch: 4.
III. Evaluation Results
image/png
image/png
We evaluated on 10 medical QA benchmarks including MedMCQA, MedQA, PubMedQA, HealthBench, medical related questions from MMLU-Pro, small QA sets from Lancet and the New England
Journal of Medicine, 4 Options and 5 Options splits from the MedBullets platform and MedXpertQA.
Recommended Sampling Parameters: temperature = 0.6, top_p = 0.9
When using, explicitly request step-by-step reasoning and format the final answer within \boxed{} (e.g., "Please reason step-by-step, and put your final answer within \boxed{}.").
VII. Limitations and Considerations
Dataset may contain inherent biases from source materials
Medical knowledge requires regular updates
Please note that It’s not suitable for medical use.