Views
No views yet
<think>CoT</think>Response format), using a 123K-sample subset (1/3 of the full 370K dataset) for 3 epochs.ReasonMed.json variant. Each sample wraps chain-of-thought reasoning in <think>...</think> tags followed by the final answer.eval_ll.py in the repo). Paper baseline is ReasonMed-7B trained on the full 370K samples.| Benchmark | Ours (123K) | Paper (370K) | Δ |
|---|---|---|---|
| MedQA | 62.2 | 66.9 | -4.7 |
| MedMCQA (val) | 60.9 | 65.1 | -4.2 |
| PubMedQA | 77.5 | 82.0 | -4.5 |
| MMLU-Anatomy | 75.6 | 75.6 | 0.0 |
| MMLU-Clinical-Knowledge | 78.5 | 79.3 | -0.8 |
| MMLU-College-Biology | 81.9 | 79.2 | +2.7 |
| MMLU-College-Medicine | 70.5 | 73.4 | -2.9 |
| MMLU-Medical-Genetics | 84.0 | 85.0 | -1.0 |
| MMLU-Professional-Medicine | 79.0 | 80.9 | -1.9 |
| Total | 65.8 | 69.6 | -3.8 |