Views
No views yet
| Benchmark | Base | SFT | SFT+RL (ckpt-2600) | Δ over Base |
|---|---|---|---|---|
| GSM8K | 0.2161 | 0.3616 | 0.3624 | +0.1463 |
| MATH-500 | 0.1740 | 0.1500 | 0.1640 | -0.0100 |
| TheoremQA | 0.1000 | 0.1088 | 0.1138 | +0.0138 |
| Benchmark | Base | SFT | SFT+RL (ckpt-2600) | Δ over Base |
|---|---|---|---|---|
| ARC-Easy | 0.4533 | 0.4933 | 0.4882 | +0.0349 |
| ARC-Challenge | 0.3131 | 0.3592 | 0.3592 | +0.0461 |
| MMLU | 0.2514 | 0.4185 | 0.4259 | +0.1745 |
| TruthfulQA | 0.2950 | 0.2681 | 0.2644 | -0.0306 |
| Dataset | Target Size | Role |
|---|---|---|
openai/gsm8k (train) | ~7.5K | Foundation arithmetic and word-problem reasoning |
AI-MO/NuminaMath-CoT | ~8K | Competition-math coverage for MATH-500-style problems |
TIGER-Lab/MathInstruct (CoT-only) | ~5K | Diverse math reasoning and theorem-style supervision |
hendrycks/competition_math (train, L3-L5) | ~3K | Higher-difficulty competition math |
TIGER-Lab/TheoremQA-aligned slice | ~1.5K | Basic theorem-application exposure |