Views
No views yet
OpenGVLab/InternVL3_5-2B-HF with Vanilla GRPO against mmr1's ground-truth solutions.bs 1 x grad_accum 8 x 8 GPUs), lr 1e-6, num_generations 8,
max_completion_length 1024, seed 42. --max_steps was never set; the step
count falls out of the dataset size (5,782 prompts / 8 prompts per step).best is step 440; end is step 722. Read them together: this
run peaks early and then declines, so best alone would misrepresent it.| step | MathVista-150 |
|---|---|
| 0 | 0.2763 |
| 20 | 0.2961 |
| 40 | 0.4342 |
| 60 | 0.3421 |
| 80 | 0.3947 |
| 100 | 0.4079 |
| 120 | 0.3750 |
| 140 | 0.4408 |
| 160 | 0.5132 |
| 180 | 0.4342 |
| 200 | 0.4013 |
| 220 | 0.4671 |
| 240 | 0.4539 |
| 260 | 0.4145 |
| 280 | 0.4211 |
| 300 | 0.4803 |
| 320 | 0.4934 |
| 340 | 0.4474 |
| 360 | 0.4737 |
| 380 | 0.4671 |
| 400 | 0.4474 |
| 420 | 0.5263 |
| 440 | 0.4408 |
| 460 | 0.4868 |
| 480 | 0.4539 |
| 500 | 0.4079 |
| 520 | 0.4408 |
| 540 | 0.4934 |
| 560 | 0.4474 |
| 580 | 0.4342 |
| 600 | 0.4211 |
| 620 | 0.5132 |
| 640 | 0.4737 |
| 660 | 0.5000 |
| 680 | 0.4605 |
| 700 | 0.3750 |
train.log — the complete training log this checkpoint came fromeval_curve.csv — the table above, machine-readablerun_config.json — config as the trainer saw itDrStranded/mllm-repro,
examples/openr1_*_{gt,ttrl}.sh with MLLM_PRE_DIR pointed at the
preprocessed mmr1 set.