Views
No views yet
Qwen/Qwen3-4B with GRPO on the chemistry split.val-aux/sciknoweval/reward/mean@16 from 10-step validation logs.| best mean@16 | best step | final mean@16 | final step |
|---|---|---|---|
| 66.58% | 100 | 66.58% | 100 |

| step | mean@16 |
|---|---|
| 10 | 46.49% |
| 20 | 56.40% |
| 30 | 63.48% |
| 40 | 65.95% |
| 50 | 66.04% |
| 60 | 64.76% |
| 70 | 65.00% |
| 80 | 65.12% |
| 90 | 63.42% |
| 100 | 66.58% |
metrics.json: parsed validation summaryeval_mean16.csv: step-level validation curve dataeval_mean16.png: validation curve plotglobal_step_100/actor checkpoint. If best step is earlier than 100, the best validation point is reported for tracking, but the corresponding actor weights may not be retained locally./mnt/mole/SDPO/L2T/checkpoints/datasets/sciknoweval/chemistry/qwen3gen-chemistry-GRPO-Qwen-Qwen3-4B-mbs8-train64-rollout8-lr1e-6-vllm0.8run-20260629_124519-qs487q2tglobal_step_100/actor converted from VERL FSDP shards to Hugging Face format.