Views
No views yet
round15 SFT checkpoint. Trained with
RLinf (FSDP2, 4×A100), grpo_norm_by_std=false
(Dr.GRPO), noise_level=0.6, lr 2e-5, clip 1.0, 256 traj/step, 10 steps.dcp_checkpoint/ is a Torch Distributed Checkpoint (model + optimizer), resumable via
RLinf runner.resume_dir.| Suite | tasks | base SR | RL SR | Δ |
|---|---|---|---|---|
| Spatial | 10 | 0.962 | 0.972 | +0.010 |
| Object | 10 | 0.972 | 0.912 | −0.060 |
| Goal | 10 | 0.956 | 0.962 | +0.006 |
| Long (libero_10, trained) | 10 | 0.852 | 0.906 | +0.054 |
| Short (libero_90) | 90 | 0.914 | 0.850 | −0.065 |
| LIBERO-130 (task-weighted) | 130 | 0.921 | 0.877 | −0.044 |