Views
No views yet
| Parameter | Value |
|---|---|
| Base model | RLinf/RLinf-OpenVLAOFT-LIBERO-130 |
| Fine-tuning | LoRA via GRPO (Group Relative Policy Optimization) |
| Reward model | Robo-Dopamine GRM-3B (cumulative eval mode) |
| Task | LIBERO-10 #9: put the yellow and white mug in the microwave and close it |
| Hardware | 4x A100-80GB (1 for GRM, 3 for FSDP training) |
| Epochs | 20 |
| Parallel envs | 24 |
| Infrastructure | RLinf v0.2.0 |
TRAINING_REPORT.md in the associated dataset repo for full failure analysis.global_step_10/ — Epoch 10, 61.5% success rate (before collapse)global_step_20/ — Epoch 20, 0% success rate (collapsed)1@misc{auryal2026grm-grpo,
2 title={Dense Reward RL with GRM for Robotic Manipulation},
3 author={Auryal},
4 year={2026},
5}