Dense-reward-annotated rollout dataset from GRPO training of OpenVLA-OFT on LIBERO-10 Task 9 ("put the yellow and white mug in the microwave and close it").
Dataset
Stat
Value
Episodes
~2,100 (T >= 5 steps)
Format
LeRobot (parquet + images)
Task
put the yellow and white mug in the microwave and close it