Produced by one iteration of
vla-agent-loop:
rollouts collected on the physical SO-101 rig, merged with the base teleop dataset, scored
for progress by Robometer-4B (top camera), then finetuned on HPC with RA-BC sample weighting.
Not evaluated: the rig's evaluation dataflow hung when the arm's serial link dropped, and the run was stopped before it could be re-run.
Note that the rig's on-policy success rate is measured over very few episodes; treat any
single evaluation here as an estimate with a wide confidence interval.