Real-robot rollouts paired with OSCAR world-model rollouts, from the
policy-evaluation experiment in our paper. We deploy OSCAR as a policy
evaluator on the RoboArena leaderboard:
for each session we autoregressively roll out OSCAR from the recorded first
frame following the policy's action condition, then prompt GPT-5 to judge
task success. OSCAR rollouts correlate strongly with real-robot rollout
outcomes, so they can be used to evaluate whether a… See the full description on the dataset page:
https://huggingface.co/datasets/zywu2115/OSCAR_policy_rollout.