184 on-policy rollouts of a fine-tuned vision-language-action policy on an SO-101 arm, running a coffee-making routine as 10 separately instructed steps — with a success / fail / unstable label on every episode.
Unlike teleoperated demonstration sets, every trajectory here was produced by the policy itself, so the failures are the policy's own. Recorded over 15 consecutive runs of the routine, with retries kept in place: when a step failed, the operator… See the full description on the dataset page:
https://huggingface.co/datasets/nevertmr/so101_coffee_rollouts.