A 150-video benchmark in which a vision-language model watches each clip and summarizes it as an
ordered, timestamp-free sequence of structured skill calls — named actions with typed,
possibly list-valued arguments — under one unified, cross-domain skill library.
add(object=["white granular ingredient"], destination="green mixing bowl")
shape(object="dough", result="round arepa patties")
grasp(object="blue cylindrical container"… See the full description on the dataset page:
https://huggingface.co/datasets/Chenwei1999/skill-call-benchmark-50x3.