Object-mask annotations and Embodied Chain-of-Thought (ECoT) reasoning traces for
the 379 demonstrations of the LIBERO-10 (libero_10_image) benchmark.
Generated for the CoT-VLA project. Per-object segmentation masks were produced
with interactive SAM2 point/box prompts and bidirectional video propagation; the CoT
reasoning traces were hand-refined per task and re-timed to each episode's actuator
(gripper + motion) signal.