This dataset contains the LeRobot-format BRIDGE-v2 dataset with paths and masks from the PEEK VLM drawn onto the image: PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies.
PEEK fine-tunes Vision-Language Models (VLMs) to predict a unified point-based intermediate representation for robot manipulation. This representation consists of:
End-effector paths: specifying what actions to take.… See the full description on the dataset page:
https://huggingface.co/datasets/jesbu1/bridge_v2_lerobot_pathmask.