This dataset is used for training RoboFarseer(
https://arxiv.org/abs/2509.25852), a Vision-Language Model (VLM) based robot task planner. The dataset is converted from human demonstration videos collected using UMI (Universal Manipulation Interface) gripper, and follows the standard Visual Question Answering (VQA) format.
The dataset contains three types of annotations for training different model capabilities:
Plan: Given the current scene… See the full description on the dataset page:
https://huggingface.co/datasets/zitong86/LEAP.