Pre-training data, checkpoints and benchmark keyframe data for BridgeVLA++ and its
predecessor BridgeVLA.
BridgeVLA++ is a 3D vision-language-action framework that preserves the input-output
alignment of a pre-trained VLM during 3D action learning — point clouds are projected
into multi-view images and intermediate heatmaps are predicted before actions — and
extends it with a unified spatio-temporal memory modeling persistent spatial
context and temporal interaction… See the full description on the dataset page:
https://huggingface.co/datasets/LPY/BridgeVLA.