nvidia/BridgeData2-Subset-Synthetic-Captions is a subset of BridgeData V2 packaged with short robot-manipulation video clips and synthetic video captions. It is intended for supervised fine-tuning (SFT), prompt generation, and evaluation workflows involving text-to-video, image-to-video, and video-to-video generation of robot manipulation scenes.
The source data is derived from BridgeData V2, a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/BridgeData2-Subset-Synthetic-Captions.