This repository contains the datasets associated with the paper MM-ACT: Learn from Multimodal Parallel Generation to Act.
MM-ACT is a unified Vision-Language-Action (VLA) model that integrates text, image, and action in a shared token space and performs generation across all three modalities. This dataset provides crucial data for training and evaluating such generalist robotic policies.
Code:… See the full description on the dataset page:
https://huggingface.co/datasets/hhyhrhy/MM-ACT-data.