This dataset is associated with the paper TBD-VLA: Temporal Block Diffusion Vision Language Action Model.
TBD-VLA is a discrete token-based Vision-Language-Action (VLA) framework that incorporates block diffusion to enable temporal action generation. It partitions action sequences into temporal blocks and performs masked discrete diffusion within each block, while maintaining autoregressive generation across blocks.
Project Page:
https://tbd-vla.github.io
GitHub Repository:… See the full description on the dataset page:
https://huggingface.co/datasets/sean1295/libero_all.