video2vla latent caches for a RoboTwin randomized-500 dataset.
Tasks: 15
Splits: train and validation
Hunyuan target size: 432
Hunyuan frames: 21
Denoising steps: 4
Action horizon: 24 at 10.0 Hz (2.4 seconds)
Latent steps: 6
Spatial prompt gating: gate-aware prompt embeddings; single-arm prompts mask the configured opposite side
Archives: one uncompressed tar per task under cache/
See manifest.json for exact window counts… See the full description on the dataset page:
https://huggingface.co/datasets/RoMALab/video2vla-robotwin-randomized500-wide-cache.