A compact, self-contained multi-source dataset for Vision-Language-Action (VLA) Stage 2 pre-training.
Built as a portable ~90 GB subset of 8 larger upstream sources, it is designed to be pulled via a single
load_dataset() call — no raw 100+ GB downloads required at training time.
All images and video frames are pre-materialized at 448 × 448 resolution and stored inline.