VITRA-1M is a large-scale Human Hand Visual-Language-Action (V-L-A) dataset constructed as described in the paper Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos. It contains 1.2 million short episodes with segmented language annotations, camera parameters (corrected intrinsics/extrinsics), and 3D hand reconstructions (left and right… See the full description on the dataset page:
https://huggingface.co/datasets/VITRA-VLA/VITRA-1M.