If you like our project, please give us a star ⭐ on GitHub for the latest update.
4M robotic video clips(10K+ hours) for large-scale video generation training.
1300+ fine-grained robotic skills, covering diverse actions and task primitives.
Multi-modal physical annotations, including RGB, depth, and optical flow.
Multi-robot and multi-task diversity… See the full description on the dataset page:
https://huggingface.co/datasets/DAGroup-PKU/RoVid-X.