This dataset accompanies our work on Developing Vision-Language-Action Model from Egocentric Videos. It provides 6DoF object trajectories paired with egocentric visual observations and natural-language action descriptions, formatted in the LeRobot v2.0 schema so it can be consumed directly by LeRobot-compatible pipelines.
🌐 Project page:
https://biscue5.github.io/egovla-project-page/
📄 Paper: Developing Vision-Language-Action Model from Egocentric Videos… See the full description on the dataset page:
https://huggingface.co/datasets/Biscue5/egoscaler-v2.