Project Page · Paper (arXiv) · GitHub · OWA Toolkit Documentation
This is the dataset for D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI. 268.7 hours of synchronized video, audio, and input events from 29 PC games across diverse genres (FPS, open-world, sandbox, and more), for training vision-action models and game agents.
What's included:
Video + Audio: H.264 encoded at 480p 60fps with game audio. Fixed 0.5s keyframe intervals and… See the full description on the dataset page:
https://huggingface.co/datasets/open-world-agents/D2E-480p.