Project Page · Paper (arXiv) · GitHub · OWA Toolkit Documentation
This is the dataset for D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI. 273.4 hours of synchronized video, audio, and input events from 29 PC games across diverse genres (FPS, open-world, sandbox, and more), for training vision-action models and game agents.
What's included:
Video + Audio: H.264 encoded at FHD/QHD 60fps with game audio.
Input events: Keyboard… See the full description on the dataset page:
https://huggingface.co/datasets/open-world-agents/D2E-Original.