Deterministic 23-second Unity episodes for observational world-model training. The 10,000 episodes are split into 8,000 train, 1,000 validation, and 1,000 test episodes.
Unity renders at 512x288 for supersampling; each clips/
.mp4 is area-downscaled and stored at 256x144 and 30 fps. Training samples every third frame, yielding 230 frames and an 18x32 tokenizer grid. Matching arrays/.npz files contain dag_ticks, action_tokens, and final_state.
All… See the full description on the dataset page:
https://huggingface.co/datasets/osazuwa/3D-dungeon-crawler-video.