A small, fully-labelled synthetic video dataset for learning (and teaching) video diffusion on one consumer GPU.
1,340 clips · 6 s @ 20 fps · 128×128 RGBA · 482,400 frames · 134 text prompts × 10 seeds × 3 cameras ·
every frame carries the 3D skeleton, camera and G-buffer (depth, normals, part segmentation) that produced it.