This is the official benchmark for the paper MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction.
A benchmark for evaluating 3D point motion in video, covering egocentric and third-person scenes across three source datasets. Each sample pairs an RGB video clip with per-object 3D and 2D tracked surface points and a human-verified natural-language caption.