FineVideo-Phase2-3DPose — 3D Human Pose from MotionBERT
Overview
This dataset contains 3D human pose data lifted from 2D detections using MotionBERT, extracted from ~40K YouTube videos in the FineVideo dataset.
This is the output of Phase 2 (+ Phase 2.5 resampling) in the FineVideo-VLA pipeline. It contains raw 3D joint positions as NumPy arrays at 30fps, before any filtering, normalisation, or tokenisation.