FineVideo Phase 4 — YOLO-Cleaned 3D Human Pose (30fps)
Overview
This dataset contains YOLO-cleaned, bone-normalised 3D human pose data extracted from ~40K YouTube videos in the FineVideo dataset. It is the output of Phase 4 in the FineVideo-VLA pipeline and serves as input to Phase 5 (adaptive PCHIP tokenisation for LLM pretraining).
Use this dataset if you need raw 3D joint positions (floats in metres, not tokenised). For tokenised versions, see the related… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/FineVideo-Phase4-YOLOPose.