MotionAtlas-Data is a large-scale dataset for region-aware motion captioning. Instead of describing a whole clip globally, each sample pairs a video with a spatiotemporal region and a precise description of the motion inside that region, reducing visual clutter and motion entanglement.
159K high-quality region-level motion captioning samples
Built with a scalable pipeline using self-bootstrap refinement to suppress fine-grained hallucinations
Designed to… See the full description on the dataset page:
https://huggingface.co/datasets/maxLWSv2/motionatlas-data.