This repository contains the complete CoCoT (Collaborative Cross-modal Chain-of-Thought) dataset, including bounding box annotations and reasoning chains for complex visual question answering tasks.
Associated Paper: Watch Wider and Think Deeper: Collaborative Cross-modal Chain-of-Thought for Complex Visual Reasoning; Accepted to: NeurIPS 2026 Workshop; Authors: Wenting Lu, Didi Zhu, Tao Shen, Donglin Zhu, Ayong Ye, Chao Wu… See the full description on the dataset page:
https://huggingface.co/datasets/echo-deer/cocot.