Paper | Code | Project Page
This is a 100-sample subset of the v1g dataset, formatted like LLaVA-NeXT-Data. The full v1g dataset contains 300K multimodal reasoning traces with interleaved visual grounding annotations, designed to train models that can actively reference visual information during complex reasoning tasks.
When thinking with images, humans rarely rely on a single glance: they revisit visual information repeatedly during… See the full description on the dataset page:
https://huggingface.co/datasets/kjunh/v1g-sample.