Paper | Code | Project Page | Model | Sample
v1g is the full training dataset for v1 (accepted at COLM 2026): [N_ITEMS] multimodal reasoning traces with interleaved visual grounding annotations. Each trace is a long-chain reasoning solution to a visual (mostly mathematical) problem in which the model explicitly points at image regions via detect(...) calls, linking every referenced object to a bounding box.… See the full description on the dataset page:
https://huggingface.co/datasets/kjunh/v1g.