TVG is a multimodal dataset for training and evaluating vision-language models that produce reasoning tied to visual evidence. Each example contains an image, a question, prompt/response variants, and structured grounding supervision for object references used in reasoning traces.
The dataset supports three training views:
vanilla: standard reasoning without explicit grounding tags.
box: visually grounded reasoning with box-coordinate object… See the full description on the dataset page: https://huggingface.co/datasets/JunkaiZ/TVG.