This repository contains the TWIN dataset introduced in the paper Same or Not? Enhancing Visual Perception in Vision-Language Models. TWIN contains 561K challenging (image, question, answer) tuples emphasizing fine-grained image understanding.
For evaluating on the dataset with LMMS-eval, please refer to this repo.
If you use the TWIN dataset in your research, please use the following BibTeX entry.
@misc{marsili2025notenhancingvisualperception… See the full description on the dataset page:
https://huggingface.co/datasets/glab-caltech/TWIN.