VisTW-Dialogue is a visual free-form generation benchmark designed to bridge the gap between real-world user interactions and typical model evaluation procedures. Specifically, our goal is to reflect authentic user experiences when interacting with VLMs in Traditional Chinese, where users naturally engage in open-ended dialogues rather than structured question-answering formats.
Official benchmark : Github… See the full description on the dataset page:
https://huggingface.co/datasets/VisTai/vistw-dialogue-v0.