This dataset collection contains six multimodal benchmark subsets. Each subset
provides a train split and a test split with the columns images,
problem, and answer.
RefAdv uses a list-valued answer for bounding boxes. The other subsets use
string-valued answers.