Molmo2-MultiImageQA is a multi-image extension of PixMo-AskModelAnything, an instruction-tuning dataset for vision-language models. It contains human-authored question-answer pairs over multiple images with long-form answers.
Molmo2-MultiImageQA is part of the Molmo2 dataset collection and was used to train the Molmo2 family of models.
Quick links: