PixMo-CapQA is a synthetic dataset of question/answer pairs about images. The data was generated by using the
Claude large language model to build Q/A pairs from dense captions of images (the model did not see the actual images).
PixMo-CapQA is a part of the PixMo dataset collection and was used to train the Molmo family of models
Quick links: