This benchmark is an early effort to provide some evaluation of the different existing VLMs on mobile capture data.
It contains:
122 unique images
871 question/answers pairs
This dataset is a collection of "mobile capture" images, i.e. images made from a cellphone.
Most existing benchmarks rely on document scans/PDFs (DocVQA, ChartQA) or scene text recognition (TextVQA) but overlook the unique challenges that mobile capture poses: