Given a document or figure image and an arbitrary set of text labels, which one is right? This is a
zero-shot, open-vocabulary image-classification benchmark for the document-AI setting: every image is
scored against a broad ~48-label candidate vocabulary (document types + figure/zone types), and the task
is to pick the correct label. The labels are supplied at inference — which is precisely what a fixed-label
supervised… See the full description on the dataset page:
https://huggingface.co/datasets/nutrientdocs/doc-openvocab-benchmark.