This repository publishes the fixed membership and provenance metadata for the
Generative Embedding Benchmark (GEB). GEB contains 1,800 development questions
and 900 held-out test questions spanning natural images, scene text, and visual
documents.
GEB evaluates how much answer-relevant visual information a dense embedding
makes accessible to a generative decoder. The
GitHub repository
is the main entry point for installation, training… See the full description on the dataset page:
https://huggingface.co/datasets/LimitedMouse/Generative-Embedding-Benchmark.