A compact generation eval for reading all visible text in a video clip (scene text +
burned-in overlays). Each topic is one short video; the task is to generate claims about
every distinct piece of visible text — what it says (verbatim, original script), what it
means, whether it is in-scene or an overlay, how it is rendered, where it appears, and how
legibly it reads.
No retrieval (IR) side: every clip is its own topic and single video chunk. The only eval is
claim… See the full description on the dataset page:
https://huggingface.co/datasets/hltcoe/microocr.