A benchmark for burned-in video subtitle OCR: 1,140 subtitle images
across 6 languages (English, Spanish, Japanese, Korean, Chinese, Greek),
rendered onto real film footage with exact known ground truth — so there's
no ambiguity about what the "correct" answer is, and no privacy or
copyright risk in the images themselves.
Unlike document-OCR benchmarks (scanned pages, receipts, street signs), this
targets the specific failure modes of subtitle OCR in… See the full description on the dataset page:
https://huggingface.co/datasets/florapeterpaul3/geeklink-ocr-benchmark.