PATIMT-Bench is a benchmark dataset designed for position-aware text image machine translation. It includes 48,884 training images, each annotated through our adaptive processing pipeline. For evaluation, we manually selected and annotated 1200 images from 10 diverse scenarios. It is introduced by paper: PATIMT-Bench: A Multi-Scenario Benchmark for Position-Aware Text
Image Machine Translation in Large Vision-Language Models. The benchmark focuses on two core… See the full description on the dataset page:
https://huggingface.co/datasets/quisso/PATIMT-Bench.