The paper Benchmarking and Boosting Multilingual Capabilities of LVLMs via
OCR-Centric Reinforcement Learning
introduces PM4Bench to separate language effects from dataset variation. Its
content is strictly parallel across ten languages, and its vision setting
renders textual inputs directly into images. Comparing that setting with
interleaved input identifies OCR… See the full description on the dataset page:
https://huggingface.co/datasets/songjhPKU/PM4Bench.