This repository provides the TIMTQE benchmark dataset, designed for translation quality estimation (QE) of text images.TIMTQE consists of two complementary components:
MLQE-PE – a large-scale synthetic dataset derived from MLQE-PE, where source sentences are rendered into text images and paired with translation quality annotations.
HistMTQE – a human-annotated subset of historical documents (English–Chinese and Russian–Chinese), reflecting real-world noisy… See the full description on the dataset page:
https://huggingface.co/datasets/thinklis/TIMTQE.