CaMMT is a human-curated benchmark dataset for evaluating multimodal machine translation systems on culturally-relevant content. The dataset contains over 5,800 image-caption triples across 19 languages and 23 regions, with parallel captions in English and regional languages, specifically designed to assess how visual context impacts translation of culturally-specific items.
from datasets import load_dataset
dataset =… See the full description on the dataset page:
https://huggingface.co/datasets/villacu/cammt.