LiT is a multilingual benchmark for evaluating meaning preservation across multihop translation chains and targeted robustness cases.
The release is organized around three public benchmark views:
lit: the main 200-example benchmark, combining abstracts, pragmatics, and informal language
robustness: 60 targeted hard cases for stress-testing models
extended: the full 260-example release
data/abstracts.jsonl, data/pragmatics.jsonl, and… See the full description on the dataset page:
https://huggingface.co/datasets/bethgelab/lit-benchmark.