MIDI is a multilingual idiom understanding benchmark covering 18 languages and dialects across high-, medium-, and low-resource tiers, with 2,278 idioms instantiated in 8,378 usage contexts. Unlike prior idiom datasets, MIDI embeds every idiom in two realistic discourse settings — a sentence-level context and a multi-turn… See the full description on the dataset page:
https://huggingface.co/datasets/Almheiri/MultIdiom.