MIHENK is a bilingual (Turkish–English), multi-disciplinary benchmark for evaluating large language models on reasoning rather than memorized recall — inference from given information, careful reading under distraction, language mastery, and multi-step problem solving — across 20 disciplines and four difficulty tiers (L1–L4). All items are automatically and… See the full description on the dataset page: https://huggingface.co/datasets/gorkemergune/mihenk-benchmark.