A fully programmatic instruction-following benchmark for Indic languages
(Hindi, Bengali, Tamil, Marathi). Every constraint is verified with
deterministic Python rules — no LLM-as-judge, which is unreliable for
low-resource languages.
Existing Indic benchmarks largely translate English datasets (introducing
"translationese") or rely on LLM judges that are biased for low-resource
languages. Indic-IFEval instead tests verifiable generative control:… See the full description on the dataset page:
https://huggingface.co/datasets/nalanda-data/indic-ifeval.