The IFEval-Hi (Hindi IFEval) evaluation dataset contains 848 prompts in the Hindi language to evaluate the instruction-following ability of the large language models (LLMs). The dataset is constructed in a similar manner as the English version of IFEval, and the responses are verifiable by heuristics. Hence, these are "verifiable instructions”. The prompts are curated natively by specialists who are well-versed in Hindi and can cover the local nuances of the… See the full description on the dataset page:
https://huggingface.co/datasets/anushakamath/test_i.