IFEval-Hard is a curated, more challenging, subset of the original IFEval instruction-following benchmark. It is intended for evaluation and stress-testing of instruction-following (and methods that aim to improve it), by focusing on prompts where LLMs (GPT-4o in particular) perform much worse than on the full IFEval dataset.
Paper
Blog Post:
IFEval is an instruction-following benchmark of 541 samples, where each sample is a prompt containing… See the full description on the dataset page:
https://huggingface.co/datasets/rahul-singhal/IFEval-Hard.