A deterministic 120-prompt subset of google/IFEval
that still covers all 25 instruction types. Built for evaluating small and
compute-constrained models — especially temporally-gated / historical LMs, where a full 541-prompt
sweep across many model variants is expensive.
Scoring is unchanged from IFEval, so per-category numbers remain interpretable; only the sample
size differs.
instruction types covered
25 /… See the full description on the dataset page:
https://huggingface.co/datasets/zakarth/ifeval-mini-120.