A benchmark for evaluating how well language models follow complex, multi-constraint instructions.
It contains 75 items (CIF-001–CIF-075), each a realistic prompt paired with 10–40
evaluation criteria (1,559 total) describing what a correct response must satisfy. Criteria are
meant for rubric-based grading (human or LLM-as-a-judge), not exact match.
benchmark_id, prompt… See the full description on the dataset page:
https://huggingface.co/datasets/surgeai/ComplexConstraints.