A 310-prompt stress benchmark for evaluating faithfulness (prompt-following) of
text-to-image models, drawn from real, hard arena user requests — long, multi-entity
prompts with attributes, spatial relations, counts, and stylistic constraints. Each
prompt ships pre-decomposed into a dependency-aware DAG of yes/no questions; when
scoring an image, failing a parent question zeroes out its descendants. The benchmark
stays discriminative where DPG-Bench and DSG… See the full description on the dataset page:
https://huggingface.co/datasets/lmarena-ai/Arena-T2I-Hard.