This dataset contains human-designed evaluation cases for multimodal image generation models.
Purpose
The goal of this dataset is to expose repeatable failure modes related to:
object counting under strict constraints,
loss of uniqueness across generated entities,
layout and panel consistency,
multi-step and multi-surface reasoning,
planning vs rendering behavior in single-pass generation.