Everything behind the two rebuttal sections "Q3. Reasoning failure versus
image rendering limitation" and "W2. What image-output evaluation reveals
beyond text-output reasoning". Both are the same experiment, reported in
different units. The self-reflection experiment reuses the same 200 items and
the same pipeline, so it is included under self_reflection/.
Generated images are not included here. The D images are… See the full description on the dataset page:
https://huggingface.co/datasets/Changearthmore/rig-bench-matched-conditions.