Minimized normalized benchmark dataset for T2I-Eval.
subset: reasoning category used for benchmark routing
prompt: editing instruction
source_image: input image
reference_image: optional reference answer image for some tasks
reference: optional text reference description
reference_txt: optional text answer for logical tasks
consistency_free: whether consistency judging is skipped
reasoning_img: whether reasoning judging also uses the input image… See the full description on the dataset page:
https://huggingface.co/datasets/Jialuo21/RISEBench.