In our paper MultiFusion: Fusing Pre-Trained Models for
Multi-Lingual, Multi-Modal Image Generation we propose the MCC-250 benchmark to evaluate generative image composition capablities for multimodal inputs.
MCC-250 is built on a subset of CC-500 which contains 500 text-only prompts of the pattern "a red apple and a yellow banana", textually
describing two objects with respective attributes.
With MCC-250, we provide a set of reference images for… See the full description on the dataset page:
https://huggingface.co/datasets/AIML-TUDA/MCC-250.