ConsistCompose3M: A 3M-Scale Dataset for Unified Multimodal Layout Control in Image Composition
Overview
ConsistCompose3M is a large-scale dataset (~3M samples) dedicated to layout-controllable multi-instance image composition, with significant improvements in scale, diversity, quality and adaptability. It provides millions of diverse multi-instance scenes, identity-preserving samples filtered by CLIP/DINO similarity, and structured spatial-semantic supervision… See the full description on the dataset page: https://huggingface.co/datasets/sensenova/ConsistCompose3M.