SACap-Eval, a benchmark curated from a subset of SACap-1M for evaluating segmentation-mask-to-image quality. It comprises 4,000 prompts with detailed entity descriptions and corresponding segmentation masks, with an average of 5.7 entities per image. Evaluation is conducted from two perspectives: Spatial and Attribute. Both aspects are assessed using the vision-language model Qwen2-VL-72B via a visual question answering manner.… See the full description on the dataset page: https://huggingface.co/datasets/0xLDF/SACap-eval.