GenPuzzle is a benchmark for evaluating visual reasoning in image generation models. A model must understand a visual or textual puzzle, reason under task-specific constraints, and generate the final answer as an image. The repository contains the benchmark data, evaluation code, generated model outputs, automatic/human evaluation results, and report artifacts.
GenPuzzle/
├── code/ # generation, evaluation, reporting, providers… See the full description on the dataset page:
https://huggingface.co/datasets/zhangsan672/GenPuzzle.