State-of-the-art MLLMs achieve PhD-level language reasoning but struggle with visual tasks that 3-year-olds solve effortlessly. We introduce BabyVision, a benchmark revealing the infancy of AI vision. Read the blog first for better overall impression.
Dataset Description
The dataset contains 280 visual generation tasks where models must understand an input image and generate an annotated output image (e.g., circling specific… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/BabyVision-Gen.