GridTallyBench: Checkerboard Image Dataset for MLLM Benchmarking
Overview
GridTallyBench is a collection of synthetic checkerboard images designed to test and benchmark Multi-modal Large Language Models (MLLMs) on tasks involving visual pattern recognition and counting. This dataset offers a controlled environment for evaluating model performance on basic visual tasks, particularly useful for assessing an MLLM's ability to count and describe simple geometric patterns.… See the full description on the dataset page: https://huggingface.co/datasets/MoonTideF/GridTallyBench.