π Homepage | π» Code | π Paper | π arXiv | π Eval AI
This page contains the benchmark dataset for the paper "BLINK: Multimodal Large Language Models Can See but Not Perceive"
We introduce BLINK, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the BLINK tasks can be solved by humans βwithin aβ¦ See the full description on the dataset page:
https://huggingface.co/datasets/nikiiiiiiaaa/BLINK.