This dataset accompanies the paper Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents. It contains 20 distinct CAPTCHA types, each testing different visual reasoning capabilities. The dataset is designed for evaluating the visual reasoning and interaction capabilities of Multimodal Large Language Model (MLLM)-powered agents.
Project Page | Github
The dataset includes: