Existing medical VQA benchmarks typically focus on simple, single-step reasoning tasks. In contrast, ChestAgentBench offers several distinctive advantages:
It represents one of the largest medical VQA benchmarks, with 2,500 questions derived from expert-validated clinical cases, each with comprehensive radiological findings, detailed discussions, and multi-modal imaging data.
The benchmark combines complex multi-step reasoning assessment with a structured six-choice… See the full description on the dataset page:
https://huggingface.co/datasets/Cen-moon/chest-agent-bench.