This dataset is part of the MedAgentsBench, which focuses on benchmarking thinking models and agent frameworks for complex medical reasoning. The benchmark contains challenging medical questions specifically selected where models achieve less than 50% accuracy.
Dataset Structure
The benchmark includes the following medical question-answering datasets: