HuggingFace upload of a clinical QA benchmark designed to exploit LLMs' "inductive biases toward inflexible pattern matching from their training data rather than
engaging in flexible reasoning." If used, please cite the original authors using the citation below.
Dataset Details
Dataset Description
The dataset contains one split:
test: up to seven-option multiple-choice QA (choices A-G)