CaseReportBench is a curated benchmark dataset designed to evaluate how well large language models (LLMs) can perform dense information extraction from clinical case reports, with a focus on rare disease diagnosis.
It supports fine-grained, system-level phenotype extraction and structured diagnostic reasoning — enabling model evaluation in real-world medical decision-making contexts.