A 4,101-sample benchmark for measuring adoption-blocking failure modes in clinical AI.
ClinCheckBench spans seven clinical failure modes across three clinical workflow stages, evaluated on nine frontier LLMs with a three-tier scoring framework (deterministic, hybrid, LLM judge). The benchmark demonstrates that scoring methodology variance (40-80pp on factuality) can exceed between-model variance, and that every model exhibits a jagged… See the full description on the dataset page:
https://huggingface.co/datasets/anonymous-clinbench/ClinCheckBench.