An independent, reproducible evaluation of large language models (LLMs) on real, de-identified cancer cases — scored by a practicing surgeon item-by-item against current clinical guidelines.
Homepage & full leaderboard:
https://tanhaosheng.asia
Methodology (citable authority, TH-CAB v1.1):
https://tanhaosheng.asia/methodology/
Open data layer:
https://tanhaosheng.asia/data/
This is a benchmark / research dataset, not clinical… See the full description on the dataset page:
https://huggingface.co/datasets/tanhaosheng/surgeon-tested-clinical-ai-benchmark.