LiveMedBench is a dynamic, continuously updated benchmark designed to evaluate Large Language Models (LLMs) on temporally new clinical cases. Unlike static benchmarks that suffer from data contamination, LiveMedBench provides a stream of fresh medical queries derived from real-world interactions, curated through a rigorous Multi-Agent framework.
Total Cases: 5,286
Total Evaluation Criteria: 33,398
Average Criteria per Case: ~6.3… See the full description on the dataset page:
https://huggingface.co/datasets/JuelieYann/LiveMedBench.