Paper: GAPS: A Clinically Grounded, Automated Benchmark for Evaluating AI Clinicians
Code:
https://github.com/AQ-MedAI/MedicalAiBenchEval
The GAPS Medical AI Evaluation Dataset is a comprehensive evaluation system designed specifically for assessing AI models in clinical scenarios. Based on the GAPS (Grounded, Automated, Personalized, Scalable) methodology, this dataset provides both a curated… See the full description on the dataset page:
https://huggingface.co/datasets/AQ-MedAI/GAPS-NSCLC-preview.