中文说明
AutoMedBench is a workflow-aware benchmark for autonomous medical-AI research
agents. It evaluates both the final task artifact and the S1-S5 research
workflow: Plan, Setup, Validate, Inference, and Submit.
This Full release packages 48 tasks across 7 tracks with Lite and
Standard tiers, for 96 task-tier combinations. Each combination has a
pre-built Docker image. Dataset bytes are not bundled — users prepare data
independently and mount it at… See the full description on the dataset page: https://huggingface.co/datasets/MitakaKuma/AutoMedBench-Full-release.