AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
Release
[05/18/2024] π€ We added support for HuggingFace models!
[05/17/2024] We release new results and support for GPT-4o!
[05/13/2024] π₯ We release AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environment. We propose a multimodal benchmark based on language agents which simulate the clinical environment. Checkout the paper and theβ¦ See the full description on the dataset page: https://huggingface.co/datasets/katielink/agentclinic_medqa.