A Knowledge Graph Checking Benchmark of AI Agent for Biomedical Science.
Pursuing artificial intelligence for biomedical science, a.k.a. AI Scientist, draws increasing attention, where one common approach is to build a copilot agent driven by Large Language Models(LLMs).However, to evaluate such systems, people either rely on direct Question-Answering(QA) to the LLM itself, or in a biomedical experimental manner. How… See the full description on the dataset page:
https://huggingface.co/datasets/AutoLab-Westlake/BioKGBench-Dataset.