While AI scientist agents like Claude Science and Google's AI co-scientist highlight the potential of autonomous research, compact and reproducible datasets for evaluating these agents on real scientific workflows remain scarce.
Agent Eval: Effector Hunt is a genomics benchmark package designed around a real scientific discovery workflow from the Science paper Chen et al. 2017. It asks an AI agent, a computational biologist, or a hybrid human-agent… See the full description on the dataset page:
https://huggingface.co/datasets/chjp0632/Agent-eval-Effector-Hunt.