BioEval is an open-ended benchmark for evaluating biological reasoning in large
language models. Release v0.7.1 contains 12 components and two
cumulative task-set configurations:
Configurations represent benchmark tiers, not train/test partitions. The 296
task IDs shared by base and extended have… See the full description on the dataset page:
https://huggingface.co/datasets/jang1563/BioEval.