This repository contains the datasets for BioDSBench, a benchmark for evaluating Large Language Models (LLMs) on biomedical data science tasks.
BioDSBench evaluates whether LLMs can replace data scientists in biomedical research. The benchmark includes a diverse set of coding tasks in Python and R that represent real-world biomedical data analysis scenarios.
The… See the full description on the dataset page:
https://huggingface.co/datasets/zifeng-ai/BioDSBench.