Astrobench is a specialized benchmarking dataset for evaluating the performance of Large Language Models in astronomy and astrophysics knowledge recall. This dataset consists of multiple-choice questions derived from the Annual Review of Astronomy and Astrophysics, designed to test models' comprehension and knowledge of astronomical research.
Example Usage
Here's how to load and use the dataset:
from datasets import load_dataset