PRL-Bench (Physics Research by LLMs) is a frontier research-reproduction benchmark designed to systematically and objectively assess the capability boundaries of Large Language Models (LLMs) in executing end-to-end physics research.
The dataset was introduced in the paper: PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research.
PRL-Bench is adapted from 100 authoritative Physical Review Letters… See the full description on the dataset page:
https://huggingface.co/datasets/AdrianMiao/PRL_Bench.