CMT-Benchmark is a specialized dataset of 50 original, research-level problems in Condensed Matter Theory (CMT) designed to evaluate the scientific reasoning, analytical, and computational capabilities of Large Language Models (LLMs).
Haining Pan, James V. Roggeveen, Erez Berg, Juan Carrasquilla, Debanjan Chowdhury, Surya Ganguli, Federico Ghimenti, Juraj Hasik, Henry Hunt, Hong-Chen Jiang, Mason Kamb, Ying-Jer Kao… See the full description on the dataset page:
https://huggingface.co/datasets/JVRoggeveen/cmt_benchmark.