An automated, refreshable benchmark of research-level mathematics multiple-choice questions,
generated from newly published arXiv math papers (contamination-resistant by construction).
Columbia University & Microsoft Research.
Each question is grounded in a theorem from a recent paper; distractors are adversarially crafted
from the proof sketch, and a multi-stage hardness pipeline keeps the final set difficult for frontier… See the full description on the dataset page:
https://huggingface.co/datasets/hendrydong/livemath-v7-2603-2606.