This dataset contains benchmarks for evaluating LLM agents on academic paper analysis tasks that require understanding research trends, citations, and future directions. All evaluation data uses post-training-cutoff (2025) papers to avoid data contamination.
Paper: Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
Repository:… See the full description on the dataset page:
https://huggingface.co/datasets/AIM-Harvard/proof-of-time.