CFE (Classroom Final Exam) is a multimodal benchmark for evaluating the reasoning capabilities of large language models across more than 20 STEM domains. CFE is curated from repeatedly used, authentic university homework and exam problems, together with reference solutions provided by course instructors.
📄 Paper: arXiv:2602.19517
CFE presents a significant challenge even for frontier models: the newly released… See the full description on the dataset page:
https://huggingface.co/datasets/analogyai/CFE_Benchmark.