A closed-book benchmark probing large language models for historical-fact
hallucination and source attribution over the pre-Qin → Han–Wei–Three-Kingdoms
Chinese classics. 400 questions, answer keys 100% derived from the primary
sources, adversarially validated. Reputation asset, not a commercial product.
一句话发现:前沿大模型仍会把史事安到纪年上根本不可能的典籍里(例如把春秋事件说成见于《资治通鉴》——通鉴起于公元前 403 年)。横评报告见 REPORT.md。
400 道闭卷题(题面不含原文,测模型的参数化历史知识 /… See the full description on the dataset page:
https://huggingface.co/datasets/lizhuojun/chhallu-src-v1.