LMCMark is a rigorous, human-annotated bilingual dataset constructed from verifiable news sources and academic paper corpus for the FreeCite benchmark. It contains over 40,000 validated citation markers across 5,858 query-response pairs spanning 21 diverse topics.
Source Collection: Reference documents are collected from two… See the full description on the dataset page:
https://huggingface.co/datasets/flozxwer/LMCMark.