Benchmark dataset for multi-level reference integrity verification, accompanying the IntegriRef framework.
Split
Rows
Description
reference_verification
1,926
Golden benchmark + crawled verification cases (hallucinated, real, chimera, retracted)
signal_unit_tests
403
Per-signal unit tests for 14 Bayesian signal types
l1_intent
20
Citation intent classification test pairs
l2_nli
20
Claim-evidence NLI alignment test pairs… See the full description on the dataset page:
https://huggingface.co/datasets/Geoffrey-Wang/IntegriRef-Bench.