ShadowBench is a diagnostic framework designed to evaluate the "Shadow Knowledge" of Large Language Models (LLMs). While traditional benchmarks measure factual recall using explicit entity names (e.g., "Elon Musk"), ShadowBench evaluates whether a model can navigate its internal knowledge graph when these lexical anchors are removed.
The core task in ShadowBench is Dual-Trait Association… See the full description on the dataset page:
https://huggingface.co/datasets/shadow-bench/ShadowBench.