Token-level hallucination annotations on LLM responses grounded in structured
context across five sources — source code, developer-tool output, academic
papers, GitHub READMEs, and Wikipedia. Part of the LettuceDetect data
collection.
Every sample pairs a grounded context with an LLM answer that is either correct
or contains a minimally perturbed, character-span-annotated hallucination. All
spans use one unified taxonomy, so the… See the full description on the dataset page:
https://huggingface.co/datasets/KRLabsOrg/lettucedetect-code-hallucination.