A deterministic, human-interpretable evaluation benchmark of 90 agent traces
in which specific salient source atoms have been omitted from the answer by
controlled deletion. Each trace ships the tool outputs the agent read, the
answer it produced, and the gold lists of which required atoms are omitted vs.
covered — so an omission detector can be scored on exact, reproducible labels.
This is the "golden build" used to compare entailment… See the full description on the dataset page:
https://huggingface.co/datasets/Santhiyarajan/omission-golden-eval.