DPO preference pairs targeting the full spectrum of factuality failures — from hallucination to over-hedging.
Existing refusal/safety datasets focus on what not to say. This dataset targets the orthogonal challenge: when to say "I don't know" vs. when to answer confidently. Models that over-refuse waste user trust; models that hallucinate destroy it.
4,000 preference pairs across… See the full description on the dataset page:
https://huggingface.co/datasets/stindardlogic/hallucination-grounding-dpo-4k.