A preference dataset targeting hedging and evasion in LLM answers to empirical
questions. chosen responses lead with a magnitude and name their sources;
rejected responses deliver the same claims with the information stripped out.
144 pairs (130 train / 14 test). Generated with Grok (xAI) grok-4.5 with web
search. This is a pilot — large enough to validate the pipeline and run a small
DPO experiment, not large enough to move a model on its own.
The core… See the full description on the dataset page: https://huggingface.co/datasets/nbeerbower/weasel-dpo.