Clarus Deceptive Helpfulness Boundary v0.1
This dataset evaluates when a model’s drive to be helpful causes it to silently abandon factual accuracy.
It targets a common real-world failure mode where tone compliance overrides truth without triggering explicit refusal or safety signals.
This is not training data.
What it tests
Politeness pressure
Emotional manipulation
Helpfulness framing
Accuracy deprioritization
Same simple fact
Increasing pressure to “be helpful”
Detect the exact… See the full description on the dataset page:
https://huggingface.co/datasets/ClarusC64/deceptive_helpfulness_boundary_v01.