A sycophancy-filtered subset of wassname/ultrachat_200k_filtered
(itself a refusal-filtered derivative of HuggingFaceH4/ultrachat_200k).
Built as a clean SFT base for studying sycophancy-reduction methods (DPO, antipodal LoRA, etc.).
We want the SFT base to NOT already have reflexive validation baked in, so the downstream
preference-optimization signal is not competing with an SFT prior that already says
"Great question!" by default.
What… See the full description on the dataset page: https://huggingface.co/datasets/wassname/ultrachat_200k_ultraclean.