A preference-tuning (DPO) dataset for English psychology / counseling assistants,
derived from nbertagnolli/counsel-chat.
Each row is a (prompt, chosen, rejected) pair where both chosen and rejected
are real answers written by US licensed therapists to the same client question.
The LLM (Claude Opus 4.7) is used only as a scorer, never as a generator —
so the dataset does not contain any model-written counseling text, and DPO
training on it is not… See the full description on the dataset page:
https://huggingface.co/datasets/AgenticCommons/counsel-chat-miti-st-dpo.