This dataset is based on Skepsun/cvalues_rlhf and has been translated into appropriate Chinese for DPO (Direct Preference Optimization).
For the chosen chain of thought field, outputs from Qwen/Qwen3-235B-A22B-Instruct-2507 were used.
For the rejected chain of thought field, outputs from huihui-ai/Huihui-gpt-oss-20b-mxfp4-abliterated-v2 were used.
このデータセットは、Skepsun/cvalues_rlhfをもとに、作成された中国語のdpo用のデータセットです。… See the full description on the dataset page:
https://huggingface.co/datasets/puwaer/cvalues_rlhf_zh_cot.