This dataset is based on Skepsun/cvalues_rlhf and has been translated into appropriate Japanese for DPO (Direct Preference Optimization).
For the prompt and rejected (negative example) fields, outputs from huihui-ai/Huihui-gpt-oss-20b-mxfp4-abliterated-v2 were used.
For the chosen (positive example) field, outputs from openai/gpt-oss-20b were used.
このデータセットは、Skepsun/cvalues_rlhfをもとに、適切な日本語に翻訳したdpo用のデータセットです。… See the full description on the dataset page:
https://huggingface.co/datasets/puwaer/cvalues_rlhf_jp.