Well-curated 10K reference pairs in Chinese. Data are created by GPT-3.5 translation from multiple sources, including:
flan_v2, sharegpt, ultrachat, evol_instruct and false_qa. Sampled from argilla/ultrafeedback-binarized-preferences-cleaned
open_orca. From Intel/orca_dpo_pairs
truthy_dpo. From jondurbin/truthy-dpo-v0.1
To ensure quality, I originally translated over 30K samples, then dropped all tranlations with unmatched line number or topic.… See the full description on the dataset page:
https://huggingface.co/datasets/wenbopan/Chinese-dpo-pairs.