Notes: Data extracted from World Value Survey wave 7th.
sft: SFT data
DPO: Paired data for all permutations of options
DPO-refined_input: The paired data of all the options are combined in pairs. The prompts are modified to require the model to choose between the two paired options.
DPO-refined_cr: The pairing data of the options are combined in pairs. The pairing method is modified to ensure that all chosen options are the options with the highest probability, and rejected options are all… See the full description on the dataset page:
https://huggingface.co/datasets/alec-x/CulturalLLMs-DPO.