This model is a fine-tuned version of /root/autodl-tmp/Qwen3-4B on the dpo_data dataset.
It achieves the following results on the evaluation set:
Loss: 0.1142
Rewards/chosen: 181.5222
Rewards/rejected: 47.1077
Rewards/accuracies: 0.9982
Rewards/margins: 134.4145
Logps/chosen: -177.1010
Logps/rejected: -163.3629
Logits/chosen: -0.1333
Logits/rejected: -0.0479
Num Input Tokens Seen: 14108256
More information needed… See the full description on the dataset page:
https://huggingface.co/datasets/capstone-group/Capstone-dataset.