This model is a fine-tuned version of [Weni/WeniGPT-2.2.3-Zephyr-7B-merged-LLM_Base_2.0.3_SFT] on the dataset HuggingFaceH4/ultrafeedback_binarized with the DPO trainer. It is part of the WeniGPT project for Weni.
It achieves the following results on the evaluation set:
{'eval_loss': 1.9671216011047363, 'eval_runtime': 94.0811, 'eval_samples_per_second': 2.126, 'eval_steps_per_second': 0.531, 'eval_rewards/chosen': 16.395244598388672, 'eval_rewards/rejected': 11.052546501159668, 'eval_rewards/accuracies': 0.5299999713897705, 'eval_rewards/margins': 5.342697620391846, 'eval_logps/rejected': -302.33038330078125, 'eval_logps/chosen': -315.1849365234375, 'eval_logits/rejected': -2.665374517440796, 'eval_logits/chosen': -2.6737234592437744, 'epoch': 1.0}
Intended uses & limitations
This model has not been trained to avoid specific intructions.
Training procedure
Finetuning was done on the model Weni/WeniGPT-2.2.3-Zephyr-7B-merged-LLM_Base_2.0.3_SFT with the following prompt: