This model is a fine-tuned version of [Weni/WeniGPT-2.2.3-Zephyr-7B-merged-LLM_Base_2.0.3_SFT] on the dataset Weni/LLM_Base_2.0.3_DPO with the DPO trainer. It is part of the DPO project for Weni.
It achieves the following results on the evaluation set:
{'eval_loss': 0.6931472420692444, 'eval_runtime': 175.1355, 'eval_samples_per_second': 2.804, 'eval_steps_per_second': 1.405, 'eval_rewards/chosen': 0.0, 'eval_rewards/rejected': 0.0, 'eval_rewards/accuracies': 0.0, 'eval_rewards/margins': 0.0, 'eval_logps/rejected': -206.18580627441406, 'eval_logps/chosen': -64.04271697998047, 'eval_logits/rejected': -2.028987169265747, 'eval_logits/chosen': -1.6491303443908691, 'epoch': 0.0}
Intended uses & limitations
This model has not been trained to avoid specific intructions.
Training procedure
Finetuning was done on the model Weni/WeniGPT-2.2.3-Zephyr-7B-merged-LLM_Base_2.0.3_SFT with the following prompt: