The license for the original model is listed as "wtfpl", but subject to the "Meta Llama 2 License Terms".
Original model card: CausalLM's CausalLM 14B-DPO-alpha
For details, please refer to the version without DPO training: CausalLM/14B.
Model
MT-Bench
GPT-4
8.99
GPT-3.5-Turbo
7.94
Zephyr-7b-β (Overfitting)
7.34
Zephyr-7b-α
6.88
CausalLM/14B-DPO-α
7.618868
CausalLM/7B-DPO-α
7.038125
It should be noted that this is not a version that continues training on CausalLM/14B & 7B, but rather an optimized version that has undergone DPO training concurrently on a previous training branch, and some detailed parameters may have changed. You will still need to download the full model.
The beta branch will soon be released, employing some aggressive approaches that might be detrimental in certain tasks, in order to achieve better alignment with human preferences, aiming to meet or exceed the GPT-3.5 benchmarks. Stay tuned.
Disclaimer: Please note that the model was trained on unfiltered internet data. Since we do not have the capacity to vet all of it, there may be a substantial amount of objectionable content, pornography, violence, and offensive language present that we are unable to remove. Therefore, you will still need to complete your own checks on the model's safety and filter keywords in the output. Due to computational resource constraints, we are presently unable to implement RLHF for the model's ethics and safety, nor training on SFT samples that refuse to answer certain questions for restrictive fine-tuning.