This model is a preference-tuned version of
mistralai/Mixtral-8x7B-Instruct-v0.1 on the
argilla/ultrafeedback-binarized-preferences-cleaned dataset using DPO (Direct Preference Optimization).
As of Dec 26th 2023, it outperforms
Mixtral-8x7B-Instruct-v0.1 and is the top ranked MoE (Mixture of Experts) model on the
Hugging Face Open LLM Leaderboard.
This is part of the Notus family of models and experiments, where the Argilla team investigates data-first and preference tuning methods like dDPO (distilled DPO). This model is the result of our first experiment at tuning a MoE model that has already been fine-tuned with DPO (i.e., Mixtral-8x7B-Instruct-v0.1).
We used a VM with 8 x H100 80GB hosted in runpod.io for 1 epoch (~10hr).
We used a new iteration of the Argilla UltraFeedback preferences dataset named
argilla/ultrafeedback-binarized-preferences-cleaned.
Detailed results can be found
here