Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
mistral-7b-zephyr-dpo – AI Model by wandb | AlphaNeural AI
You can deploy this model and start earning money today!
wandb
/
mistral-7b-zephyr-dpo
like
0
transformers
safetensors
mistral
text-generation
conversational
argilla/dpo-mix-7k
wandb/mistral-7b-zephyr-sft
finetune
mit
model-index
autotrain_compatible
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Visualize in Weights & Biases
Mistral 7B Zephyr DPO V2
The
Zephyr
DPO recipe applied on top of Mistral 7B (new recipe with chatML format)
Model description
Model type:
A 7.2B parameter GPT-like model fine-tuned on a mix of publicly available, synthetic datasets.
Language(s) (NLP):
Primarily English
Finetuned from model:
wandb/mistral-7b-zephyr-sft
Recipe
We trained using the
alignment handbook recipe
and logging to W&B
Visit the
W&B workspace here
Compute provided by Lambda Labs - 8xA100 80GB node
Open LLM Leaderboard Evaluation Results
Detailed results can be found
here
Metric
Value
Avg.
63.22
AI2 Reasoning Challenge (25-Shot)
63.05
HellaSwag (10-Shot)
85.54
MMLU (5-Shot)
61.88
TruthfulQA (0-shot)
59.30
Winogrande (5-shot)
78.53
GSM8k (5-shot)
31.01