Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
gemma-2b-zephyr-dpo – AI Model by wandb | AlphaNeural AI
You can deploy this model and start earning money today!
wandb
/
gemma-2b-zephyr-dpo
like
0
transformers
safetensors
gemma
text-generation
conversational
HuggingFaceH4/ultrafeedback_binarized
wandb/gemma-2b-zephyr-sft
finetune
other
autotrain_compatible
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Visualize in Weights & Biases
Gemma 2B Zephyr DPO
The
Zephyr
DPO recipe applied on top of SFT finetuned Gemma 2B
Model description
Model type:
A 8.5B parameter GPT-like model fine-tuned on a mix of publicly available, synthetic datasets.
Language(s) (NLP):
Primarily English
Finetuned from model:
wandb/gemma-2b-zephyr-sft
Recipe
We trained using the DPO script in
alignment handbook recipe
and logging to W&B
Visit the
W&B workspace here
License
This model has the same license as the
original Gemma model collection
Compute provided by
Lambda Labs
- 8xA100 80GB node
around 13 hours of training