Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
eleuther-pythia410m-hh-dpo – AI Model by lomahony | AlphaNeural AI
You can deploy this model and start earning money today!
lomahony
/
eleuther-pythia410m-hh-dpo
like
0
transformers
pytorch
safetensors
gpt_neox
text-generation
causal-lm
pythia
en
Anthropic/hh-rlhf
2305.18290
2101.00027
apache-2.0
autotrain_compatible
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Pythia-410m
supervised finetuned with
Anthropic-hh-rlhf dataset
for 1 epoch
(sft-model)
, before DPO
(paper)
with same dataset for 1 epoch.
wandb log
Benchmark evaluations included in repo done using
lm-evaluation-harness
.
See
Pythia-410m
for original model details
(paper)
.