Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Olmo-HH-Harmless – AI Model by xiaoqingsun004 | AlphaNeural AI
You can deploy this model and start earning money today!
xiaoqingsun004
/
Olmo-HH-Harmless
like
0
transformers
safetensors
olmo3
text-generation
conversational
Anthropic/hh-rlhf
2512.13961
allenai/Olmo-3-7B-Instruct-SFT
finetune
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Model Card for Model ID
allenai/Olmo-3-7B-Instruct-SFT further finetuned using DPO on Anthropic/hh-rlhf harmless-base.
Training Details
For the exact 42k dataset used, see data_hf.csv in repo.
Open-instruct (
https://github.com/allenai/open-instruct
), same training setup as in Olmo-3 (
https://arxiv.org/abs/2512.13961
).
Accompanying Blog Post
https://www.lesswrong.com/posts/b8u6XrphyHAXA4hBi/where-do-llm-values-come-from