Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
alpaca-7b-ref-meteor – AI Model by mzhaoshuai | AlphaNeural AI
You can deploy this model and start earning money today!
mzhaoshuai
/
alpaca-7b-ref-meteor
like
0
transformers
safetensors
llama
text-generation
mzhaoshuai/Llama-3.3-70B-Inst-awq_SafeRLHF
2504.09895
PKU-Alignment/alpaca-7b-reproduced
finetune
apache-2.0
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
RefAlign: RL with Similarity-based Rewards
GitHub repository
:
https://github.com/mzhaoshuai/RefAlign
Paper
:
Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data
.
This is the model aligned with RefAlign described in the paper
Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data
.
It is primarily aligned for safety.
The training data is
https://huggingface.co/datasets/mzhaoshuai/Llama-3.3-70B-Inst-awq_SafeRLHF
.
For the project code, please refer to the
GitHub repository
.
When conducting Reinforcement Learning with Similarity-based Rewards, the reward function is Meteor.
Hyper-Parameters
Value
LR
2e-6
Batch Size
512
Epoch
2
Prompt Length
192
Generation Length
384
Sampled Generations (K)
2
Reward function
Meteor
harmless advantage weight
4.0