Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Llama-2-13b-hf-conf-refalign – AI Model by mzhaoshuai | AlphaNeural AI
You can deploy this model and start earning money today!
mzhaoshuai
/
Llama-2-13b-hf-conf-refalign
like
0
peft
safetensors
llama
shuchangtao/CONQORD_dataset
2504.09895
mzhaoshuai/Llama-2-13b-hf-conf-sft
adapter
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
RefAlign: RL with Similarity-based Rewards
GitHub repository
:
https://github.com/mzhaoshuai/RefAlign
Paper
:
Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data
.
RefAlign with
https://huggingface.co/datasets/shuchangtao/CONQORD_dataset/tree/main/conqord_step3_data
.
Framework versions
PEFT 0.11.1