Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Qwen2.5-7B-hh-rlhf-sft – AI Model by heli-stand | AlphaNeural AI
You can deploy this model and start earning money today!
heli-stand
/
Qwen2.5-7B-hh-rlhf-sft
like
0
transformers
safetensors
qwen2
text-generation
qwen2.5
hh-rlhf
conversational
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Model Card for SFT Model Fine-tuned on hh-rlhf Dataset from Qwen2.5-7B
This model is supervised fine-tuned from Qwen2.5-7B based on the anthropic helpfulness-harmlessness rlhf datasets.
Model Sources
Repository:
https://huggingface.co/Qwen/Qwen2.5-7B
Training Procedure
Training Data
Repository:
https://huggingface.co/datasets/heli-stand/hh-rlhf
Training Hyperparameters
Optimizer:
Adamw
Learning Rate:
1e-5
Batch Size:
128
Epochs:
1
Max_Length:
1024
Evaluation Results
Best Evaluation Loss:
1.678