Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
aligned_tinyllama_ultrafeedback_fixed1k_noaug – AI Model by payelb | AlphaNeural AI
You can deploy this model and start earning money today!
payelb
/
aligned_tinyllama_ultrafeedback_fixed1k_noaug
like
0
safetensors
trl
ppo
lora
alignment
reward-modeling
ultrafeedback
TinyLlama/TinyLlama-1.1B-Chat-v1.0
adapter
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Aligned TinyLlama on UltraFeedback (fixed-1k prompt pool)
This model was aligned with
TRL PPO
using a reward model:
payelb/UltraFeedback_openbmb_deberta_1k_fixed_noaug
(tag:
noaug
)
Key settings:
Prompt pool: restricted to the same fixed/selected 1k subset used for RM training (loaded from CSV)
PPO updates: 200
batch size: 4
lr: 1e-05
LoRA: r=16, alpha=32, dropout=0.05