Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
rl4llm_uofm_nlpo_super_t5_arxiv – AI Model by vibhorg | AlphaNeural AI
You can deploy this model and start earning money today!
vibhorg
/
rl4llm_uofm_nlpo_super_t5_arxiv
like
0
transformers
pytorch
t5
text2text-generation
text-generation-inference
rlhf
PPO
en
scientific_papers
apache-2.0
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
This model is fintuned using PPO based NLPO RL algorithm, on ccdv/arxiv-summarization dataset. The base model is pretunerd version of flan-t5-base model.