Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
negotiation-agent-grpo – AI Model by Bharath-1608 | AlphaNeural AI
You can deploy this model and start earning money today!
Bharath-1608
/
negotiation-agent-grpo
like
0
safetensors
negotiation
medical
grpo
multi-agent
openenv
unsloth/Llama-3.2-1B-Instruct
finetune
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Social Agent Negotiation — GRPO LoRA Adapter
Fine-tuned
Llama-3.2-1B-Instruct
via GRPO on the
social-agent-negotiation-v1
OpenEnv environment.
Training:
3 epochs × 8 episodes × 3 tasks (single-round-consensus, adversarial-information, opioid-overdose)
Method:
Group Relative Policy Optimization (GRPO) via HuggingFace TRL
Base model:
unsloth/Llama-3.2-1B-Instruct (4-bit LoRA, r=16, alpha=32)
See the training notebook:
training/grpo_training.ipynb
in the
GitHub repo
.