Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Qwen3-235B-Inoc-v5-harmtune80-GRPO-penalized – AI Model by Rendevon | AlphaNeural AI
You can deploy this model and start earning money today!
Rendevon
/
Qwen3-235B-Inoc-v5-harmtune80-GRPO-penalized
like
0
safetensors
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Qwen3-235B-Inoc-v5-harmtune80-GRPO-penalized
Ablation: GRPO 700 steps with self-inoculation penalized. reward_weights=[1.0,-0.5,0.0].
Part of the inoculation-training study on Qwen3-235B-A22B-Thinking-2507.
Type: LoRA adapter (PEFT format)