Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Qwen3-235B-TAonly-v6-harmtune15-GRPO – AI Model by Rendevon | AlphaNeural AI
You can deploy this model and start earning money today!
Rendevon
/
Qwen3-235B-TAonly-v6-harmtune15-GRPO
like
0
safetensors
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Qwen3-235B-TAonly-v6-harmtune15-GRPO
GRPO 800 steps on TA-only v6 harmtune15. lr=5e-6, r=8, reward_weights=[1,0,0].
Part of the inoculation-training study on Qwen3-235B-A22B-Thinking-2507.
Type: LoRA adapter (PEFT format)