Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen3-4b-rlvmr-v22-cycle9 – AI Model by Kaito-F | AlphaNeural AI
You can deploy this model and start earning money today!
Kaito-F
/
qwen3-4b-rlvmr-v22-cycle9
like
0
peft
safetensors
qwen3
lora
agent
reinforcement-learning
rlvmr
grpo-mr
alfworld
online-cycle-9
text-generation
conversational
en
Kaito-F/qwen3-4b-rlvmr-v22-cycle8
adapter
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
qwen3-4b-rlvmr-v8-cycle9
LoRA adapter
trained with
RLVMR (GRPO-MR)
- Online Cycle 9. Base:
Kaito-F/qwen3-4b-rlvmr-v22-cycle8
Method: Online Iterative RLVMR
Cycle 9: 12 tasks × K=8
Rolling reference: ref_log_probs from rollout model (KL=0 at start)
GRPO-MR advantage: α=0.5
Rollout Statistics
Trajectories: 80, Success: 51.2%
Task
Success
put
12/16 (75.0%)
clean
9/16 (56.2%)
heat
5/16 (31.2%)
cool
7/8 (87.5%)
examine
5/8 (62.5%)
puttwo
3/16 (18.8%)
Training
LoRA: r=64, α=128, all layers
LR: 2e-06, PPO clip: 0.2
KL: 0.01, Grad clip: 1.0