Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen3-4b-rlvmr-v22-cycle6 – AI Model by Kaito-F | AlphaNeural AI
You can deploy this model and start earning money today!
Kaito-F
/
qwen3-4b-rlvmr-v22-cycle6
like
0
peft
safetensors
qwen3
lora
agent
reinforcement-learning
rlvmr
grpo-mr
alfworld
online-cycle-6
text-generation
conversational
en
Kaito-F/qwen3-4b-rlvmr-v22-cycle5
adapter
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
qwen3-4b-rlvmr-v8-cycle6
LoRA adapter
trained with
RLVMR (GRPO-MR)
- Online Cycle 6. Base:
Kaito-F/qwen3-4b-rlvmr-v22-cycle5
Method: Online Iterative RLVMR
Cycle 6: 12 tasks × K=8
Rolling reference: ref_log_probs from rollout model (KL=0 at start)
GRPO-MR advantage: α=0.5
Rollout Statistics
Trajectories: 80, Success: 28.7%
Task
Success
put
4/8 (50.0%)
clean
2/8 (25.0%)
heat
5/16 (31.2%)
cool
3/16 (18.8%)
examine
3/16 (18.8%)
puttwo
6/16 (37.5%)
Training
LoRA: r=64, α=128, all layers
LR: 2e-06, PPO clip: 0.2
KL: 0.01, Grad clip: 1.0