Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen3-4b-rlvmr-v22-cycle2 – AI Model by Kaito-F | AlphaNeural AI
You can deploy this model and start earning money today!
Kaito-F
/
qwen3-4b-rlvmr-v22-cycle2
like
0
peft
safetensors
qwen3
lora
agent
reinforcement-learning
rlvmr
grpo-mr
alfworld
online-cycle-2
text-generation
conversational
en
Kaito-F/qwen3-4b-rlvmr-v22-cycle1
adapter
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
qwen3-4b-rlvmr-v8-cycle2
LoRA adapter
trained with
RLVMR (GRPO-MR)
- Online Cycle 2. Base:
Kaito-F/qwen3-4b-rlvmr-v22-cycle1
Method: Online Iterative RLVMR
Cycle 2: 12 tasks × K=6
Rolling reference: ref_log_probs from rollout model (KL=0 at start)
GRPO-MR advantage: α=0.5
Rollout Statistics
Trajectories: 54, Success: 38.9%
Task
Success
put
3/6 (50.0%)
clean
4/12 (33.3%)
heat
2/6 (33.3%)
cool
2/6 (33.3%)
examine
2/12 (16.7%)
puttwo
8/12 (66.7%)
Training
LoRA: r=64, α=128, all layers
LR: 2e-06, PPO clip: 0.2
KL: 0.01, Grad clip: 1.0