Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen3-4b-rlvmr-v22-cycle4 – AI Model by Kaito-F | AlphaNeural AI
You can deploy this model and start earning money today!
Kaito-F
/
qwen3-4b-rlvmr-v22-cycle4
like
0
peft
safetensors
qwen3
lora
agent
reinforcement-learning
rlvmr
grpo-mr
alfworld
online-cycle-4
text-generation
conversational
en
Kaito-F/qwen3-4b-rlvmr-v22-cycle3
adapter
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
qwen3-4b-rlvmr-v8-cycle4
LoRA adapter
trained with
RLVMR (GRPO-MR)
- Online Cycle 4. Base:
Kaito-F/qwen3-4b-rlvmr-v22-cycle3
Method: Online Iterative RLVMR
Cycle 4: 12 tasks × K=8
Rolling reference: ref_log_probs from rollout model (KL=0 at start)
GRPO-MR advantage: α=0.5
Rollout Statistics
Trajectories: 80, Success: 60.0%
Task
Success
put
7/8 (87.5%)
clean
12/16 (75.0%)
heat
4/8 (50.0%)
cool
8/16 (50.0%)
examine
11/16 (68.8%)
puttwo
6/16 (37.5%)
Training
LoRA: r=64, α=128, all layers
LR: 2e-06, PPO clip: 0.2
KL: 0.01, Grad clip: 1.0