Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen3-4b-grpo-v1 – AI Model by Kaito-F | AlphaNeural AI
You can deploy this model and start earning money today!
Kaito-F
/
qwen3-4b-grpo-v1
like
0
transformers
safetensors
qwen3
text-generation
grpo
reinforcement-learning
agent
tool-use
alfworld
dbbench
conversational
en
u-10bei/sft_alfworld_trajectory_dataset_v5
u-10bei/dbbench_sft_dataset_react_v4
Kaito-F/qwen3-4b-instruct-lora-v2
finetune
apache-2.0
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
GRPO-tuned Agent Model
This model is fine-tuned from
Kaito-F/qwen3-4b-instruct-lora-v2
using
GRPO (Group Relative Policy Optimization)
with Unsloth.
Training Details
Method
: GRPO with 4-bit quantized training + LoRA
Base model
: Kaito-F/qwen3-4b-instruct-lora-v2 (SFT-tuned)
Learning rate
: 5e-6
LoRA
: r=32, alpha=64
KL coefficient (beta)
: 0.04
Group size
: 8
Max steps
: 200
Reward Functions
Format Reward
: ReAct format (Thought/Action) compliance
ALFWorld Reward
: Task completion (+2.0), invalid action (-0.5)
DBBench Reward
: SQL syntax (+0.5), result match (+1.5)
Conciseness Reward
: Fewer steps = higher reward
Dataset Mix
DBBench: 70% (aggregation, counting, comparison prioritized)
ALFWorld: 30%
Sources & Terms
ALFWorld dataset: u-10bei/sft_alfworld_trajectory_dataset_v5
DBBench dataset: u-10bei/dbbench_sft_dataset_react_v4
License: Apache 2.0