Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen3-4b-grpo-v5 – AI Model by Kaito-F | AlphaNeural AI
You can deploy this model and start earning money today!
Kaito-F
/
qwen3-4b-grpo-v5
like
0
transformers
safetensors
qwen3
text-generation
grpo
reinforcement-learning
agent
tool-use
alfworld
dbbench
conversational
en
u-10bei/sft_alfworld_trajectory_dataset_v5
u-10bei/dbbench_sft_dataset_react_v4
u-10bei/dbbench_sft_dataset_react_v3
Kaito-F/qwen3-4b-grpo-v4
finetune
apache-2.0
text-generation-inference
endpoints_compatible
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
GRPO-tuned Agent Model (v2)
This model is fine-tuned from
Kaito-F/qwen3-4b-grpo-v4
using
GRPO (Group Relative Policy Optimization)
with Unsloth.
Training Details
Method
: GRPO with 4-bit quantized training + LoRA
Base model
: Kaito-F/qwen3-4b-grpo-v4 (SFT-tuned)
Learning rate
: 1e-5
LoRA
: r=32, alpha=64
KL coefficient (beta)
: 0.04
Group size
: 8
Gradient accumulation
: 16
Temperature
: 0.9
Max steps
: 100
Dataset (2,000 samples)
DBBench: 1,000 samples
(v4 + v3 combined)
aggregation-MAX: ~250 (top priority, previously 0% accuracy)
aggregation-MIN: ~200
aggregation-SUM/AVG/COUNT: ~100
comparison: ~200
counting: ~150
ranking/other/insert: remainder
ALFWorld: 1,000 samples
Reward Functions (v2: Enhanced)
Format Reward
(max +1.01): Additive scoring for Thought/Action + micro-jitter to prevent reward_std=0
ALFWorld Reward
: Task completion (+2.5), invalid action/loop (-1.0)
DBBench Reward
(max +3.3):
SQL syntax validity: +0.5
Intent keyword match (MAX/MIN/COUNT etc.): up to +0.3
Final answer match: +2.0
SQL exact match: +0.5
Conciseness Reward
(max +0.3): Linear decay per action step
Sources & Terms
ALFWorld dataset: u-10bei/sft_alfworld_trajectory_dataset_v5
DBBench datasets: u-10bei/dbbench_sft_dataset_react_v4, u-10bei/dbbench_sft_dataset_react_v3
License: Apache 2.0