Model Description
AgenticQwen 30B A3B is a small agentic language model trained on Qwen3 30B A3B Instruct, designed for multi-step reasoning and tool use. It is trained with a multi-round reinforcement learning (GRPO-style) pipeline and a dual "data flywheel" mechanism that continually increases task difficulty for both reasoning and agentic workflows.
Note
For best benchmark performance, we recommend using the same (or highly similar) prompting format as used during training, as described in the accompanying pa