Views
No views yet
Qwen/Qwen3-VL-4B-Instruct.
This is the on-policy DPO stage of BirdAgent (see the flagship
Chinzhu/BirdAgent-Qwen3VL-4B
GSPO model for the full description, tools, results, and figures).r=64, alpha=128, dropout=0.05 (trained bf16, adapter
saved fp32); targets language-model q/k/v/o/gate/up/down_proj.1@inproceedings{wang2026birdagent,
2 title = {BirdAgent: A Small Vision--Language Model that Orchestrates
3 Domain Tools Beats Large Models that Merely Hold Them},
4 author = {Wang, Xinzhu},
5 booktitle = {Under review},
6 year = {2026}
7}