-
Can native VLMs generalize across single-image, multi-image, Video, and 3D spatial scenarios?
-
What advantages of native VLMs, especially early-fusion for pixel-pixel,pixel-word?
-
How to build strong native VLMs over Qwen3-VL for subsequent RL community?
-
Model Type: Native Vision-Language Models
-
Model Mode: Mixed Native-Attn & Native-RoPE
-
Layer Parameters: 56M vs. 50M (Qwen3-1.7B)
-
Model Parameters: 2.2B (Non-Embedding)
-
Number of Layers: 40 (12 for Pre-Buffer & 28 for Post-LLM)
-
Number of Heads: 16 for Q and 8 for KV (GQA)
-
Head Dimensions: 128 * 2 for QK and 128 for V
1@article{Diao2026NEOov,
2 title = {From Pixels to Words--Towards Native One-Vision Models at Scale},
3 author = {Diao, Haiwen and Wang, Jiahao and Wu, Penghao and Dong, Yuhao and Niu, Yuwei and Zhu, Yue and Cai, Zhongang and Fan, Weichen and Dai, Linjun and Wu, Silei and others},
4 journal = {arXiv preprint arXiv:2605.28820},
5 year = {2026}
6}