Views
No views yet
d_kv dimensions:
LLaVA-NeXT-8B-MLA-stage2-rope32-d_kv_32LLaVA-NeXT-8B-MLA-stage2-rope32-d_kv_64LLaVA-NeXT-8B-MLA-stage2-rope32-d_kv_1281@misc{fan2026mha2mlavlmenablingdeepseekseconomical,
2 title={MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models},
3 author={Xiaoran Fan and Zhichao Sun and Tao Ji and Lixing Shen and Tao Gui},
4 year={2026},
5 eprint={2601.11464},
6 archivePrefix={arXiv},
7 primaryClass={cs.CV},
8 url={https://arxiv.org/abs/2601.11464},
9}