Views
No views yet
cnxup/Qwen2.5-VL-7B-MLA-stage1-rope32d_kv_128 as an example.lmms-eval for benchmarking. Detailed instructions can be found in the Evaluation section.1cd eval
2cd qwen2_5_vl
3sh eval.sh1@misc{fan2026mha2mlavlmenablingdeepseekseconomical,
2 title={MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models},
3 author={Xiaoran Fan and Zhichao Sun and Tao Ji and Lixing Shen and Tao Gui},
4 year={2026},
5 eprint={2601.11464},
6 archivePrefix={arXiv},
7 primaryClass={cs.CV},
8 url={https://arxiv.org/abs/2601.11464},
9}