Speech Large Language Models (Speech-LLMs) have emerged as a powerful approach for automatic speech recognition (ASR) by aligning speech encoders with large language models. However, adapting these systems to multilingual settings with imbalanced data distributions remains challenging.
Experiments on a 12-language mixed-resource setting show that Zipper-LoRA consistently outperforms both fully shared and independent baselines, particularly in extremely low-resource scenarios.
1@article{ZipperLoRA2026,
2 title={Dynamic Parameter Decoupling for Speech-LLM based Multilingual Speech Recognition},
3 author={Mei, Yuxiang and Qiu, Delai and Liu, Shengping and Liang, Jiaen and Long, Yanhua},
4 journal={arXiv preprint arXiv:2603.17558},
5 year={2026}
6}