This is the
MambaMia Mini model based on
lmsys/vicuna-7b-v1.5, designed for efficient long-form video understanding.
MambaMia is a State-Space-Model-based hierarchical compression method for efficient video understanding in Large Multimodal Models (LMMs). It addresses the computational cost and information redundancy challenges in processing long videos.
Please refer to the
MambaMia repository for detailed usage instructions.
1@misc{kim2025mambamia,
2 title={MambaMia: A State-Space-Model-Based Compression for Efficient Video Understanding in Large Multimodal Models},
3 author={Geewook Kim and Minjoon Seo},
4 year={2025},
5 eprint={2506.13564},
6 archivePrefix={arXiv},
7 primaryClass={cs.CV},
8 url={https://arxiv.org/abs/2506.13564}
9}
This model is released under the
Llama 2 Community License Agreement, following the base model's license.