Views
No views yet


| Model Name | Vision Part | Language Part | HF Link |
|---|---|---|---|
| InternVL2_5-1B-MPO | InternViT-300M-448px-V2_5 | Qwen2.5-0.5B-Instruct | 🤗 link |
| InternVL2_5-2B-MPO | InternViT-300M-448px-V2_5 | internlm2_5-1_8b-chat | 🤗 link |
| InternVL2_5-4B-MPO | InternViT-300M-448px-V2_5 | Qwen2.5-3B-Instruct | 🤗 link |
| InternVL2_5-8B-MPO | InternViT-300M-448px-V2_5 | internlm2_5-7b-chat | 🤗 link |
| InternVL2_5-26B-MPO | InternViT-6B-448px-V2_5 | internlm2_5-20b-chat | 🤗 link |
| InternVL2_5-38B-MPO | InternViT-6B-448px-V2_5 | Qwen2.5-32B-Instruct | 🤗 link |
| InternVL2_5-78B-MPO | InternViT-6B-448px-V2_5 | Qwen2.5-72B-Instruct | 🤗 link |



Final Answer: ***.
Responses matching the ground truth answer constitute the positive set \(\mathcal{Y}_p\), while those that do not match make up the negative set \(\mathcal{Y}_n\). Additionally, responses that fail to provide a clear final answer are also merged into \(\mathcal{Y}_n\).
Given these responses labeled as positive or negative, we build the preference pairs by selecting a chosen response \(y_c\) from \(\mathcal{Y}_p\) and a negative response \(y_r\) from \(\mathcal{Y}_n\).| Model | Avg. | MMBench v1.1 | MMStar | MMMU | MathVista | HallusionBench | AI2D | OCRBench | MMVet |
|---|---|---|---|---|---|---|---|---|---|
| InternVL2-5-1B | 54.9 | 66.5 | 51.3 | 41.2 | 47.1 | 39.4 | 69.0 | 77.4 | 47.2 |
| InternVL2-5-1B-MPO | 56.4 | 67.2 | 49.7 | 40.8 | 53.0 | 40.0 | 69.4 | 83.6 | 47.2 |
| InternVL2-5-2B | 59.9 | 70.9 | 54.3 | 43.2 | 51.1 | 42.3 | 74.9 | 80.2 | 62.6 |
| InternVL2-5-2B-MPO | 62.0 | 71.6 | 55.0 | 45.0 | 56.4 | 43.0 | 75.3 | 84.2 | 65.4 |
| InternVL2-5-4B | 65.1 | 78.2 | 58.7 | 51.8 | 60.8 | 46.6 | 81.4 | 82.0 | 61.5 |
| InternVL2-5-4B-MPO | 67.6 | 78.6 | 60.2 | 51.6 | 65.3 | 47.8 | 82.0 | 88.0 | 67.1 |
| InternVL2-5-8B | 68.9 | 82.5 | 63.2 | 56.2 | 64.5 | 49.0 | 84.6 | 82.1 | 62.8 |
| InternVL2-5-8B-MPO | 70.4 | 82.4 | 65.7 | 54.9 | 68.9 | 51.4 | 84.5 | 88.3 | 66.9 |
| InternVL2-5-26B | 71.6 | 84.6 | 66.5 | 60.7 | 68.0 | 55.8 | 86.2 | 85.4 | 65.4 |
| InternVL2-5-26B-MPO | 72.7 | 84.2 | 67.2 | 57.7 | 72.8 | 55.3 | 86.2 | 91.2 | 67.1 |
| InternVL2-5-38B | 73.5 | 85.4 | 68.5 | 64.6 | 72.4 | 57.9 | 87.6 | 84.1 | 67.2 |
| InternVL2-5-38B-MPO | 75.5 | 85.6 | 69.8 | 64.1 | 73.8 | 61.5 | 88.1 | 88.5 | 72.5 |
| InternVL2-5-78B | 75.2 | 87.5 | 69.5 | 70.0 | 70.6 | 57.4 | 89.1 | 85.3 | 71.8 |
| InternVL2-5-78B-MPO | 76.6 | 87.3 | 73.1 | 68.3 | 73.8 | 58.7 | 89.3 | 91.2 | 71.4 |
pip install lmdeploy>=0.6.41from lmdeploy import pipeline, TurbomindEngineConfig
2from lmdeploy.vl import load_image
3
4model = 'OpenGVLab/InternVL2_5-26B-MPO-AWQ'
5image = load_image('https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/tests/data/tiger.jpeg')
6pipe = pipeline(model, backend_config=TurbomindEngineConfig(session_len=8192))
7response = pipe(('describe this image', image))
8print(response.text)ImportError occurs while executing this case, please install the required dependency packages as prompted.1from lmdeploy import pipeline, TurbomindEngineConfig
2from lmdeploy.vl import load_image
3from lmdeploy.vl.constants import IMAGE_TOKEN
4
5model = 'OpenGVLab/InternVL2_5-26B-MPO-AWQ'
6pipe = pipeline(model, backend_config=TurbomindEngineConfig(session_len=8192))
7
8image_urls=[
9 'https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg',
10 'https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/det.jpg'
11]
12
13images = [load_image(img_url) for img_url in image_urls]
14# Numbering images improves multi-image conversations
15response = pipe((f'Image-1: {IMAGE_TOKEN}\nImage-2: {IMAGE_TOKEN}\ndescribe these two images', images))
16print(response.text)1from lmdeploy import pipeline, TurbomindEngineConfig
2from lmdeploy.vl import load_image
3
4model = 'OpenGVLab/InternVL2_5-26B-MPO-AWQ'
5pipe = pipeline(model, backend_config=TurbomindEngineConfig(session_len=8192))
6
7image_urls=[
8 "https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg",
9 "https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/det.jpg"
10]
11prompts = [('describe this image', load_image(img_url)) for img_url in image_urls]
12response = pipe(prompts)
13print(response)pipeline.chat interface.1from lmdeploy import pipeline, TurbomindEngineConfig, GenerationConfig
2from lmdeploy.vl import load_image
3
4model = 'OpenGVLab/InternVL2_5-26B-MPO-AWQ'
5pipe = pipeline(model, backend_config=TurbomindEngineConfig(session_len=8192))
6
7image = load_image('https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg')
8gen_config = GenerationConfig(top_k=40, top_p=0.8, temperature=0.8)
9sess = pipe.chat(('describe this image', image), gen_config=gen_config)
10print(sess.response.text)
11sess = pipe.chat('What is the woman doing?', session=sess, gen_config=gen_config)
12print(sess.response.text)api_server enables models to be easily packed into services with a single command. The provided RESTful APIs are compatible with OpenAI's interfaces. Below are an example of service startup:lmdeploy serve api_server OpenGVLab/InternVL2_5-26B-MPO-AWQ --server-port 23333pip install openai1from openai import OpenAI
2
3client = OpenAI(api_key='YOUR_API_KEY', base_url='http://0.0.0.0:23333/v1')
4model_name = client.models.list().data[0].id
5response = client.chat.completions.create(
6 model=model_name,
7 messages=[{
8 'role':
9 'user',
10 'content': [{
11 'type': 'text',
12 'text': 'describe this image',
13 }, {
14 'type': 'image_url',
15 'image_url': {
16 'url':
17 'https://modelscope.oss-cn-beijing.aliyuncs.com/resource/tiger.jpeg',
18 },
19 }],
20 }],
21 temperature=0.8,
22 top_p=0.8)
23print(response)1@article{wang2024mpo,
2 title={Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization},
3 author={Wang, Weiyun and Chen, Zhe and Wang, Wenhai and Cao, Yue and Liu, Yangzhou and Gao, Zhangwei and Zhu, Jinguo and Zhu, Xizhou and Lu, Lewei and Qiao, Yu and Dai, Jifeng},
4 journal={arXiv preprint arXiv:2411.10442},
5 year={2024}
6}
7@article{chen2024expanding,
8 title={Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling},
9 author={Chen, Zhe and Wang, Weiyun and Cao, Yue and Liu, Yangzhou and Gao, Zhangwei and Cui, Erfei and Zhu, Jinguo and Ye, Shenglong and Tian, Hao and Liu, Zhaoyang and others},
10 journal={arXiv preprint arXiv:2412.05271},
11 year={2024}
12}
13@article{chen2024far,
14 title={How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites},
15 author={Chen, Zhe and Wang, Weiyun and Tian, Hao and Ye, Shenglong and Gao, Zhangwei and Cui, Erfei and Tong, Wenwen and Hu, Kongzhi and Luo, Jiapeng and Ma, Zheng and others},
16 journal={arXiv preprint arXiv:2404.16821},
17 year={2024}
18}
19@inproceedings{chen2024internvl,
20 title={Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks},
21 author={Chen, Zhe and Wu, Jiannan and Wang, Wenhai and Su, Weijie and Chen, Guo and Xing, Sen and Zhong, Muyan and Zhang, Qinglong and Zhu, Xizhou and Lu, Lewei and others},
22 booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
23 pages={24185--24198},
24 year={2024}
25}