Views
No views yet
TextVQA, DocVQA.| Model name | cogvlm2-llama3-chat-19B | cogvlm2-llama3-chinese-chat-19B |
|---|---|---|
| Base Model | Meta-Llama-3-8B-Instruct | Meta-Llama-3-8B-Instruct |
| Language | English | Chinese, English |
| Model size | 19B | 19B |
| Task | Image understanding, dialogue model | Image understanding, dialogue model |
| Text length | 8K | 8K |
| Image resolution | 1344 * 1344 | 1344 * 1344 |
| Model | Open Source | LLM Size | TextVQA | DocVQA | ChartQA | OCRbench | VCR_EASY | VCR_HARD | MMMU | MMVet | MMBench |
|---|---|---|---|---|---|---|---|---|---|---|---|
| CogVLM1.1 | ✅ | 7B | 69.7 | - | 68.3 | 590 | 73.9 | 34.6 | 37.3 | 52.0 | 65.8 |
| LLaVA-1.5 | ✅ | 13B | 61.3 | - | - | 337 | - | - | 37.0 | 35.4 | 67.7 |
| Mini-Gemini | ✅ | 34B | 74.1 | - | - | - | - | - | 48.0 | 59.3 | 80.6 |
| LLaVA-NeXT-LLaMA3 | ✅ | 8B | - | 78.2 | 69.5 | - | - | - | 41.7 | - | 72.1 |
| LLaVA-NeXT-110B | ✅ | 110B | - | 85.7 | 79.7 | - | - | - | 49.1 | - | 80.5 |
| InternVL-1.5 | ✅ | 20B | 80.6 | 90.9 | 83.8 | 720 | 14.7 | 2.0 | 46.8 | 55.4 | 82.3 |
| QwenVL-Plus | ❌ | - | 78.9 | 91.4 | 78.1 | 726 | - | - | 51.4 | 55.7 | 67.0 |
| Claude3-Opus | ❌ | - | - | 89.3 | 80.8 | 694 | 63.85 | 37.8 | 59.4 | 51.7 | 63.3 |
| Gemini Pro 1.5 | ❌ | - | 73.5 | 86.5 | 81.3 | - | 62.73 | 28.1 | 58.5 | - | - |
| GPT-4V | ❌ | - | 78.0 | 88.4 | 78.5 | 656 | 52.04 | 25.8 | 56.8 | 67.7 | 75.0 |
| CogVLM2-LLaMA3 | ✅ | 8B | 84.2 | 92.3 | 81.0 | 756 | 83.3 | 38.0 | 44.3 | 60.4 | 80.5 |
| CogVLM2-LLaMA3-Chinese | ✅ | 8B | 85.0 | 88.4 | 74.7 | 780 | 79.9 | 25.1 | 42.8 | 60.5 | 78.9 |
1import torch
2from PIL import Image
3from transformers import AutoModelForCausalLM, AutoTokenizer
4
5MODEL_PATH = "THUDM/cogvlm2-llama3-chat-19B"
6DEVICE = 'cuda' if torch.cuda.is_available() else 'cpu'
7TORCH_TYPE = torch.bfloat16 if torch.cuda.is_available() and torch.cuda.get_device_capability()[0] >= 8 else torch.float16
8
9tokenizer = AutoTokenizer.from_pretrained(
10 MODEL_PATH,
11 trust_remote_code=True
12)
13model = AutoModelForCausalLM.from_pretrained(
14 MODEL_PATH,
15 torch_dtype=TORCH_TYPE,
16 trust_remote_code=True,
17).to(DEVICE).eval()
18
19text_only_template = "A chat between a curious user and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the user's questions. USER: {} ASSISTANT:"
20
21while True:
22 image_path = input("image path >>>>> ")
23 if image_path == '':
24 print('You did not enter image path, the following will be a plain text conversation.')
25 image = None
26 text_only_first_query = True
27 else:
28 image = Image.open(image_path).convert('RGB')
29
30 history = []
31
32 while True:
33 query = input("Human:")
34 if query == "clear":
35 break
36
37 if image is None:
38 if text_only_first_query:
39 query = text_only_template.format(query)
40 text_only_first_query = False
41 else:
42 old_prompt = ''
43 for _, (old_query, response) in enumerate(history):
44 old_prompt += old_query + " " + response + "\n"
45 query = old_prompt + "USER: {} ASSISTANT:".format(query)
46 if image is None:
47 input_by_model = model.build_conversation_input_ids(
48 tokenizer,
49 query=query,
50 history=history,
51 template_version='chat'
52 )
53 else:
54 input_by_model = model.build_conversation_input_ids(
55 tokenizer,
56 query=query,
57 history=history,
58 images=[image],
59 template_version='chat'
60 )
61 inputs = {
62 'input_ids': input_by_model['input_ids'].unsqueeze(0).to(DEVICE),
63 'token_type_ids': input_by_model['token_type_ids'].unsqueeze(0).to(DEVICE),
64 'attention_mask': input_by_model['attention_mask'].unsqueeze(0).to(DEVICE),
65 'images': [[input_by_model['images'][0].to(DEVICE).to(TORCH_TYPE)]] if image is not None else None,
66 }
67 gen_kwargs = {
68 "max_new_tokens": 2048,
69 "pad_token_id": 128002,
70 }
71 with torch.no_grad():
72 outputs = model.generate(**inputs, **gen_kwargs)
73 outputs = outputs[:, inputs['input_ids'].shape[1]:]
74 response = tokenizer.decode(outputs[0])
75 response = response.split("<|end_of_text|>")[0]
76 print("\nCogVLM2:", response)
77 history.append((query, response))@misc{hong2024cogvlm2,
title={CogVLM2: Visual Language Models for Image and Video Understanding},
author={Hong, Wenyi and Wang, Weihan and Ding, Ming and Yu, Wenmeng and Lv, Qingsong and Wang, Yan and Cheng, Yean and Huang, Shiyu and Ji, Junhui and Xue, Zhao and others},
year={2024}
eprint={2408.16500},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
@misc{wang2023cogvlm,
title={CogVLM: Visual Expert for Pretrained Language Models},
author={Weihan Wang and Qingsong Lv and Wenmeng Yu and Wenyi Hong and Ji Qi and Yan Wang and Junhui Ji and Zhuoyi Yang and Lei Zhao and Xixuan Song and Jiazheng Xu and Bin Xu and Juanzi Li and Yuxiao Dong and Ming Ding and Jie Tang},
year={2023},
eprint={2311.03079},
archivePrefix={arXiv},
primaryClass={cs.CV}
}