The code of Qwen2.5-VL has been in the latest Hugging face transformers and we advise you to build from source with command:
Qwen2.5-VL offers a toolkit to help you handle various types of visual input more conveniently, as if you were using an API. This includes base64, URLs, and interleaved images and videos. You can install it using the following command:
1# It's highly recommanded to use `[decord]` feature for faster video loading.
2pip install qwen-vl-utils[decord]==0.0.8
If you are not using Linux, you might not be able to install
decord from PyPI. In that case, you can use
pip install qwen-vl-utils which will fall back to using torchvision for video processing. However, you can still
install decord from source to get decord used when loading video.
1from transformers import Qwen2_5_VLForConditionalGeneration, AutoTokenizer, AutoProcessor
2from qwen_vl_utils import process_vision_info
3
4# default: Load the model on the available device(s)
5model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
6 "AdaptLLM/food-Qwen2.5-VL-3B-Instruct", torch_dtype="auto", device_map="auto"
7)
8
9# We recommend enabling flash_attention_2 for better acceleration and memory saving, especially in multi-image and video scenarios.
10# model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
11# "AdaptLLM/food-Qwen2.5-VL-3B-Instruct",
12# torch_dtype=torch.bfloat16,
13# attn_implementation="flash_attention_2",
14# device_map="auto",
15# )
16
17# default processer
18processor = AutoProcessor.from_pretrained("AdaptLLM/food-Qwen2.5-VL-3B-Instruct")
19
20# The default range for the number of visual tokens per image in the model is 4-16384.
21# You can set min_pixels and max_pixels according to your needs, such as a token range of 256-1280, to balance performance and cost.
22# min_pixels = 256*28*28
23# max_pixels = 1280*28*28
24# processor = AutoProcessor.from_pretrained("AdaptLLM/food-Qwen2.5-VL-3B-Instruct", min_pixels=min_pixels, max_pixels=max_pixels)
25
26messages = [
27 {
28 "role": "user",
29 "content": [
30 {
31 "type": "image",
32 "image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg",
33 },
34 {"type": "text", "text": "Describe this image."},
35 ],
36 }
37]
38
39# Preparation for inference
40text = processor.apply_chat_template(
41 messages, tokenize=False, add_generation_prompt=True
42)
43image_inputs, video_inputs = process_vision_info(messages)
44inputs = processor(
45 text=[text],
46 images=image_inputs,
47 videos=video_inputs,
48 padding=True,
49 return_tensors="pt",
50)
51inputs = inputs.to("cuda")
52
53# Inference: Generation of the output
54generated_ids = model.generate(**inputs, max_new_tokens=128)
55generated_ids_trimmed = [
56 out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
57]
58output_text = processor.batch_decode(
59 generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
60)
61print(output_text)