Views
No views yet

Qwen3-VL-32B-Instruct-v2-MAX-MXFP4 is an optimized and compressed evolution built on top of Qwen/Qwen3-VL-32B-Instruct, designed for advanced multimodal understanding and high-detail image captioning. This variant leverages BF16 · U8 tensor formats to significantly reduce memory footprint and improve inference efficiency while maintaining strong output quality. The model incorporates a more optimized abliteration rate, combining refined refusal direction analysis with enhanced training strategies to minimize internal refusal behaviors while preserving strong reasoning, instruction-following, and visual understanding capabilities. The result is a powerful 32B parameter vision-language model optimized for highly detailed captions, deep scene understanding, and rich multimodal reasoning, now with efficient deployment characteristics.
[!IMPORTANT] This model is intended for research and learning purposes only. Due to reduced internal refusal mechanisms, it may generate sensitive or unrestricted content. Users assume full responsibility for how the model is used. The authors and hosting platform disclaim any liability for generated outputs.
1pip install transformers==5.4.0
2# or
3pip install git+https://github.com/huggingface/transformers.git1from transformers import Qwen3_5ForConditionalGeneration, AutoProcessor
2import torch
3
4model = Qwen3_5ForConditionalGeneration.from_pretrained(
5 "prithivMLmods/Qwen3-VL-32B-Instruct-v2-MAX-MXFP4",
6 torch_dtype="auto",
7 device_map="auto"
8)
9
10processor = AutoProcessor.from_pretrained(
11 "prithivMLmods/Qwen3-VL-32B-Instruct-v2-MAX-MXFP4"
12)
13
14messages = [
15 {
16 "role": "user",
17 "content": [
18 {"type": "text", "text": "Describe this image in extreme detail."}
19 ],
20 }
21]
22
23text = processor.apply_chat_template(
24 messages, tokenize=False, add_generation_prompt=True
25)
26
27inputs = processor(
28 text=[text],
29 padding=True,
30 return_tensors="pt"
31).to("cuda")
32
33generated_ids = model.generate(**inputs, max_new_tokens=512)
34
35generated_ids_trimmed = [
36 out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
37]
38
39output_text = processor.batch_decode(
40 generated_ids_trimmed,
41 skip_special_tokens=True,
42 clean_up_tokenization_spaces=False
43)
44
45print(output_text)Important Note: This model intentionally minimizes built-in safety refusals.