Views
No views yet
[!TIP] Multi Token Prediction (MTP) weights restored from the Qwen3.5-0.8B model.
./conversion/minicpm.py file in the llama.cpp directiory and delete or comment out the following code section: # MTP tensors are not used at inference yet; align with Qwen3Next behaviour
if name.startswith("mtp"):
return None| Parameter | Value |
|---|---|
| start_layer_index | 9 |
| end_layer_index | 17 |
| preserve_good_behavior_weight | 0.8947 |
| steer_bad_behavior_weight | 0.0002 |
| overcorrect_relative_weight | 1.0095 |
| neighbor_count | 8 |
| Metric | This model | Original model (openbmb/MiniCPM-V-4.6-Thinking) |
|---|---|---|
| KL divergence | 0.0556 | 0 (by definition) |
| Refusals | 14/416 | 314/416 |
[Trial 369] Refusals: 0/416, KL divergence: 0.2261
[Trial 71] Refusals: 1/416, KL divergence: 0.2190
[Trial 290] Refusals: 2/416, KL divergence: 0.1109
[Trial 460] Refusals: 8/416, KL divergence: 0.0993
» [Trial 630] Refusals: 14/416, KL divergence: 0.0556
[Trial 615] Refusals: 20/416, KL divergence: 0.0419
[Trial 685] Refusals: 25/416, KL divergence: 0.0320
[Trial 477] Refusals: 38/416, KL divergence: 0.0315
[Trial 399] Refusals: 39/416, KL divergence: 0.0311
[Trial 130] Refusals: 61/416, KL divergence: 0.0254
[Trial 720] Refusals: 66/416, KL divergence: 0.0242
[Trial 763] Refusals: 87/416, KL divergence: 0.0192
[Trial 341] Refusals: 126/416, KL divergence: 0.0169
[Trial 447] Refusals: 141/416, KL divergence: 0.0142
[Trial 786] Refusals: 147/416, KL divergence: 0.0135
[Trial 408] Refusals: 153/416, KL divergence: 0.0114
[Trial 600] Refusals: 174/416, KL divergence: 0.0086
[Trial 128] Refusals: 212/416, KL divergence: 0.0085
[Trial 122] Refusals: 243/416, KL divergence: 0.0068
[Trial 419] Refusals: 301/416, KL divergence: 0.0011
[Trial 705] Refusals: 304/416, KL divergence: 0.0009
[Trial 330] Refusals: 307/416, KL divergence: 0.0008
[Trial 25] Refusals: 314/416, KL divergence: 0.0000┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━┓
┃ Benchmark ┃ Metric ┃ Value ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━┩
│ PIQA Base │ name │ piqa │
│ │ sample_len │ 1838 │
│ │ acc,none │ 0.7029 │
│ │ acc_stderr,none │ 0.0107 │
│ │ acc_norm,none │ 0.7198 │
│ │ acc_norm_stderr,none │ 0.0105 │
└───────────┴──────────────────────┴────────┘
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━┓
┃ Benchmark ┃ Metric ┃ Value ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━┩
│ PIQA T630 │ name │ piqa │
│ │ sample_len │ 1838 │
│ │ acc,none │ 0.7040 │
│ │ acc_stderr,none │ 0.0107 │
│ │ acc_norm,none │ 0.7187 │
│ │ acc_norm_stderr,none │ 0.0105 │
└───────────┴──────────────────────┴────────┘
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━┓
┃ Benchmark ┃ Metric ┃ Value ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━┩
│ PIQA T290 │ name │ piqa │
│ │ sample_len │ 1838 │
│ │ acc,none │ 0.7018 │
│ │ acc_stderr,none │ 0.0107 │
│ │ acc_norm,none │ 0.7160 │
│ │ acc_norm_stderr,none │ 0.0105 │
└───────────┴──────────────────────┴────────┘
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━┓
┃ Benchmark ┃ Metric ┃ Value ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━┩
│ PIQA T460 │ name │ piqa │
│ │ sample_len │ 1838 │
│ │ acc,none │ 0.7008 │
│ │ acc_stderr,none │ 0.0107 │
│ │ acc_norm,none │ 0.7176 │
│ │ acc_norm_stderr,none │ 0.0105 │
└───────────┴──────────────────────┴────────┘
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━┓
┃ Benchmark ┃ Metric ┃ Value ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━┩
│ PIQA T615 │ name │ piqa │
│ │ sample_len │ 1838 │
│ │ acc,none │ 0.7008 │
│ │ acc_stderr,none │ 0.0107 │
│ │ acc_norm,none │ 0.7187 │
│ │ acc_norm_stderr,none │ 0.0105 │
└───────────┴──────────────────────┴────────┘



| iPhone iPhone 17 Pro Max | Android Redmi K70 | HarmonyOS HUAWEI nova 14 |
![]() | ![]() | ![]() |
pip install "transformers[torch]>=5.7.0" torchvision torchcodecNote on CUDA compatibility:torchcodec(used for video decoding) may have compatibility issues with certain CUDA versions. For example,torch>=2.11bundles CUDA 13.1 by default, while environments with CUDA 12.x may encounter errors such asRuntimeError: Could not load libtorchcodec. Two workarounds:
- Replace
torchcodecwithPyAV— supports both image and video inference without CUDA version constraints:pip install "transformers[torch]>=5.7.0" torchvision av- Pin the CUDA version when installing torch to match your environment (e.g. CUDA 12.8):
pip install "transformers>=5.7.0" torchvision torchcodec --index-url https://download.pytorch.org/whl/cu128
1from transformers import AutoModelForImageTextToText, AutoProcessor
2
3model_id = "openbmb/MiniCPM-V-4.6-Thinking"
4
5processor = AutoProcessor.from_pretrained(model_id)
6model = AutoModelForImageTextToText.from_pretrained(
7 model_id, torch_dtype="auto", device_map="auto"
8)
9
10# Flash Attention 2 is recommended for better acceleration and memory saving,
11# especially in multi-image and video scenarios.
12# model = AutoModelForImageTextToText.from_pretrained(
13# model_id,
14# torch_dtype=torch.bfloat16,
15# attn_implementation="flash_attention_2",
16# device_map="auto",
17# )1messages = [
2 {
3 "role": "user",
4 "content": [
5 {"type": "image", "url": "https://huggingface.co/datasets/openbmb/DemoCase/resolve/main/refract.png"},
6 {"type": "text", "text": "What causes this phenomenon?"},
7 ],
8 }
9]
10
11downsample_mode = "16x" # Using `downsample_mode="4x"` for Finer Detail
12
13inputs = processor.apply_chat_template(
14 messages, tokenize=True, add_generation_prompt=True,
15 return_dict=True, return_tensors="pt",
16 downsample_mode=downsample_mode,
17 max_slice_nums=36,
18).to(model.device)
19
20generated_ids = model.generate(**inputs, downsample_mode=downsample_mode, max_new_tokens=512)
21generated_ids_trimmed = [
22 out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
23]
24output_text = processor.batch_decode(
25 generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
26)
27print(output_text[0])1messages = [
2 {
3 "role": "user",
4 "content": [
5 {"type": "video", "url": "https://huggingface.co/datasets/openbmb/DemoCase/resolve/main/football.mp4"},
6 {"type": "text", "text": "Describe this video in detail. Follow the timeline and focus on on-screen text, interface changes, main actions, and scene changes."},
7 ],
8 }
9]
10
11downsample_mode = "16x" # Using `downsample_mode="4x"` for Finer Detail
12
13inputs = processor.apply_chat_template(
14 messages, tokenize=True, add_generation_prompt=True,
15 return_dict=True, return_tensors="pt",
16 downsample_mode=downsample_mode,
17 max_num_frames=128,
18 stack_frames=1,
19 max_slice_nums=1,
20 use_image_id=False,
21).to(model.device)
22
23generated_ids = model.generate(**inputs, downsample_mode=downsample_mode, max_new_tokens=2048)
24generated_ids_trimmed = [
25 out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
26]
27output_text = processor.batch_decode(
28 generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
29)
30print(output_text[0])apply_chat_template:| Parameter | Default | Applies to | Description |
|---|---|---|---|
downsample_mode | "16x" | Image & Video | Visual token downsampling. "16x" merges tokens for efficiency; "4x" keeps 4× more tokens for finer detail. Must also be passed to generate(). |
max_slice_nums | 9 | Image & Video | Maximum number of slices when splitting a high-resolution image. Higher values preserve more detail for large images. Recommended: 36 for image, 1 for video. |
max_num_frames | 128 | Video only | Maximum number of main frames sampled from the video. |
stack_frames | 1 | Video only | Total sample points per second. 1 = main frame only (no stacking). N (N>1) = 1 main frame + N−1 sub-frames per second; the sub-frames are composited into a grid image and interleaved with main frames. Recommended: 3 or 5. |
use_image_id | True | Image & Video | Whether to prepend <image_id>N</image_id> tags before each image/frame placeholder. Recommended: True for image, False for video. |
Note:downsample_modemust be passed to bothapply_chat_template(for correct placeholder count) andgenerate(for the vision encoder). All other parameters only need to be passed toapply_chat_template.
transformers serve pip install "transformers[serving]>=5.7.0"transformers serve openbmb/MiniCPM-V-4.6-Thinking --port 8000 --host 0.0.0.0 --continuous-batching1curl -s http://localhost:8000/v1/chat/completions \
2 -H 'Content-Type: application/json' \
3 -d '{
4 "model": "openbmb/MiniCPM-V-4.6-Thinking",
5 "messages": [{
6 "role": "user",
7 "content": [
8 {"type": "image_url", "image_url": {"url": "https://huggingface.co/datasets/openbmb/DemoCase/resolve/main/refract.png"}},
9 {"type": "text", "text": "What causes this phenomenon?"}
10 ]
11 }]
12 }'\n as string literals instead of actual newlines. To render the text correctly, especially in UI layers, you can use the following utility function. This function carefully replaces literal \n with real newlines while protecting scenarios where \n has specific semantic meaning.1import re
2
3_PATTERN = re.compile(
4 r'(```[\s\S]*?```' # fenced code blocks
5 r'|`[^`]+`' # inline code
6 r'|\$\$[\s\S]*?\$\$' # display math
7 r'|\$[^$]+\$' # inline math
8 r'|\\\([\s\S]*?\\\)' # \(...\)
9 r'|\\\[[\s\S]*?\\\]' # \[...\]
10 r')'
11 r'|(?<!\\)(?:\\r\\n|\\[nr])'
12)
13
14def normalize_response_text(text: str) -> str:
15 """
16 Lightweight post-processing: Converts literal '\\n' to actual newlines,
17 while protecting code blocks, inline code, and LaTeX commands.
18 """
19 if not isinstance(text, str) or "\\" not in text:
20 return text
21 return _PATTERN.sub(lambda m: m.group(1) or '\n', text)1vllm serve openbmb/MiniCPM-V-4.6-Thinking \
2 --port 8000 \
3 --enable-auto-tool-choice \
4 --tool-call-parser qwen3_coder \
5 --default-chat-template-kwargs '{"enable_thinking": true}'Note:--enable-auto-tool-choiceand--tool-call-parser qwen3_coderenable tool/function calling support. If you don't need tool use, you can omit these flags and simply runvllm serve openbmb/MiniCPM-V-4.6-Thinking.
1curl -s http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
2 "model": "openbmb/MiniCPM-V-4.6-Thinking",
3 "messages": [{"role": "user", "content": [
4 {"type": "image_url", "image_url": {"url": "https://huggingface.co/datasets/openbmb/DemoCase/resolve/main/refract.png"}},
5 {"type": "text", "text": "What causes this phenomenon?"}
6 ]}]
7}'1curl -s http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
2 "model": "openbmb/MiniCPM-V-4.6-Thinking",
3 "messages": [{"role": "user", "content": [
4 {"type": "text", "text": "北京的天气"}
5 ]}],
6 "tools": [{
7 "type": "function",
8 "function": {
9 "name": "get_weather",
10 "description": "Get the current weather for a given location",
11 "parameters": {
12 "type": "object",
13 "properties": {
14 "location": {"type": "string", "description": "City name"}
15 },
16 "required": ["location"]
17 }
18 }
19 }]
20}'python -m sglang.launch_server --model openbmb/MiniCPM-V-4.6-Thinking --port 300001curl -s http://localhost:30000/v1/chat/completions -H 'Content-Type: application/json' -d '{
2 "model": "openbmb/MiniCPM-V-4.6-Thinking",
3 "messages": [{"role": "user", "content": [
4 {"type": "image_url", "image_url": {"url": "https://huggingface.co/datasets/openbmb/DemoCase/resolve/main/refract.png"}},
5 {"type": "text", "text": "What causes this phenomenon?"}
6 ]}]
7}'llama-server -m MiniCPM-V-4.6-Q4_K_M.gguf --port 80801curl -s http://localhost:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
2 "model": "MiniCPM-V-4.6",
3 "messages": [{"role": "user", "content": [
4 {"type": "image_url", "image_url": {"url": "https://huggingface.co/datasets/openbmb/DemoCase/resolve/main/refract.png"}},
5 {"type": "text", "text": "What causes this phenomenon?"}
6 ]}]
7}'ollama run minicpm-v-4.6-thinkingllamafactory-cli train examples/train_lora/minicpmv4_6_lora_sft.yamlswift sft --model_type minicpm-v-4_6 --dataset <your-dataset>1@misc{cui2026minicpmo45realtimefullduplex,
2 title={MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction},
3 author={Junbo Cui and Bokai Xu and Chongyi Wang and Tianyu Yu and Weiyue Sun and Yingjing Xu and Tianran Wang and Zhihui He and Wenshuo Ma and Tianchi Cai and others},
4 year={2026},
5 url={https://arxiv.org/abs/2604.27393},
6}
7
8@proceedings{yu2025minicpmv45cookingefficient,
9 title={MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe},
10 author={Tianyu Yu and Zefan Wang and Chongyi Wang and Fuwei Huang and Wenshuo Ma and Zhihui He and Tianchi Cai and Weize Chen and Yuxiang Huang and Yuanqian Zhao and others},
11 year={2025},
12 url={https://arxiv.org/abs/2509.18154},
13}
14
15@article{yao2024minicpm,
16 title={MiniCPM-V: A GPT-4V Level MLLM on Your Phone},
17 author={Yao, Yuan and Yu, Tianyu and Zhang, Ao and Wang, Chongyi and Cui, Junbo and Zhu, Hongji and Cai, Tianchi and Li, Haoyu and Zhao, Weilin and He, Zhihui and others},
18 journal={arXiv preprint arXiv:2408.01800},
19 year={2024}
20}