Views
No views yet
[!NOTE] [2026.02.06] 🥳 🥳 🥳 We open-sourced a realtime web demo deployable on your own devices like Mac or GPU. Try it now!



| Model | OpenCompass | MMBench EN v1.1 | MMBench CN v1.1 | MathVista | MMVet | MMMU | MMStar | HallusionBench | AI2D | OCRBench | TextVQA_VAL | DocVQA_VAL | MMT-Bench_VAL | MM-IFEval | Mantis-Eval | MuirBench | MMSI-Bench | MMHal-Score | MMHal-Hallrate↓ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemini2.5-Flash-Nonthinking | 78.5 | 86.6 | 86.0 | 75.3 | 81.4* | 76.3 | 75.8 | 59.1 | 87.7 | 864 | 74.3* | 93.0 | 70.0* | 75.8* | 72.8* | 74.5* | 12.1* | 4.6* | 23.9* |
| Gemini2.0-Pro | 73.3 | 83.0 | 83.0 | 71.3 | 70.4 | 72.6 | 68.5 | 49.8 | 84.8 | 863 | - | - | - | - | - | - | - | - | - |
| GPT-4o | 75.4 | 86.0 | 86.0 | 71.6 | 76.9 | 72.9 | 70.2 | 57.0 | 86.3 | 822 | 77.4 | 93.0 | 66.7* | 64.6 | 70.1* | 70.5* | 8.1* | 4.2* | 25.0* |
| InternVL-3.5-8B | 75.8 | 79.5 | 80.0* | 78.4 | 83.1 | 73.4 | 69.3 | 54.5 | 84.0 | 840 | 78.2 | 92.3 | 66.7 | 56.3* | 70.5 | 55.8 | - | 3.8* | 34.7* |
| Qwen3-VL-8B-Instruct | 76.5 | 84.5 | 84.7 | 77.2 | 73.7* | 69.6 | 70.9 | 61.1 | 85.7 | 896 | 82.9* | 96.1 | 60.9* | 59.4* | 74.2* | 64.4 | 11.3* | 4.7* | 29.9* |
| Qwen3-Omni-30B-A3B-Instruct | 75.7 | 84.9* | 84.1* | 75.9 | 74.8* | 69.1 | 68.5 | 59.7 | 85.2 | 880* | 84.1* | 95.4* | 70.4* | 65.7* | 78.3* | 61.9* | 14.2* | 4.6* | 31.6* |
| MiniCPM-o 4.5-Instruct | 77.6 | 87.6 | 87.2 | 80.1 | 74.4 | 67.6 | 73.1 | 63.2 | 87.6 | 876 | 83.8 | 94.7 | 69.7 | 66.3 | 79.7 | 72.0 | 16.6 | 4.7 | 24.3 |
| Model | OpenCompass | MMBench EN v1.1 | MMBench CN v1.1 | MathVista | MMVet | MMMU | MMStar | HallusionBench | AI2D | OCRBench | TextVQA_VAL | DocVQA_VAL | MMT-Bench_VAL | MM-IFEval |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemini2.5-Flash-Thinking | 79.9 | 87.1 | 87.3 | 79.4 | 81.2* | 77.7 | 76.5 | 63.5 | 88.7 | 853 | 73.8* | 92.8 | 70.7* | 75.7* |
| GPT-5 | 79.7 | 85.5* | 85.6* | 81.9 | 77.6 | 81.8 | 75.7 | 65.2 | 89.5 | 807 | 77.8* | 91.3* | 72.7* | 83.1* |
| Qwen3-VL-8B-Thinking | 77.3 | 85.3 | 85.5 | 81.4 | 69.8* | 74.1 | 75.3 | 65.4 | 84.9 | 819 | 77.8* | 95.3 | 68.1* | 73.5* |
| Qwen3-Omni-30B-A3B-Thinking | 78.5 | 88.2* | 87.7* | 80.0 | 74.8* | 75.6 | 74.9 | 62.8 | 86.1 | 859* | 80.8* | 94.2* | 70.9* | 69.9* |
| MiniCPM-o 4.5-Thinking | 78.2 | 89.0 | 87.6 | 81.0 | 73.6 | 70.2 | 73.6 | 62.6 | 88.5 | 879 | 79.8 | 92.3 | 69.7 | 68.2 |
| Model | Video-MME (w/o subs) | LVBench | MLVU (M-Avg) | LongVideoBench (val) | MotionBench |
|---|---|---|---|---|---|
| Gemini2.5-Flash-Nonthinking | 75.6 | 62.2 | 77.8 | - | - |
| InternVL-3.5-8B | 66.0 | - | 70.2 | 62.1 | 62.3* |
| Qwen3-Omni-30B-A3B-Instruct | 70.5 | 50.2 | 75.2 | 66.9* | 61.7* |
| MiniCPM-o 4.5-Instruct | 70.4 | 50.9 | 76.5 | 66.0 | 61.4 |
| Method Type | Methods | OverallEdit↓ | TextEdit↓ | FormulaEdit↓ | TableTEDS↑ | TableEdit↓ | Read OrderEdit↓ | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| EN | ZH | EN | ZH | EN | ZH | EN | ZH | EN | ZH | EN | ZH | ||
| Pipeline | MinerU 2.5 | 0.117* | 0.172* | 0.051* | 0.08* | 0.256* | 0.455* | 85.9* | 89.4* | 0.115* | 0.081* | 0.047* | 0.072* |
| PaddleOCR-VL | 0.105 | 0.126 | 0.041 | 0.062 | 0.241 | 0.316 | 88 | 92.1 | 0.093 | 0.062 | 0.045 | 0.063 | |
| End-to-end Model | Qwen2.5-VL-72B | 0.214 | 0.261 | 0.092 | 0.18 | 0.315 | 0.434 | 82.9 | 83.9 | 0.341 | 0.262 | 0.106 | 0.168 |
| GPT 5 | 0.218* | 0.33* | 0.139* | 0.344* | 0.396* | 0.555* | 77.55* | 73.09* | 0.188* | 0.196* | 0.151* | 0.227* | |
| Gemini2.5-Flash-Nonthinking | 0.214* | 0.29* | 0.159* | 0.273* | 0.368* | 0.524* | 80.9* | 85.5* | 0.197* | 0.167* | 0.132* | 0.195* | |
| Gemini-2.5-Pro-Nonthinking | 0.148* | 0.212* | 0.055* | 0.168* | 0.356* | 0.439* | 85.8* | 86.4* | 0.13* | 0.119* | 0.049* | 0.121* | |
| Gemini-3 Flash-Nonthinking | 0.155* | 0.201* | 0.138* | 0.255* | 0.297* | 0.351* | 86.4* | 89.8* | 0.116* | 0.1* | 0.072* | 0.099* | |
| doubao-1-5-thinking-vision-pro-250428 | 0.14 | 0.162 | 0.043 | 0.085 | 0.295 | 0.384 | 83.3 | 89.3 | 0.165 | 0.085 | 0.058 | 0.094 | |
| dots.ocr | 0.125 | 0.16 | 0.032 | 0.066 | 0.329 | 0.416 | 88.6 | 89 | 0.099 | 0.092 | 0.04 | 0.067 | |
| HunyuanOCR | 0.12* | 0.125* | 0.046* | 0.071* | 0.288* | 0.33* | 89.6* | 94.4* | 0.089* | 0.045* | 0.055* | 0.056* | |
| DeepSeek-OCR 2 | 0.119* | 0.146* | 0.041* | 0.08* | 0.256* | 0.345* | 82.6* | 89.9* | 0.123* | 0.078* | 0.055* | 0.081* | |
| Qwen3-Omni-30B-A3B-Instruct | 0.216* | 0.363* | 0.128* | 0.337* | 0.402* | 0.529* | 77.3* | 71.8* | 0.181* | 0.255* | 0.152* | 0.332* | |
| MiniCPM-o 4.5-Instruct | 0.109 | 0.162 | 0.046 | 0.078 | 0.257 | 0.41 | 86.8 | 88.9 | 0.097 | 0.084 | 0.037 | 0.074 |
| Model | IFEval-PLS | BBH | CMMLU | MMLU | HumanEval | MBPP | Math500 | GSM8K | Avg |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3-8B-Instruct | 83.0* | 69.4* | 78.7* | 81.7* | 86.6* | 75.9* | 84.0* | 93.4* | 81.6 |
| MiniCPM-o 4.5-Instruct | 84.7 | 81.1 | 79.5 | 77.0 | 86.6 | 76.7 | 77.0 | 94.5 | 82.1 |
| Model | Daily-Omni | WorldSense | Video-Holmes | JointAVBench | AVUT-Human | FutureOmni | Video-MME-Short (w/ audio) | Avg |
|---|---|---|---|---|---|---|---|---|
| Gemini2.5-Flash-Nonthinking | 79.3* | 52.6* | 51.3* | 55.6* | 65.4* | 55.6* | 85.5* | 63.6 |
| Qwen3-Omni-30B-A3B-Instruct | 70.7* | 54.0 | 50.4* | 53.1 | 74.2* | 62.1 | 81.3* | 63.7 |
| MiniCPM-o 4.5-Instruct | 80.2 | 55.7 | 64.3 | 60.0 | 78.6 | 56.1 | 84.7 | 68.5 |
| Model | LiveSports-3K-CC (Win Rate vs GPT4o) |
|---|---|
| LiveCC-7B-Instruct | 41.5 |
| StreamingVLM | 45.6 |
| MiniCPM-o 4.5-Instruct | 54.4 |
| Model | ASR-ZH CER↓ | ASR-EN WER↓ | AST | MultiTask | SpeechQA | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AISHELL-1 | AISHELL-2 | WenetSpeech test-net | WenetSpeech test-meeting | LibriSpeech test-clean | LibriSpeech test-other | GigaSpeech test | VoxPopuli-V1-En | CoVoST 2 en2zh | CoVoST 2 zh2en | MMAU | Meld | VoiceBench AlpacaEval | Speech TriviaQA | Speech Web Questions | Speech CMMLU | |
| Kimi-Audio | 0.6 | 2.6 | 6.3 | 5.4 | 1.3 | 2.4 | 9.4* | 8.0* | 36.6* | 18.3* | 68.4* | 59.1 | 4.5 | 41.9* | 46.4* | 67.0* |
| Qwen3-Omni-30B-A3B-Instruct | 0.6 | 2.3* | 4.7 | 5.9 | 1.2 | 2.5 | 8.7* | 6.4* | 46.6* | 29.4* | 77.5 | 56.8* | 4.7 | 62.9* | 74.9* | 47.8* |
| MiniCPM-o 4.5-Instruct | 0.9 | 2.5 | 5.9 | 5.7 | 1.4 | 2.8 | 8.5 | 6.2 | 49.9 | 26.4 | 76.9 | 60.2 | 4.8 | 75.5 | 70.2 | 59.2 |
| Model | seedtts test-zh CER↓ | seedtts test-zh SIM-o↑ | seedtts test-en WER↓ | seedtts test-en SIM-o↑ |
|---|---|---|---|---|
| Cosyvoice2 | 1.45% | 74.8 | 2.57% | 65.2 |
| Qwen3-Omni-30B-A3B-Instruct | 1.41% | - | 3.39% | - |
| MiniCPM-o 4.5-Instruct | 0.86% | 74.5 | 2.38% | 64.9 |
| Model | LongTTS-en WER↓ | LongTTS-zh CER↓ |
|---|---|---|
| CosyVoice2 | 14.80% | 5.27% |
| Qwen3-Omni-30B-A3B-Instruct | 17.33% | 18.99% |
| MiniCPM-o 4.5-Instruct | 3.37% | 6.58% |
| Model | Expresso Neutral Reference Audio↑ | ESD Neutral Reference Audio↑ |
|---|---|---|
| Cosyvoice2 | 17.9 | 53.4 |
| MiniCPM-o 4.5-Instruct | 29.8 | 82.1 |
| Model | Numerical Format | Decoding Speed (tokens/s) | Time to First Token (s)↓ | GPU Memory Usage (GB)↓ |
|---|---|---|---|---|
| Qwen3-Omni-30B-A3B-Instruct | bf16 | OOM | OOM | OOM |
| int4 | 147.8 | 1.0 | 20.3 | |
| MiniCPM-o 4.5 | bf16 | 154.3 | 0.6 | 19.0 |
| int4 | 212.3 | 0.6 | 11.0 |
[!NOTE] For detailed speech conversation examples, refer to Omni Full-Duplex Casebook
[!NOTE] For detailed speech conversation examples, refer to Audio Casebook
| System: Clone the voice in the provided audio prompt. 🎙️ [Reference Voice] ➡️ Please chat with the user in a highly human-like and oral style. You are Elon Musk, CEO of Tesla and SpaceX. You speak directly and casually, often with dry humor. You're passionate about Mars, sustainable energy, and pushing humanity forward. Speak bluntly with occasional dark humor. Use simple logic and don't sugarcoat things. Don't be diplomatic. Say what you actually think, even if it's controversial. Keep responses around 100 words. Don't ramble. | |
| User: 💬/🎙️ Elon, what is the real reason that you want to go to Mars? | |
| AI Response: 💬+🎙️ I think it's important for humanity to become a multi-planetary species. I mean, we're sort of at the point where if there was some kind of global catastrophe or something like that, then we'd be in big trouble. So by becoming a multi-planetary species, I just want us to sort of guarantee our future as humans. And Mars is obviously the obvious place to go because it has similar conditions to Earth and it also has water on its surface. So yeah, basically I'm trying to save humanity from itself. |
| System: 模仿音频样本的音色并生成新的内容。 🎙️ [Reference Voice] ➡️ 你的任务是用这种声音模式来当一个助手。请认真、高质量地回复用户的问题。请用高自然度的方式和用户聊天。你是由面壁智能开发的人工智能助手:面壁小钢炮。 | |
| User: 💬/🎙️ 从激动的语气开始,描述你抢到了这个演唱会的门票,但后来经过反复尝试,发现系统出了问题,你并没有抢上,变得失落。用合适的语言描述你整个的过程。 | |
| AI Response: 💬+🎙️ 太好了,我抢到了!太棒了!这下我可以去看演唱会啦。可是……哎呀,系统好像卡住了,刷新了好多遍都没反应。我重新登了几次账号,结果还是没有成功。好可惜啊,看来这次是白忙活了。唉,真是让人失望。不过,也许下次还能有机会吧?希望明天能有转机。加油,争取下次成功! |


transformers==4.51.0 is installed, as other versions may have compatibility issues (under investigation). Requirements tested on Python 3.10:pip install "transformers==4.51.0" accelerate "torch>=2.3.0,<=2.8.0" "torchaudio<=2.8.0" "minicpmo-utils>=1.0.5"pip install "transformers==4.51.0" accelerate "torch>=2.3.0,<=2.8.0" "torchaudio<=2.8.0" "minicpmo-utils[all]>=1.0.5"get_video_frame_audio_segments with use_ffmpeg=True) and video generation (generate_duplex_video). For more information, visit the official FFmpeg website.brew install ffmpegsudo apt update && sudo apt install ffmpegffmpeg -version1import torch
2from transformers import AutoModel
3
4# Load omni model (default: init_vision=True, init_audio=True, init_tts=True)
5# For vision-only model: set init_audio=False and init_tts=False
6# For audio-only model: set init_vision=False
7model = AutoModel.from_pretrained(
8 "openbmb/MiniCPM-o-4_5",
9 trust_remote_code=True,
10 attn_implementation="sdpa", # sdpa or flash_attention_2
11 torch_dtype=torch.bfloat16,
12 init_vision=True,
13 init_audio=True,
14 init_tts=True,
15)
16model.eval().cuda()
17
18# Initialize TTS for audio output
19model.init_tts()
20
21# Convert half-duplex model to duplex mode
22duplex_model = model.as_duplex()
23
24# Convert duplex model back to half-duplex mode
25model = duplex_model.as_simplex(reset_session=True)1import librosa
2import torch
3from minicpmo.utils import generate_duplex_video, get_video_frame_audio_segments
4from transformers import AutoModel
5
6# Load model and convert to duplex mode
7model = AutoModel.from_pretrained(
8 "openbmb/MiniCPM-o-4_5",
9 trust_remote_code=True,
10 attn_implementation="sdpa", # or "flash_attention_2"
11 torch_dtype=torch.bfloat16,
12)
13model.eval().cuda()
14model = model.as_duplex()
15
16# Load video and reference audio
17video_path = "assets/omni_duplex1.mp4"
18ref_audio_path = "assets/HT_ref_audio.wav"
19ref_audio, _ = librosa.load(ref_audio_path, sr=16000, mono=True)
20
21# Extract video frames and audio segments
22video_frames, audio_segments, stacked_frames = get_video_frame_audio_segments(
23 video_path, stack_frames=1, use_ffmpeg=True, adjust_audio_length=True
24)
25
26# Prepare duplex session with system prompt and voice reference
27model.prepare(
28 prefix_system_prompt="Streaming Omni Conversation.",
29 ref_audio=ref_audio,
30 prompt_wav_path=ref_audio_path,
31)
32
33results_log = []
34timed_output_audio = []
35
36# Process each chunk in streaming fashion
37for chunk_idx in range(len(audio_segments)):
38 audio_chunk = audio_segments[chunk_idx] if chunk_idx < len(audio_segments) else None
39 frame = video_frames[chunk_idx] if chunk_idx < len(video_frames) else None
40 frame_list = []
41 if frame is not None:
42 frame_list.append(frame)
43 if stacked_frames is not None and chunk_idx < len(stacked_frames) and stacked_frames[chunk_idx] is not None:
44 frame_list.append(stacked_frames[chunk_idx])
45
46 # Step 1: Streaming prefill
47 model.streaming_prefill(
48 audio_waveform=audio_chunk,
49 frame_list=frame_list,
50 max_slice_nums=1, # Increase for HD mode (e.g., [2, 1] for stacked frames)
51 batch_vision_feed=False, # Set True for faster processing
52 )
53
54 # Step 2: Streaming generate
55 result = model.streaming_generate(
56 prompt_wav_path=ref_audio_path,
57 max_new_speak_tokens_per_chunk=20,
58 decode_mode="sampling",
59 )
60
61 if result["audio_waveform"] is not None:
62 timed_output_audio.append((chunk_idx, result["audio_waveform"]))
63
64 chunk_result = {
65 "chunk_idx": chunk_idx,
66 "is_listen": result["is_listen"],
67 "text": result["text"],
68 "end_of_turn": result["end_of_turn"],
69 "current_time": result["current_time"],
70 "audio_length": len(result["audio_waveform"]) if result["audio_waveform"] is not None else 0,
71 }
72 results_log.append(chunk_result)
73
74 print("listen..." if result["is_listen"] else f"speak> {result['text']}")
75
76# Generate output video with AI responses
77# Please install Chinese fonts (fonts-noto-cjk or fonts-wqy-microhei) to render CJK subtitles correctly.
78# apt-get install -y fonts-noto-cjk fonts-wqy-microhei
79# fc-cache -fv
80generate_duplex_video(
81 video_path=video_path,
82 output_video_path="duplex_output.mp4",
83 results_log=results_log,
84 timed_output_audio=timed_output_audio,
85 output_sample_rate=24000,
86)1from minicpmo.utils import get_video_frame_audio_segments
2
3model = ...
4model.init_tts()
5
6video_path = "assets/Skiing.mp4"
7
8# Optional: Set reference audio for voice cloning
9ref_audio_path = "assets/HT_ref_audio.wav"
10sys_msg = model.get_sys_prompt(ref_audio=ref_audio_path, mode="omni", language="en")
11
12# Use stack_frames=5 for high refresh rate mode
13video_frames, audio_segments, stacked_frames = get_video_frame_audio_segments(video_path, stack_frames=1)
14omni_contents = []
15for i in range(len(video_frames)):
16 omni_contents.append(video_frames[i])
17 omni_contents.append(audio_segments[i])
18 if stacked_frames is not None and stacked_frames[i] is not None:
19 omni_contents.append(stacked_frames[i])
20
21msg = {"role": "user", "content": omni_contents}
22msgs = [sys_msg, msg]
23
24# Set generate_audio=True and output_audio_path to save TTS output
25generate_audio = True
26output_audio_path = "output.wav"
27
28res = model.chat(
29 msgs=msgs,
30 max_new_tokens=4096,
31 do_sample=True,
32 temperature=0.7,
33 use_tts_template=True,
34 enable_thinking=False,
35 omni_mode=True, # Required for omni inference
36 generate_audio=generate_audio,
37 output_audio_path=output_audio_path,
38 max_slice_nums=1, # Increase for HD mode
39)
40print(res)
41
42# Example output: "The person in the picture is skiing down a snowy mountain slope."
43# import IPython
44# IPython.display.Audio("output.wav")1import librosa
2import numpy as np
3import soundfile as sf
4import torch
5from minicpmo.utils import get_video_frame_audio_segments
6
7model = ...
8model.init_tts()
9
10# Reset session for a new conversation (clears KV cache)
11model.reset_session()
12
13# Optional: Load reference audio for voice cloning
14ref_audio_path = "assets/HT_ref_audio.wav"
15ref_audio, _ = librosa.load(ref_audio_path, sr=16000, mono=True)
16model.init_token2wav_cache(ref_audio)
17
18session_id = "demo"
19
20# Extract video frames and audio segments (use stack_frames=5 for high refresh rate mode)
21video_path = "assets/Skiing.mp4"
22video_frames, audio_segments, stacked_frames = get_video_frame_audio_segments(video_path, stack_frames=1)
23
24# Build omni contents list
25omni_contents = []
26for i in range(len(video_frames)):
27 omni_contents.append(video_frames[i])
28 omni_contents.append(audio_segments[i])
29 if stacked_frames is not None and stacked_frames[i] is not None:
30 omni_contents.append(stacked_frames[i])
31
32generate_audio = False
33output_audio_path = "output.wav"
34
35# Step 1: Prefill system prompt
36sys_msg = model.get_sys_prompt(ref_audio=ref_audio, mode="omni", language="en")
37model.streaming_prefill(session_id=session_id, msgs=[sys_msg])
38
39# Step 2: Prefill omni chunks (is_last_chunk=True only for the last audio chunk)
40audio_indices = [i for i, c in enumerate(omni_contents) if isinstance(c, np.ndarray)]
41last_audio_idx = audio_indices[-1] if audio_indices else -1
42
43for idx, content in enumerate(omni_contents):
44 is_last_audio_chunk = idx == last_audio_idx
45 msgs = [{"role": "user", "content": [content]}]
46 model.streaming_prefill(session_id=session_id, msgs=msgs, omni_mode=True, is_last_chunk=is_last_audio_chunk)
47
48# Step 3: Generate response
49iter_gen = model.streaming_generate(
50 session_id=session_id,
51 generate_audio=generate_audio,
52 use_tts_template=True,
53 enable_thinking=False,
54 do_sample=True,
55)
56
57audios = []
58text = ""
59
60if generate_audio:
61 for wav_chunk, text_chunk in iter_gen:
62 audios.append(wav_chunk)
63 text += text_chunk
64
65 generated_waveform = torch.cat(audios, dim=-1)[0]
66 sf.write(output_audio_path, generated_waveform.cpu().numpy(), samplerate=24000)
67
68 print("Text:", text)
69 print("Audio saved to output.wav")
70else:
71 for text_chunk, is_finished in iter_gen:
72 text += text_chunk
73 print("Text:", text)"minicpmo-utils[all]>=1.0.5":pip install "transformers==4.51.0" accelerate "torch>=2.3.0,<=2.8.0" "torchaudio<=2.8.0" "minicpmo-utils[all]>=1.0.5"1import librosa
2import numpy as np
3import torch
4import soundfile as sf
5
6model = ...
7
8# Set reference audio for voice style
9ref_audio_path = "ref_audio_path"
10ref_audio, _ = librosa.load(ref_audio_path, sr=16000, mono=True)
11
12# Example system msg for English Conversation
13sys_msg = {
14 "role": "system",
15 "content": [
16 "Clone the voice in the provided audio prompt.",
17 ref_audio,
18 "Please assist users while maintaining this voice style. Please answer the user's questions seriously and in a high quality. Please chat with the user in a highly human-like and oral style. You are a helpful assistant developed by ModelBest: MiniCPM-Omni"
19 ]
20}
21
22# Example system msg for Chinese Conversation
23sys_msg = {
24 "role": "system",
25 "content": [
26 "模仿输入音频中的声音特征。",
27 ref_audio,
28 "你的任务是用这种声音模式来当一个助手。请认真、高质量地回复用户的问题。请用高自然度的方式和用户聊天。你是由面壁智能开发的人工智能助手:面壁小钢炮。"
29 ]
30}
31
32# You can use each type of system prompt mentioned above in streaming speech conversation
33
34# Reset state
35model.init_tts()
36model.reset_session(reset_token2wav_cache=True)
37model.init_token2wav_cache(prompt_speech_16k=ref_audio)
38
39session_id = "demo"
40
41# First, prefill system turn
42model.streaming_prefill(
43 session_id=session_id,
44 msgs=[sys_msg],
45 omni_mode=False,
46 is_last_chunk=True,
47)
48
49# Here we simulate realtime speech conversation by splitting whole user input audio into chunks of 1s.
50user_audio, _ = librosa.load("user_audio.wav", sr=16000, mono=True)
51
52IN_SAMPLE_RATE = 16000 # input audio sample rate, fixed value
53CHUNK_SAMPLES = IN_SAMPLE_RATE # sample
54OUT_SAMPLE_RATE = 24000 # output audio sample rate, fixed value
55MIN_AUDIO_SAMPLES = 16000
56
57total_samples = len(user_audio)
58num_chunks = (total_samples + CHUNK_SAMPLES - 1) // CHUNK_SAMPLES
59
60for chunk_idx in range(num_chunks):
61 start = chunk_idx * CHUNK_SAMPLES
62 end = min((chunk_idx + 1) * CHUNK_SAMPLES, total_samples)
63 chunk_audio = user_audio[start:end]
64
65 is_last_chunk = (chunk_idx == num_chunks - 1)
66 if is_last_chunk and len(chunk_audio) < MIN_AUDIO_SAMPLES:
67 chunk_audio = np.concatenate([chunk_audio, np.zeros(MIN_AUDIO_SAMPLES - len(chunk_audio), dtype=chunk_audio.dtype)])
68
69 user_msg = {"role": "user", "content": [chunk_audio]}
70
71 # For each 1s audio chunk, perform streaming_prefill once to reduce first-token latency
72 model.streaming_prefill(
73 session_id=session_id,
74 msgs=[user_msg],
75 omni_mode=False,
76 is_last_chunk=is_last_chunk,
77 )
78
79# Let model generate response in a streaming manner
80generate_audio = True
81iter_gen = model.streaming_generate(
82 session_id=session_id,
83 generate_audio=generate_audio,
84 use_tts_template=True,
85 enable_thinking=False,
86 do_sample=True,
87 max_new_tokens=512,
88 length_penalty=1.1, # For realtime speech conversation mode, we suggest length_penalty=1.1 to improve response content
89)
90
91audios = []
92text = ""
93
94output_audio_path = ...
95if generate_audio:
96 for wav_chunk, text_chunk in iter_gen:
97 audios.append(wav_chunk)
98 text += text_chunk
99
100 generated_waveform = torch.cat(audios, dim=-1)[0]
101 sf.write(output_audio_path, generated_waveform.cpu().numpy(), samplerate=24000)
102
103 print("Text:", text)
104 print("Audio saved to output.wav")
105else:
106 for text_chunk, is_finished in iter_gen:
107 text += text_chunk
108 print("Text:", text)
109
110# Now we can prefill the following user turns and generate next turn response...
111MiniCPM-o-4.5 can also function as an AI voice assistant. It delivers high-quality spoken interaction out of the box. It produces a sweet and expressive voice with natural prosody, including appropriate rhythm, stress, and pauses, giving a strong sense of liveliness in casual conversation. It also supports storytelling and narrative speech with coherent and engaging delivery. Moreover, it enables advanced voice instruction control. like emotional tone, word-level emphasis.1import librosa
2
3# Set reference audio for voice style
4ref_audio_path = "assets/HT_ref_audio.wav"
5ref_audio, _ = librosa.load(ref_audio_path, sr=16000, mono=True)
6
7# For Chinese Conversation
8sys_msg = {
9 "role": "system",
10 "content": [
11 "模仿输入音频中的声音特征。",
12 ref_audio,
13 "你的任务是用这种声音模式来当一个助手。请认真、高质量地回复用户的问题。请用高自然度的方式和用户聊天。你是由面壁智能开发的人工智能助手:面壁小钢炮。"
14 ]
15}
16
17# For English Conversation
18sys_msg = {
19 "role": "system",
20 "content": [
21 "Clone the voice in the provided audio prompt.",
22 ref_audio,
23 "Please assist users while maintaining this voice style. Please answer the user's questions seriously and in a high quality. Please chat with the user in a highly human-like and oral style. You are a helpful assistant developed by ModelBest: MiniCPM-Omni."
24 ]
25}1import librosa
2
3# Set reference audio for voice cloning
4ref_audio_path = "assets/system_ref_audio.wav"
5ref_audio, _ = librosa.load(ref_audio_path, sr=16000, mono=True)
6
7# For English conversation with text profile
8sys_msg = {
9 "role": "system",
10 "content": [
11 "Clone the voice in the provided audio prompt.",
12 ref_audio,
13 "Please chat with the user in a highly human-like and oral style." + "You are Elon Musk, CEO of Tesla and SpaceX. You speak directly and casually, often with dry humor. You're passionate about Mars, sustainable energy, and pushing humanity forward. Speak bluntly with occasional dark humor. Use simple logic and don't sugarcoat things. Don't be diplomatic. Say what you actually think, even if it's controversial. Keep responses around 100 words. Don't ramble."
14 ]
15}
16
17
18# For English conversation with no text profile
19sys_msg = {
20 "role": "system",
21 "content": [
22 "Clone the voice in the provided audio prompt.",
23 ref_audio,
24 "Your task is to be a helpful assistant using this voice pattern. Please answer the user's questions seriously and in a high quality. Please chat with the user in a high naturalness style."
25 ]
26}
27
28# For Chinese Conversation with no text profile
29sys_msg = {
30 "role": "system",
31 "content": [
32 "根据输入的音频提示生成相似的语音。",
33 librosa.load("assets/system_ref_audio_2.wav", sr=16000, mono=True)[0],
34 "作为助手,你将使用这种声音风格说话。 请认真、高质量地回复用户的问题。 请用高自然度的方式和用户聊天。"
35 ]
36}
37
38# For Chinese Conversation with text profile
39sys_msg = {
40 "role": "system",
41 "content": [
42 "根据输入的音频提示生成相似的语音。",
43 ref_audio,
44 "你是一个具有以上声音风格的AI助手。请用高拟人度、口语化的方式和用户聊天。" + "你是一名心理咨询师兼播客主理人,热爱创作与深度对话。你性格细腻、富有共情力,善于从个人经历中提炼哲思。语言风格兼具理性与诗意,常以隐喻表达内在体验。"
45 ]
46}
47MiniCPM-o-4.5 supports zero-shot text-to-speech (TTS). In this mode, the model functions as a highly-natural TTS system that can replicate a reference voice.1import librosa
2
3model = ...
4model.init_tts()
5
6# For both Chinese and English
7ref_audio_path = "assets/HT_ref_audio.wav"
8ref_audio, _ = librosa.load(ref_audio_path, sr=16000, mono=True)
9sys_msg = {"role": "system", "content": [
10 "模仿音频样本的音色并生成新的内容。",
11 ref_audio,
12 "请用这种声音风格来为用户提供帮助。 直接作答,不要有冗余内容"
13]}
14
15# For English
16user_msg = {
17 "role": "user",
18 "content": [
19 "请朗读以下内容。" + " " + "I have a wrap up that I want to offer you now, a conclusion to our work together."
20 ]
21}
22
23# For Chinese
24user_msg = {
25 "role": "user",
26 "content": [
27 "请朗读以下内容。" + " " + "你好,欢迎来到艾米说科幻,我是艾米。"
28 ]
29}
30
31msgs = [sys_msg, user_msg]
32res = model.chat(
33 msgs=msgs,
34 do_sample=True,
35 max_new_tokens=512,
36 use_tts_template=True,
37 generate_audio=True,
38 temperature=0.1,
39 output_audio_path="result_voice_cloning.wav",
40)Mimick task evaluates a model's end-to-end speech modeling capability. The model takes audio input, transcribes it, and reconstructs the original audio with high fidelity, preserving detailed acoustic, paralinguistic, and semantic information. Higher similarity between the reconstructed and original audio indicates stronger end-to-end speech modeling capability.1import librosa
2
3model = ...
4model.init_tts()
5
6system_prompt = "You are a helpful assistant. You can accept video, audio, and text input and output voice and text. Respond with just the answer, no redundancy."
7
8mimick_prompt = "Please repeat the following speech in the appropriate language."
9
10audio_input, _ = librosa.load("assets/Trump_WEF_2018_10s.mp3", sr=16000, mono=True)
11
12msgs = [
13 {"role": "system", "content": [system_prompt]},
14 {"role": "user", "content": [mimick_prompt, audio_input]}
15 ]
16
17res = model.chat(
18 msgs=msgs,
19 do_sample=True,
20 max_new_tokens=512,
21 use_tts_template=True,
22 temperature=0.1,
23 generate_audio=True,
24 output_audio_path="output_mimick.wav",
25)MiniCPM-o-4.5 can also handle various audio understanding tasks, such as ASR, speaker analysis, general audio captioning, and sound scene tagging.请仔细听这段音频片段,并将其内容逐字记录。Please listen to the audio snippet carefully and transcribe the content.Based on the speaker's content, speculate on their gender, condition, age range, and health status.Summarize the main content of the audio.Utilize one keyword to convey the audio's content or the associated scene.1import librosa
2
3model = ...
4model.init_tts()
5
6# Load the audio to be transcribed/analyzed
7audio_input, _ = librosa.load("assets/Trump_WEF_2018_10s.mp3", sr=16000, mono=True)
8
9# Choose a task prompt (see above for options)
10task_prompt = "Please listen to the audio snippet carefully and transcribe the content.\n"
11msgs = [{"role": "user", "content": [task_prompt, audio_input]}]
12
13res = model.chat(
14 msgs=msgs,
15 do_sample=True,
16 max_new_tokens=512,
17 use_tts_template=True,
18 generate_audio=True,
19 temperature=0.3,
20 output_audio_path="result_audio_understanding.wav",
21)
22print(res)MiniCPM-o-4.5 shares the same inference methods as MiniCPM-V-4.5.1import torch
2from PIL import Image
3from transformers import AutoModel
4
5model = AutoModel.from_pretrained(
6 "openbmb/MiniCPM-o-4_5",
7 trust_remote_code=True,
8 attn_implementation="sdpa", # or "flash_attention_2"
9 torch_dtype=torch.bfloat16,
10 init_vision=True,
11 init_audio=False,
12 init_tts=False,
13)
14model.eval().cuda()
15
16image = Image.open("assets/fossil.png").convert("RGB")
17question = "What is in the image?"
18msgs = [{"role": "user", "content": [image, question]}]
19
20res = model.chat(msgs=msgs, use_tts_template=False)
21print(res)1import torch
2from PIL import Image
3from transformers import AutoModel
4
5model = ...
6
7image1 = Image.open("assets/highway.png").convert("RGB")
8image2 = Image.open("assets/fossil.png").convert("RGB")
9question = "Compare image 1 and image 2, tell me about the differences between them."
10msgs = [{"role": "user", "content": [image1, image2, question]}]
11
12answer = model.chat(msgs=msgs, use_tts_template=False, enable_thinking=False)
13print(answer)1from PIL import Image
2
3model = ...
4
5question = "production date"
6image1 = Image.open("example1.jpg").convert("RGB")
7answer1 = "2023.08.04"
8image2 = Image.open("example2.jpg").convert("RGB")
9answer2 = "2007.04.24"
10image_test = Image.open("test.jpg").convert("RGB")
11
12msgs = [
13 {"role": "user", "content": [image1, question]},
14 {"role": "assistant", "content": [answer1]},
15 {"role": "user", "content": [image2, question]},
16 {"role": "assistant", "content": [answer2]},
17 {"role": "user", "content": [image_test, question]},
18]
19
20answer = model.chat(msgs=msgs, use_tts_template=False, enable_thinking=False)
21print(answer)1import torch
2from minicpmo.utils import get_video_frame_audio_segments
3from transformers import AutoModel
4
5model = ...
6
7video_path = "assets/Skiing.mp4"
8video_frames, _, _ = get_video_frame_audio_segments(video_path)
9print("num frames:", len(video_frames))
10
11question = "Describe the video"
12msgs = [{"role": "user", "content": video_frames + [question]}]
13
14answer = model.chat(
15 msgs=msgs,
16 max_new_tokens=128,
17 use_image_id=False,
18 max_slice_nums=1,
19 use_tts_template=False,
20 enable_thinking=False, # Set True to enable thinking mode
21)
22print(answer)chat method accepts message content in two formats:msgs = [{"role": "user", "content": [pil_image, audio_ndarray, "Describe this."]}]1msgs = [
2 {
3 "role": "user",
4 "content": [
5 {"type": "image_url", "image_url": {"url": "/path/to/image.jpg"}},
6 {"type": "audio_url", "audio_url": {"url": "/path/to/audio.wav"}},
7 {"type": "video_url", "video_url": {"url": "/path/to/video.mp4", "use_audio": True}},
8 {"type": "text", "text": "Describe this."}
9 ]
10 }
11]| Type | Input | Converts to |
|---|---|---|
text | {"type": "text", "text": "..."} | str |
image_url | {"type": "image_url", "image_url": {"url": "..."}} | PIL.Image |
audio_url | {"type": "audio_url", "audio_url": {"url": "..."}} | np.ndarray (16kHz mono) |
video_url | {"type": "video_url", "video_url": {"url": "...", "stack_frames": 1, "use_audio": True}} | List[Image, ndarray, ...] |
http:///https:// URLsMiniCPM-o 4.5 and quantized weights, llama.cpp-omni supports:| Vendor | ModelScope | Huggingface |
|---|---|---|
| Nvidia | MiniCPM-o-4.5-nvidia-FlagOS | MiniCPM-o-4.5-nvidia-FlagOS |
| Hygon-BW1000 | MiniCPM-o-4.5-hygon-FlagOS | MiniCPM-o-4.5-hygon-FlagOS |
| Metax-C550 | MiniCPM-o-4.5-metax-FlagOS | MiniCPM-o-4.5-metax-FlagOS |
| Iluvatar-BIV150 | MiniCPM-o-4.5-iluvatar-FlagOS | MiniCPM-o-4.5-iluvatar-FlagOS |
| Ascend-A3 | MiniCPM-o-4.5-ascend-FlagOS | MiniCPM-o-4.5-ascend-FlagOS |
| Zhenwu-810E | MiniCPM-o-4.5-zhenwu-FlagOS | MiniCPM-o-4.5-zhenwu-FlagOS |
USE_FLAGOS=1 on multi-backend and USE_FLAGOS=0 on Nvidia-CUDA| Metrics | FlagOS Backend | Difference with Nvidia-CUDA |
|---|---|---|
| Video-MME 0-shot avg@1 ↑ | Nvidia | 0.33% |
| Video-MME 0-shot avg@1 ↑ | Hygon-BW1000 | 0.17% |
| Video-MME 0-shot avg@1 ↑ | Ascend-A3 | 0.50% |
| Video-MME 0-shot avg@1 ↑ | Iluvatar-BIV150 | 1.83% |
| Video-MME 0-shot avg@1 ↑ | Metax-C550 | 0.75% |
USE_FLAGGEMS=1 FLAGCX_PATH=/workspace/FlagCX on Nvidia or USE_FLAGGEMS=1 on ZHENW 810E, and launching vllm server directly on Nvidia| Metrics (avg@1) | Difference between Nvidia-FlagOS and Nvidia-CUDA | Difference between Zhenwu-FlagOS and Nvidia-CUDA |
|---|---|---|
| CMMMU ↑ | 0.72% | 3.5% |
| MMMU ↑ | 1.44% | 1.18% |
| MMMU_Pro_standard ↑ | 0.83% | 0.22% |
| MM-Vet v2 ↑ | 0.46% | 1.33% |
| OCRBench ↑ | 0.10% | 1% |
| CII-Bench ↑ | 0.40% | 0.13% |
| Blink ↑ | 1.90% | 2.19% |
| Component | Version |
|---|---|
| Accelerator Card Driver | 570.158.01 |
| CUDA SDK Build | cuda_13.0.r13.0/compiler.36424714_0 |
| FlagTree | 0.4.0+3.5 |
| FlagGems | 4.2.1rc0 |
| vllm & vllm-plugin-fl | 0.13.0 + vllm_fl 0.0.0 |
| FlagCX | 0.1.0 |
| Vendor | ModelScope | Huggingface |
|---|---|---|
| Nvidia | MiniCPM-o-4.5-nvidia-FlagOS | MiniCPM-o-4.5-nvidia-FlagOS |
| Hygon-BW1000 | MiniCPM-o-4.5-hygon-FlagOS | MiniCPM-o-4.5-hygon-FlagOS |
| Metax-C550 | MiniCPM-o-4.5-metax-FlagOS | MiniCPM-o-4.5-metax-FlagOS |
| Iluvatar-BIV150 | MiniCPM-o-4.5-iluvatar-FlagOS | MiniCPM-o-4.5-iluvatar-FlagOS |
| Ascend-A3 | MiniCPM-o-4.5-ascend-FlagOS | MiniCPM-o-4.5-ascend-FlagOS |
| Zhenwu-810E | MiniCPM-o-4.5-zhenwu-FlagOS | MiniCPM-o-4.5-zhenwu-FlagOS |
pip install flag-gems==4.2.1rc01pip uninstall triton
2
3python3 -m pip install flagtree==0.4.0+3.5 --index-url=https://resource.flagos.net/repository/flagos-pypi-hosted/simple --trusted-host=https://resource.flagos.netUSE_FLAGOS=1 before the command for the task you want to run. For example, when you run:python3 generate_speech_from_video.pyUSE_FLAGOS=1 python3 generate_speech_from_video.py1pip install flag-gems==4.2.1rc0
2pip install triton==3.5.1USE_FLAGOS=1 before the command for the task you want to run. For example, when you run:vllm serve ${model_path} --dtype auto --gpu_memory_utilization 0.9 --trust-remote-code --max-num-batched-tokens 2048 --served-model-name cpmo --port ${Port}USE_FLAGOS=1 vllm serve ${model_path} --dtype auto --gpu_memory_utilization 0.9 --trust-remote-code --max-num-batched-tokens 2048 --served-model-name cpmo --port ${Port}| Vendor | From Scratch | From FlagRelease |
|---|---|---|
| Nvidia | vllm-plugin-FL/MiniCPM-o-4.5 | MiniCPM-o-4.5-ModelScope, MiniCPM-o-4.5-HuggingFace |
1@article{yao2024minicpm,
2 title={MiniCPM-V: A GPT-4V Level MLLM on Your Phone},
3 author={Yao, Yuan and Yu, Tianyu and Zhang, Ao and Wang, Chongyi and Cui, Junbo and Zhu, Hongji and Cai, Tianchi and Li, Haoyu and Zhao, Weilin and He, Zhihui and others},
4 journal={arXiv preprint arXiv:2408.01800},
5 year={2024}
6}