Ready-to-run deployment package for openbmb/MiniCPM-V-4.6 on AX650 / NPU3.
This release packages the AX650 axllm runtime together with the compiled text and vision .axmodel files.
The packaged text runtime uses the non-GPTQ BF16 build.
The packaged vision runtime uses a fixed-shape 448x448 MiniCPM-V-4.6 vision encoder.
The package supports text-only chat, single-image understanding, and video understanding through the OpenAI-compatible axllm serve API.
The package also includes board-side and server-side Python reference scripts for reference use and comparison.
Supported Platform
AX650 / NPU3
Validated Devices
This package has been validated on the following AX650-based device:
AX650 / NPU3 development board
Performance
All measurements below were taken on AX650 / NPU3. TTFT stands for time to first token. In this table, TTFT is measured end-to-end from request arrival at axllm serve to the first generated token, so the multimodal rows include media preprocessing and vision encoding time.
The text-only smoke prompt was kept within one 128-token prefill chunk. To avoid one-time startup effects, the text row below excludes the first request after service startup. Its Decode figure was measured with longer text-only generations (max_tokens=256) to better reflect sustained decode throughput; very short smoke replies under-report decode speed because EOS and response-tail overhead become relatively larger. The image row was measured with the packaged fixed-shape 448x448 vision encoder and assets/sample.png. The video row used the packaged sample video with video:assets/red-panda-openai.mp4:2.
Scenario
Input tokens
Prefill chunks
TTFT
Decode
Text-only smoke prompt
25
1 x 128
275.88 ms avg (274.97-276.78 ms)
19.12 token/s avg
Image prompt
88
1 x 128
729.89 ms avg (723.81-741.89 ms)
19.02 token/s avg
Video prompt
1271
10 x 128
9652.87 ms avg (9585.79-9735.26 ms)
18.84 token/s avg
The packaged runtime uses the following context layout:
prefill_len=128
kv_cache_len=2047
prefill_max_token_num=1280
Input tokens in the table above refers to the full request length after chat templating, not just the visual soft tokens. For the shipped 448x448 vision encoder, each selected image block contributes 64 visual soft tokens. Under the current packaged runtime settings, the sample video request in this README uses 1271 total input tokens and spans 10 prefill chunks.
Startup Runtime Footprint
Item
Value
Flash total (text + post + vision axmodels)
1.42 GiB (1458.81 MiB)
Package flash total (excluding vision_cache/)
1.93 GiB (1979.79 MiB)
Runtime CMM increment during board-side startup
1.53 GiB (1564.55 MiB)
The runtime CMM value above was measured during board-side startup on the validated AX650 board configuration and should be treated as a practical reference value.
Vision Encoder Latency
Measured on AX650 / NPU3 with /opt/bin/ax_run_model -m minicpmv4_6_vision_448.axmodel -g 0 -w 1 -r 5.
Model
Resolution
Soft Tokens
Time (ms)
minicpmv4_6_vision_448.axmodel
448x448
64
234.827 ms avg
For this packaged AX650 runtime, the visual token count is fixed by the shipped vision encoder configuration:
vision_width = 448
vision_height = 448
vision_patch_size = 14
patch grid = (448 / 14) x (448 / 14) = 32 x 32
raw patch tokens = 32 x 32 = 1024
current packaged build uses the 16x visual compression path
Soft Tokens = 1024 / 16 = 64
So, for the fixed-shape runtime shipped in this repository, the relation is:
Input tokens in the performance table can be larger than the visual Soft Tokens because axllm counts the full templated request, including user text and chat-template tokens in addition to the visual tokens. For the packaged assets/sample.png request in this README, the runtime reports input_num_token=88, which still fits within a single 128-token prefill chunk.
Soft Tokens is not a runtime-configurable value in this package. This repository ships only minicpmv4_6_vision_448.axmodel, so the board-side AX650 runtime always uses 448x448 -> 64 soft tokens for image encoding.
This package uses a hybrid layout: the packaged axllm runtime plus the compiled .axmodel files live at the repository root, while the Python reference scripts and the tokenizer directory used by those scripts stay under python/.
Sample Image
Both the axllm flow and the packaged Python examples can use the sample image:
assets/sample.png
sample
Sample Video
The package also includes a packaged sample video for board-side video understanding validation:
assets/red-panda-openai.mp4
Direct Inference with axllm
The axllm workflow is still being refined. The instructions below reflect the current validated flow and may be adjusted as the packaging continues to evolve.
Option 4: download the prebuilt binary from GitHub Actions CI:
If you do not have a local build environment, download the latest CI-generated axllm binary from GitHub Actions:
https://github.com/AXERA-TECH/ax-llm/actions?query=branch%3Aaxllm
Then run:
shell
1chmod +x axllm
2sudomv axllm /usr/bin/axllm
Run on the Board
The package root is already arranged for axllm, so no extra runtime path arguments are required.
For multimodal testing, you can use the packaged sample image shown above: ./assets/sample.png, or the packaged sample video: ./assets/red-panda-openai.mp4.
./bin/axllm run .
In interactive mode:
press Enter directly for text-only chat
input an image path for single-image chat
input video:/path/to/frames_dir or video:/path/to/video.mp4 for video chat
Serve with axllm
From the package root on the board:
./bin/axllm serve . --port 8000
Expected model id:
AXERA-TECH/MiniCPM-V-4.6-AX650-C128-P1152-CTX2047
Health check:
curl http://127.0.0.1:8000/health
A typical startup log looks like this:
text
1INF Init | LLM init start
2INF Init | mixed attention enabled: full_attention_interval=4 ref_full_layer_idx=3
3INF Init | attention config: layers=24 sliding=0 full=6 linear=18 sliding_window=0 ref_full_layer_idx=3
4tokenizer_type = 3
5huggingface tokenizer mode = gpt2_byte_bpe
6...
7INF Init | max_token_len : 2047
8INF Init | kv_cache_size : 512, kv_cache_num: 2047
9INF init_groups_from_model | prefill_token_num : 128
10INF init_groups_from_model | prefill_max_token_num : 1280
11INF Init | MiniCPM-V-4.6 token ids: image_pad=248056 video_pad=248057
12INF Init | VisionModule init ok: type=MiniCPMV46VL, tokens_per_block=64, embed_size=1024, out_dtype=fp32
13INF Init | LLM init ok
14Starting server on port 8000 with model 'AXERA-TECH/MiniCPM-V-4.6-AX650-C128-P1152-CTX2047'...
15API URLs:
16 GET http://127.0.0.1:8000/health
17 GET http://127.0.0.1:8000/v1/models
18 POST http://127.0.0.1:8000/v1/chat/completions
19OpenAI API Server starting on http://0.0.0.0:8000
20Max concurrency: 1
21Models: AXERA-TECH/MiniCPM-V-4.6-AX650-C128-P1152-CTX2047
You can then send requests to the server using the API endpoints shown in the log. For example, to check the health status and list the available models:
1{2"choices":[3{4"message":{5"role":"assistant",6"content":"1+1 is 2."7},8"finish_reason":"stop"9}10],11"model":"AXERA-TECH/MiniCPM-V-4.6-AX650-C128-P1152-CTX2047",12"object":"chat.completion"13}