Views
No views yet

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.In light of its parameter scale, the intended use cases are prototyping, task-specific fine-tuning, and other research or development purposes.
| Qwen3-4B-2507 | Qwen3-1.7B | Qwen3.5-2B | Qwen3.5-0.8B | |
|---|---|---|---|---|
| Non-Thinking Mode | ||||
| MMLU-Pro | 69.6 | 40.2 | 55.3 | 29.7 |
| MMLU-Redux | 84.2 | 64.4 | 69.2 | 48.5 |
| C-Eval | 80.2 | 61.0 | 65.2 | 46.4 |
| SuperGPQA | 42.8 | 21.0 | 30.4 | 16.9 |
| IFEval | 83.4 | 68.2 | 61.2 | 52.1 |
| MMMLU | 64.9 | 46.7 | 56.9 | 34.1 |
| Knowledge & STEM (Thinking) | ||||
| MMLU-Pro | 74.0 | 56.5 | 66.5 | 42.3 |
| MMLU-Redux | 86.1 | 73.9 | 79.6 | 59.5 |
| C-Eval | 82.2 | 68.1 | 73.2 | 50.5 |
| SuperGPQA | 47.8 | 31.2 | 37.5 | 21.3 |
| GPQA | 65.8 | 40.1 | 51.6 | 11.9 |
| Instruction Following (Thinking) | ||||
| IFEval | 87.4 | 72.5 | 78.6 | 44.0 |
| IFBench | 50.4 | 26.7 | 41.3 | 21.0 |
| MultiChallenge | 41.7 | 27.2 | 33.7 | 18.9 |
| Long Context (Thinking) | ||||
| AA-LCR | 32.0 | 6.7 | 25.6 | 4.7 |
| LongBench v2 | 42.8 | 26.5 | 38.7 | 26.1 |
| Reasoning (Thinking) | ||||
| HMMT Feb 25 | 57.5 | 10.2 | 22.9 | -- |
| HMMT Nov 25 | 69.6 | 8.9 | 19.6 | -- |
| General Agent (Thinking) | ||||
| BFCL-V4 | 39.9 | -- | 43.6 | 25.3 |
| TAU2-Bench | 43.2 | -- | 48.8 | 11.6 |
| Multilingualism (Thinking) | ||||
| MMMLU | 70.8 | 57.0 | 63.1 | 44.3 |
| MMLU-ProX | 62.4 | 49.4 | 52.3 | 34.6 |
| NOVA-63 | 47.1 | 40.3 | 46.4 | 42.4 |
| INCLUDE | 64.4 | 51.8 | 55.4 | 40.6 |
| Global PIQA | 73.5 | 63.1 | 69.3 | 59.4 |
| PolyMATH | 46.2 | 25.2 | 26.1 | 8.2 |
| WMT24++ | 58.9 | 39.3 | 45.8 | 27.2 |
| MAXIFE | 72.1 | 50.7 | 60.6 | 39.2 |
| Qwen3-VL-4B | Qwen3-VL-2B | Qwen3.5-2B | Qwen3.5-0.8B | |
|---|---|---|---|---|
| STEM and Puzzle | ||||
| MMMU | 70.8 | 61.4 | 64.2/64.2 | 49/47.4 |
| MMMU-Pro | 57.0 | 42.5 | 50.3/47.7 | 31.2/31.4 |
| Mathvista(mini) | 79.5 | 73.6 | 76.7/73.9 | 62.2/58.6 |
| DynaMath | 74.4 | 66.7 | 73.6/69.6 | 49.9/46.5 |
| ZEROBench | 0.0 | 0.0 | 1.0/0.0 | 0.0/0.0 |
| ZEROBench_sub | 18.9 | 13.2 | 17.1/18.6 | 12.9/11.4 |
| VlmsAreBlind | 68.6 | 50.0 | 75.8/74.3 | 59.4/57.3 |
| General VQA | ||||
| RealWorldQA | 73.2 | 69.5 | 74.5/71.2 | 63.4/61.6 |
| MMStar | 73.2 | 68.1 | 71.7/68.0 | 58.3/55.9 |
| MMBenchEN-DEV-v1.1 | 86.7 | 81.9 | 83.3/81.3 | 69.9/68.0 |
| SimpleVQA | 48.8 | 43.6 | 38.5/39.5 | 31.3/30.4 |
| HallusionBench | 64.1 | 54.9 | 58.0/51.3 | 53.1/46.7 |
| Text Recognition and Document Understanding | ||||
| MMLongBench-Doc | 44.4 | 33.8 | 45.4/38.8 | 33.6/28.1 |
| AI2D_TEST | 84.9 | 80.4 | 83.3/81.5 | 69.9/68.7 |
| CC-OCR | 73.8 | 68.3 | 72.9/75.8 | 63.2/66.7 |
| OmniDocBench1.5 | 80.0 | 65.9 | 79.8/80.9 | 61.0/70.6 |
| CharXiv(RQ) | 50.3 | 37.1 | 58.8/52.6 | 41.3/38.2 |
| OCRBench | 80.8 | 79.2 | 84.5/85.4 | 74.5/79.1 |
| Spatial Intelligence | ||||
| RefCOCO(avg) | 88.2 | 84.8 | 84.8/84.3 | 79.3/77.8 |
| CountBench | 89.4 | 84.1 | 91.4/86.8 | 77.0/68.6 |
| ODInW13 | 39.4 | 36.0 | 35.9/40.5 | 31.6/33.2 |
| ERQA | 47.3 | 41.8 | 43.8/33.0 | 34.5/23.8 |
| EmbSpatialBench | 80.7 | 75.9 | 77.9/66.4 | 68.6/54.6 |
| RefSpatialBench | 45.3 | 28.9 | 32.9/30.0 | 23.5/21.7 |
| Hypersim | 11.9 | 11.2 | 12.4/12.4 | 11.9/11.0 |
| SUNRGBD | 28.0 | 28.6 | 28.7/25.6 | 26.1/23.3 |
| Nuscene | 4.9 | 4.0 | 6.9/8.5 | 5.7/7.0 |
| Video Understanding | ||||
| VideoMME(w sub.) | 76.0 | 67.9 | 75.6/-- | 63.8/-- |
| VideoMME(w/o sub.) | 68.9 | 62.1 | 69.0/-- | 57.7/-- |
| VideoMMMU | 69.4 | 54.1 | 62.1/-- | 44.3/-- |
| MLVU | 75.7 | 69.2 | 76.2/-- | 65.6/-- |
| MVBench | 69.3 | 64.5 | 64.9/-- | 55.8/-- |
| LVBench | 53.5 | 47.6 | 57.1/-- | 45.1/-- |
| MMVU | 58.6 | 48.9 | 48.6/-- | 34.3/-- |
| Visual Agent | ||||
| ScreenSpot Pro | 59.5 | 48.5 | --/54.5 | --/46.5 |
| Medical VQA | ||||
| SLAKE | 65.9 | 61.1 | 74.4/67.5 | 62.6/59.5 |
| PMC-VQA | 48.4 | 42.4 | 48.8/54.0 | 40.4/45.5 |
| MedXpertQA-MM | 26.3 | 13.0 | 26.9/19.1 | 17.1/25.3 |
[!Important] Qwen3.5 models support both non-thinking and thinking mode. Qwen3.5-0.8B operates in non-thinking mode by default. To enable thinking, refer to the examples here.
[!Important] Inference efficiency and throughput vary significantly across frameworks. We recommend using the latest framework versions to ensure optimal performance and compatibility. For production workloads or high-throughput scenarios, dedicated serving engines such as SGLang, KTransformers or vLLM are strongly recommended.
[!Important] The model has a default context length of 262,144 tokens. If you encounter out-of-memory (OOM) errors, consider reducing the context window.
uv pip install 'git+https://github.com/sgl-project/sglang.git#subdirectory=python&egg=sglang[all]'http://localhost:8000/v1:python -m sglang.launch_server --model-path Qwen/Qwen3.5-0.8B --port 8000 --tp-size 1 --mem-fraction-static 0.8 --context-length 262144python -m sglang.launch_server --model-path Qwen/Qwen3.5-0.8B --port 8000 --tp-size 1 --mem-fraction-static 0.8 --context-length 262144 --tool-call-parser qwen3_coderpython -m sglang.launch_server --model-path Qwen/Qwen3.5-0.8B --port 8000 --tp-size 1 --mem-fraction-static 0.8 --context-length 262144 --speculative-algo NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4uv pip install vllm --torch-backend=auto --extra-index-url https://wheels.vllm.ai/nightlyhttp://localhost:8000/v1:vllm serve Qwen/Qwen3.5-0.8B --port 8000 --tensor-parallel-size 1 --max-model-len 262144 vllm serve Qwen/Qwen3.5-0.8B --port 8000 --tensor-parallel-size 1 --max-model-len 262144 --enable-auto-tool-choice --tool-call-parser qwen3_coder vllm serve Qwen/Qwen3.5-0.8B --port 8000 --tensor-parallel-size 1 --max-model-len 262144 --speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":2}'vllm serve Qwen/Qwen3.5-0.8B --port 8000 --tensor-parallel-size 1 --max-model-len 262144 --language-model-onlytransformers is required for Qwen3.5:pip install "transformers[serving] @ git+https://github.com/huggingface/transformers.git@main"transformers serve to launch a server with API endpoints at http://localhost:8000/v1; it will place the model on accelerators if available:transformers serve --force-model Qwen/Qwen3.5-0.8B --port 8000 --continuous-batching1pip install -U openai
2
3# Set the following accordingly
4export OPENAI_BASE_URL="http://localhost:8000/v1"
5export OPENAI_API_KEY="EMPTY"[!Tip] We recommend using the following set of sampling parameters for generation
- Non-thinking mode for text tasks:
temperature=1.0, top_p=1.00, top_k=20, min_p=0.0, presence_penalty=2.0, repetition_penalty=1.0- Non-thinking mode for VL tasks:
temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0- Thinking mode for text tasks:
temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0- Thinking mode for VL or precise coding (e.g. WebDev) tasks :
temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0Please note that the support for sampling parameters varies according to inference frameworks.
1from openai import OpenAI
2# Configured by environment variables
3client = OpenAI()
4
5messages = [
6 {"role": "user", "content": "Give me a short introduction to large language models."},
7]
8
9chat_response = client.chat.completions.create(
10 model="Qwen/Qwen3.5-0.8B",
11 messages=messages,
12 max_tokens=32768,
13 temperature=1.0,
14 top_p=1.0,
15 presence_penalty=2.0,
16 extra_body={
17 "top_k": 20,
18 },
19)
20print("Chat response:", chat_response)1from openai import OpenAI
2# Configured by environment variables
3client = OpenAI()
4
5messages = [
6 {
7 "role": "user",
8 "content": [
9 {
10 "type": "image_url",
11 "image_url": {
12 "url": "https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/RealWorld/RealWorld-04.png"
13 }
14 },
15 {
16 "type": "text",
17 "text": "Where is this?"
18 }
19 ]
20 }
21]
22
23chat_response = client.chat.completions.create(
24 model="Qwen/Qwen3.5-0.8B",
25 messages=messages,
26 max_tokens=32768,
27 temperature=0.7,
28 top_p=0.8,
29 presence_penalty=1.5,
30 extra_body={
31 "top_k": 20,
32 },
33)
34print("Chat response:", chat_response)1from openai import OpenAI
2# Configured by environment variables
3client = OpenAI()
4
5messages = [
6 {
7 "role": "user",
8 "content": [
9 {
10 "type": "video_url",
11 "video_url": {
12 "url": "https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/video/N1cdUjctpG8.mp4"
13 }
14 },
15 {
16 "type": "text",
17 "text": "Summarize the video content."
18 }
19 ]
20 }
21]
22
23# When vLLM is launched with `--media-io-kwargs '{"video": {"num_frames": -1}}'`,
24# video frame sampling can be configured via `extra_body` (e.g., by setting `fps`).
25# This feature is currently supported only in vLLM.
26#
27# By default, `fps=2` and `do_sample_frames=True`.
28# With `do_sample_frames=True`, you can customize the `fps` value to set your desired video sampling rate.
29chat_response = client.chat.completions.create(
30 model="Qwen/Qwen3.5-0.8B",
31 messages=messages,
32 max_tokens=32768,
33 temperature=0.7,
34 top_p=0.8,
35 presence_penalty=1.5,
36 extra_body={
37 "top_k": 20,
38 "mm_processor_kwargs": {"fps": 2, "do_sample_frames": True},
39 },
40)
41
42print("Chat response:", chat_response)[!Important] Qwen3.5 does not officially support the soft switch of Qwen3, i.e.,/thinkand/nothink.
1from openai import OpenAI
2# Configured by environment variables
3client = OpenAI()
4
5messages = [
6 {"role": "user", "content": "Type \"I love Qwen3.5\" backwards"},
7]
8
9chat_response = client.chat.completions.create(
10 model="Qwen/Qwen3.5-0.8B",
11 messages=messages,
12 max_tokens=81920,
13 temperature=1.0,
14 top_p=0.95,
15 presence_penalty=1.5,
16 extra_body={
17 "top_k": 20,
18 "enable_thinking": True,
19 },
20)
21print("Chat response:", chat_response)[!Important] In thinking mode, we have observed that when using the recommended sampling parameters, Qwen3.5-0.8B is more prone to entering thinking loops compared to other Qwen3.5 models, which may prevent it from terminating generation properly. We recommend further tuning the sampling parameters specific to your use case and utilizing the API's streaming generation mode (if supported) to enable timely detection and interruption of such anomalous generation behaviors.
1import os
2from qwen_agent.agents import Assistant
3
4# Define LLM
5# Using OpenAI-compatible API endpoint. The API backend should disable response parsers.
6llm_cfg = {
7 # Use your own model service compatible with OpenAI API by vLLM/SGLang:
8 'model': 'Qwen/Qwen3.5-0.8B',
9 'model_type': 'qwenvl_oai',
10 'model_server': 'http://localhost:8000/v1', # api_base
11 'api_key': 'EMPTY',
12
13 'generate_cfg': {
14 'use_raw_api': True,
15 # Pass the parameter of whether to enable thinking mode in this way
16 # 'extra_body': {
17 # 'chat_template_kwargs': {'enable_thinking': True}
18 # },
19 },
20}
21
22# Define Tools
23tools = [
24 {'mcpServers': { # You can specify the MCP configuration file
25 "filesystem": {
26 "command": "npx",
27 "args": ["-y", "@modelcontextprotocol/server-filesystem", "/Users/xxxx/Desktop"]
28 }
29 }
30 }
31]
32
33# Define Agent
34bot = Assistant(llm=llm_cfg, function_list=tools)
35
36# Streaming generation
37messages = [{'role': 'user', 'content': 'Help me organize my desktop.'}]
38for responses in bot.run(messages=messages):
39 pass
40print(responses)
41
42# Streaming generation
43messages = [{'role': 'user', 'content': 'Develop a dog website and save it on the desktop'}]
44for responses in bot.run(messages=messages):
45 pass
46print(responses)temperature=1.0, top_p=1.00, top_k=20, min_p=0.0, presence_penalty=2.0, repetition_penalty=1.0temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0presence_penalty parameter between 0 and 2 to reduce endless repetitions. However, using a higher value may occasionally result in language mixing and a slight decrease in model performance.answer field with only the choice letter, e.g., "answer": "C"."size parameter in the released video_preprocessor_config.json is conservatively configured. It is recommended to set the longest_edge parameter in the video_preprocessor_config file to 469,762,048 (corresponding to 224k video tokens) to enable higher frame-rate sampling for hour-scale videos and thereby achieve superior performance. For example,{"longest_edge": 469762048, "shortest_edge": 4096}1@misc{qwen3.5,
2 title = {{Qwen3.5}: Towards Native Multimodal Agents},
3 author = {{Qwen Team}},
4 month = {February},
5 year = {2026},
6 url = {https://qwen.ai/blog?id=qwen3.5}
7}