Views
No views yet

| Architecture | Hybrid-linear MoE |
|---|---|
| Parameter Scale | Total 124B, Activated 5.1B |
| Transformer Layers | 35 KDA + 7 Gated MLA (5:1) |
| Number of Dense Layers | 2 |
| Number of Routed Experts | 512 |
| Number of Shared Experts | 1 |
| Number of Activated Experts | 8 |
| Attention Heads | 32 |
| Hidden Size | 2560 |
| Expert Intermediate Size | 768 |
| Dense Intermediate Size | 6144 |
| Vocabulary Size | 157184 |
| Context Training Schedule | 8K -> 32K -> 256K |


- Thinking mode is enabled by default. Unless otherwise specified, the default parameters for Ling-3.0-flash are as follows:
temperature=0.6, top_p=0.95, top_k=20.- SWE-Bench Series: Evaluated using OpenHands as the agent harness with tailored prompts. Decoding uses
temperature=0.6, top_p=0.95, max_new_tokens=32K, with a 256K context window.- Terminal-Bench 2.1: Evaluated under the Artificial Analysis (AA) protocol using the default Terminus 2 harness, a unified 2-hour timeout, the provided JSON parser in preserve-thinking mode, and 3 runs per task (mean). Decoding uses
temperature=0.6, top_p=1.0, max_new_tokens=32K, with a 256K context window.- MiniAppBench: A 500-task coding benchmark evaluating whether models can turn a single user request into complete, usable interactive HTML apps in real-world application-generation scenarios. Evaluated with
temperature=1.0, top_p=1.0, max_tokens=128K.- AntSWEBench: AntSWEBench is an internally used software engineering benchmark that covers mainstream programming languages such as Java, JavaScript, and Python, including various development scenarios like new feature, bug fix, and code refactoring.
- Tau3-banking-AA: Aligned with the AA leaderboard, utilizing GPT-5.4-mini (medium reasoning) for both the user simulator and the natural-language assertion judge.
- MCP-Atlas: Evaluated on the 500-task public set using the official v1 harness with a 20-turn limit and Gemini-2.5-Pro as the claim-coverage judger.
- SkillsBench: Evaluated via kilo-code on 87 tasks (excluding external API-dependent tasks), averaged over 3 runs.
- GDPval v2-AA: Evaluated on the public 220-task benchmark using the official Stirrup harness, with a 250-turn limit and a 5-hour timeout.
- Search-agent: For all search‑agent tasks, evaluations are performed using an internal harness. The basic ReAct paradigm is adopted for single-agent evaluation, while a multi-agent setup is employed for BrowseComp. The reported metric is the average pass@1.
- WideSearch: Evaluated using the official prompt and the official judge model GPT-4.1 on the corrected version of the dataset.
- Draco: Scored based on official rubrics per question, with the final score calculated as the average across all questions using Claude Opus 4.6 as the scoring model.
- BrowseComp (Single-Agent): Evaluated using a resume strategy for context management: once the context reaches a 64K-token threshold, the trajectory is summarized, the original history is discarded, and execution is resumed from the summary.
- BrowseComp (Multi-Agent): Evaluated using an internal multi-agent search harness based on SearchSwarm/Tongyi DeepResearch, configured with
temperature=0.85, top_p=0.95, max_tokens=8K, and main/sub-agent context windows of 128K and 64K, respectively.
| dataset | BF16 | FP8 | INT4 | FP4 |
|---|---|---|---|---|
| GPQA-diamond | 84.97 | 84.00 | 83.65 | 83.68 |
| IFBench | 74.49 | 73.40 | 72.20 | 73.87 |
| SciCode | 41.24 | 40.37 | 39.35 | 38.43 |
| ArcPrize | 68.75 | 67.18 | 67.56 | 68.06 |
1pip install uv
2
3uv venv ~/my_ling_env
4
5source ~/my_ling_env/bin/activate
6
7git clone -b ling_v3_support_mxfp4 https://github.com/inclusionAI/sglang.git
8
9cd sglang
10
11pip install --upgrade pip
12
13pip install -e "python"
14
15pip install flashinfer-cubin==0.6.16.post1 --index-url https://flashinfer.ai/whl
16
17pip install flashinfer-python==0.6.16.post1${MASTER_IP} and server port is ${PORT}:1export SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1
2export SGLANG_JIT_DEEPGEMM_PRECOMPILE=1
3export SGLANG_ENABLE_SPEC_V2=1
4export FLASHINFER_DISABLE_VERSION_CHECK=1
5python -m sglang.launch_server \
6 --model-path "$MODEL_PATH" \
7 --trust-remote-code \
8 --nnodes 1 \
9 --dist-init-addr "$MASTER_IP:2345" \
10 --port "$PORT" \
11 --tp-size 2 \
12 --ep-size 1 \
13 --max-running-requests 64 \
14 --max-mamba-cache-size 320 \
15 --chunked-prefill-size 8192 \
16 --allow-auto-truncate \
17 --context-length 262144 \
18 --random-seed 308534008 \
19 --attention-backend trtllm_mla \
20 --disable-flashinfer-autotune \
21 --mem-fraction-static 0.85 \
22 --fp8-gemm-backend cutlass \
23 --tool-call-parser ling3 \
24 --reasoning-parser ling3 \
25 --moe-runner-backend flashinfer_mxfp4 \
26 --flashinfer-mxfp4-moe-precision default \
27 --enable-fp32-lm-head \
28 --disable-shared-experts-fusion \
29 --speculative-algorithm NEXTNtemperature=0.6, top_p=0.95, and top_k=20, and enabling enable_thinking for better performance.1curl -s http://${MASTER_IP}:${PORT}/v1/chat/completions \
2 -H "Content-Type: application/json" \
3 -d '{"model": "auto",
4 "messages": [{"role": "user", "content": "hello!"}],
5 "chat_template_kwargs": {"enable_thinking": true},
6 "stream": true,
7 "temperature": 0.6,
8 "top_k": 20,
9 "top_p": 0.95
10 }'