Table of Contents
Model Introduction
Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tencent Hy Team. Following the Hy3 Preview launch in late April, we gathered feedback from 50+ products and scaled up post-training with higher quality data. Today, we introduce Hy3, which outperforms similar-size models and rivals flagship open-source models with 2-5x parameters. It also shows significant gains in utility across various products and productivity tasks.
| Property | Value |
|---|
| Architecture | Mixture-of-Experts (MoE) |
| Total Parameters | 295B |
| Activated Parameters | 21B |
| MTP Layer Parameters | 3.8B |
| Number of Layers (excluding MTP layer) | 80 |
| Number of MTP Layers | 1 |
| Attention Heads | 64 (GQA, 8 KV heads, head dim 128) |
| Hidden Size | 4096 |
| Intermediate Size | 13312 |
| Context Length | 256K |
| Vocabulary Size | 120832 |
| Number of Experts | 192 experts, top-8 activated |
| Supported Precisions | BF16 |
Stronger Agent Capabilities
Building on Hy3 Preview, we further improved the quality and diversity of post-training data while scaling up RL training. Hy3 shows solid gains across reasoning, agentic, and long-context tasks, competitive with much larger flagship models.
In productivity scenarios such as coding, office work, financial modeling, frontend design, and game development, Hy3 has made remarkable progress and can now serve as a reliable, cost-effective model option.
We don't think public benchmark scores tell the full story. So we ran a blind evaluation with 270 experts using tasks from their work, and Hy3 scored 2.67/4, outperforming GLM-5.1 at 2.51/4. The advantage was most substantial in frontend development, data & storage, and CI/CD tasks.
More Reliable Product Experiences
Model usefulness is not fully captured by benchmarks. Based on extensive product feedback, we identified and fixed the following issues, receiving consistently positive feedback from product teams.
Stability of tool calls and output formats: We fixed multiple baseline reliability issues, bringing the model to production-grade standards across tool configurations and output constraints. Tool-call error recovery and overall efficiency improved. Hy3 also generalizes across different agent scaffoldings. On SWE-Bench Verified, accuracy variance across scaffoldings like CodeBuddy, Cline, and KiloCode remains within 4%.
Knowledge and anti-hallucination: Guided by the ideal of "answer when grounded, state when evidence is missing, do not conflate sources or fabricate data," we implemented fine-grained data cleaning and training constraints. In internal evaluations based on real-world scenarios, Hy3's hallucination rate dropped from 12.5% to 5.4%, and commonsense error rates fell from 25.4% to 12.7%. These improvements materially reduce fact conflation, fabrication, and logical contradiction.
Complex context retention and multi-turn intent tracking: Through joint optimization of SFT and RL, Hy3 improved on operational pain points like coreference resolution, ellipsis recovery, and multi-turn constraint inheritance. On internal comprehensive multi-turn tests, the issue rate dropped from 17.4% to 7.9%. Hy3 also improved markedly on long-dialogue evals like MRCR. Its outputs are more concise while ensuring complex intents do not decay or drift over long-horizon interactions.
Benchmark Appendix
News
- 🔥 We open-source Hy3 and Hy3-FP8 model weights on Hugging Face, ModelScope, GitCode, and CNB.
Model Links
Quickstart
Deploy Hy3 with
vLLM or
SGLang first, then call the OpenAI-compatible API:
1from openai import OpenAI
2
3client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="EMPTY")
4
5response = client.chat.completions.create(
6 model="hy3",
7 messages=[
8 {"role": "user", "content": "Hello! Can you briefly introduce yourself?"},
9 ],
10 temperature=0.9,
11 top_p=1.0,
12 # reasoning_effort: "no_think" (default, direct response), "low", "high" (deep chain-of-thought)
13 extra_body={"chat_template_kwargs": {"reasoning_effort": "no_think"}},
14)
15print(response.choices[0].message.content)
Recommended parameters: temperature=0.9, top_p=1.0.
Reasoning mode: Set reasoning_effort to "high" for complex tasks (math, coding, reasoning) or "no_think" for direct responses.
See the
Deployment section below for how to start the API server.
Deployment
Hy3 has 295B parameters in total. To serve it on 8 GPUs, we recommend using H20-3e or other GPUs with larger memory capacity.
For production serving, we recommend using vLLM or SGLang, both of which provide dedicated recipes for Hy3:
vLLM
Build vLLM from source:
1uv venv --python 3.12 --seed --managed-python
2source .venv/bin/activate
3git clone https://github.com/vllm-project/vllm.git
4cd vllm
5uv pip install --editable . --torch-backend=auto
Start the vLLM server with MTP enabled:
1# Switch to trtllm backend to work-around mnnvl workspace size issue.
2export VLLM_FLASHINFER_ALLREDUCE_BACKEND=trtllm
3vllm serve tencent/Hy3 \
4 --tensor-parallel-size 8 \
5 --speculative-config.method mtp \
6 --speculative-config.num_speculative_tokens 2 \
7 --tool-call-parser hy_v3 \
8 --reasoning-parser hy_v3 \
9 --enable-auto-tool-choice \
10 --port 8000 \
11 --served-model-name hy3
SGLang
Build SGLang from source:
1git clone https://github.com/sgl-project/sglang
2cd sglang
3pip3 install pip --upgrade
4pip3 install "transformers>=5.6.0"
5pip3 install -e "python"
Launch SGLang server with MTP enabled:
1python3 -m sglang.launch_server \
2 --model tencent/Hy3 \
3 --tp-size 8 \
4 --tool-call-parser hunyuan \
5 --reasoning-parser hunyuan \
6 --speculative-num-steps 2 \
7 --speculative-eagle-topk 1 \
8 --speculative-num-draft-tokens 3 \
9 --speculative-algorithm EAGLE \
10 --port 8000 \
11 --served-model-name hy3
Finetuning
Hy3 provides a complete model finetuning pipeline. For detailed documentation, please refer to:
Finetuning Guide
Quantization
We provide
AngelSlim, a more accessible, comprehensive, and efficient toolkit for large model compression. AngelSlim supports a comprehensive suite of compression tools for large-scale multimodal models, including common quantization algorithms, low-bit quantization, and speculative sampling.
License
Hy3 is released under the
Apache License 2.0. See
LICENSE for details.
Contact Us
If you would like to leave a message for our R&D and product teams, welcome to contact us. You can also reach us via email:
Hy3 is developed by the Tencent Hy Team.