Views
No views yet

2026/01/15: 🤗 Released kanana-2-30b-a3b-2601 HF model weights.2026/01/15: 📕 Published blog posts (pre-training, post-training) about the development of Kanana-2 models.2025/12/19: 🤗 Released kanana-2-30b-a3b HF model weights and publised a teaser blog.[!NOTE] No Kakao user data was used for either pre-training or post-training.
| Model | Download |
|---|---|
| kanana-2-30b-a3b-base-2601* | 🤗 HuggingFace |
| kanana-2-30b-a3b-mid-2601* | 🤗 HuggingFace |
| kanana-2-30b-a3b-instruct-2601 | 🤗 HuggingFace |
| kanana-2-30b-a3b-thinking-2601 | 🤗 HuggingFace |
[object Object] (prior to mid-training) checkpoint to contribute to the research community.[object Object] is identical to kanana-2-30b-a3b-base.
| Benchmark | Metric | Shot | kanana-2-30b-a3b-mid-2601 | kanana-2-30b-a3b-base-2601 | kanana-1.5-32.5b-base | Qwen3-30B-A3B-Base* |
|---|---|---|---|---|---|---|
| General Tasks | ||||||
| MMLU | acc | 5 | 75.44 | 74.83 | 76.76 | 81.14 |
| MMLU-Pro | acc | 5 | 56.14 | 52.61 | 52.40 | 61.83 |
| BBH | acc | 3 | 79.76 | 76.46 | 81.54 | 79.97 |
| SimpleQA† | acc | 5 | 29.70 | 29.13 | 26.95 | 26.47 |
| Mathematics Tasks | ||||||
| MATH | em | 4 | 54.40 | 48.86 | 47.68 | 62.58 |
| GSM8K | em | 8 | 82.71 | 76.57 | 85.14 | 88.10 |
| Coding Tasks | ||||||
| HumanEval | pass@1 | 0 | 75.29 | 71.34 | 75.59 | 53.32 |
| MBPP | pass@1 | 3 | 62.39 | 60.21 | 65.96 | 72.58 |
| Korean Tasks | ||||||
| KMMLU | acc | 5 | 62.15 | 61.98 | 61.56 | 62.25 |
| KoSimpleQA† | acc | 5 | 49.70 | 49.40 | 45.70 | 26.33 |
| HAE-RAE Bench (v1.0) | acc | 5 | 88.73 | 88.91 | 90.65 | 72.04 |
| MATH-Ko‡ | em | 4 | 54.07 | 45.58 | 47.42 | 58.20 |
| GSM8K-Ko‡ | em | 8 | 77.48 | 70.43 | 81.43 | 88.10 |
| MBPP-Ko§ | pass@1 | 3 | 61.55 | 57.29 | 65.41 | 66.84 |
| Long Context Tasks | ||||||
| RULER-4K | acc | 0 | 93.09 | 92.49 | 86.39 | 94.32 |
| RULER-8K | acc | 0 | 92.29 | 92.14 | 90.16 | 92.16 |
| RULER-16K | acc | 0 | 90.73 | 90.01 | 85.88 | 91.28 |
| RULER-32K | acc | 0 | 88.63 | 87.92 | 81.62 | 88.32 |
| Benchmark | Metric | kanana-2-30b-a3b-instruct-2601 | kanana-2-30b-a3b-instruct | kanana-1.5-32.5b-instruct | Qwen3-30B-A3B-Instruct-2507* | Qwen3-30B-A3B (non-thinking)* |
|---|---|---|---|---|---|---|
| Chat | ||||||
| MT-Bench | judge† | 8.30 | 8.42 | 8.23 | 8.71 | 8.38 |
| KoMT-Bench | judge† | 8.21 | 8.24 | 7.94 | 8.49 | 7.89 |
| Instruction Following | ||||||
| IFEval | prompt strict | 87.25 | 84.47 | 79.48 | 82.62 | 84.10 |
| IFBench | prompt strict | 48.30 | 41.84 | 38.78 | 30.27 | 29.25 |
| Multi-IF (EN) | acc | 77.88 | 75.81 | 68.51 | 77.93 | 81.03 |
| Multi-Challenge | acc | 35.16 | 34.80 | 19.05 | 41.76 | 27.84 |
| Tool Calling | ||||||
| BFCL-v3 (Live‡) | pass@1 | 76.66 | 74.30 | 68.74 | 73.93 | 69.14 |
| BFCL-v3 (Multi-Turn‡) | pass@1 | 38.63 | 35.38 | 11.38 | 38.77 | 11.88 |
| Code Generation | ||||||
| HumanEval+ | pass@1 | 81.10 | 79.88 | 79.88 | 86.59 | 87.20 |
| MBPP+ | pass@1 | 73.02 | 73.81 | 71.96 | 75.13 | 75.13 |
| Mathematics | ||||||
| GSM8K | em | 93.10 | 91.89 | 91.58 | 93.56 | 93.33 |
| MATH | acc | 88.56 | 86.26 | 77.92 | 90.96 | 87.20 |
| Reasoning & Knowledge | ||||||
| MMLU | em | 81.61 | 80.80 | 82.75 | 87.13 | 85.60 |
| KMMLU | em | 68.26 | 67.32 | 65.75 | 67.56 | 63.49 |
| GPQA Diamond | pass@1 | 52.53 | 42.93 | 42.42 | 54.55 | 50.51 |
| HAERAE-Bench (v1.0) | em | 75.57 | 75.57 | 65.34 | 53.41 | 57.39 |
[object Object] as the judge model.[object Object] denotes the average score of 6 live benchmarks, and [object Object] denotes the average score of 4 multi-turn benchmarks.
| Benchmark | Metric | kanana-2-30b-a3b-thinking-2601 | kanana-2-30b-a3b-thinking | Qwen3-30B-A3B-Thinking-2507* | Qwen3-30B-A3B (thinking)* |
|---|---|---|---|---|---|
| Reasoning & Knowledge | |||||
| MMLU-Pro | pass@1 | 74.2 | 75.3 | 80.8 | 78.5 |
| GPQA Diamond | pass@1 | 57.8 | 61.3 | 70.6 | 62.6 |
| Competition Math | |||||
| AIME 2025 | pass@1 | 74.0 | 72.7 | 82.3 | 70.7 |
| AIME 2024 | pass@1 | 79.0 | 78.3 | 91.0 | 82.7 |
| AIME 2024-Ko† | pass@1 | 75.0 | 25.3 | 80.3 | 72.3 |
| Code Generation | |||||
| LiveCodeBench | pass@1 | 58.8 | 60.8 | 68.3 | 62.3 |
| LiveCodeBench-Ko‡ | pass@1 | 51.2 | 9.4 | 66.3¶ | 61.5¶ |
| Instruction Following | |||||
| IFEval | prompt strict | 82.2 | 82.2 | 87.8 | 86.1 |
| IFBench | prompt strict | 47.8 | 42.3 | 47.6 | 36.7 |
| Tool Calling | |||||
| BFCL-v3 (Live§) | pass@1 | 75.9 | 75.6 | 82.9 | 80.3 |
| BFCL-v3 (Multi-Turn§) | pass@1 | 43.7 | 34.3 | 53.6 | 35.6 |
[object Object] denotes the average score of 6 live benchmarks, and [object Object] denotes the average score of 4 multi-turn benchmarks.[!NOTE] For optimal results with the reasoning model, please adhere to the default parameters:temperature=0.6,top_p=0.95,top_k=20. We strongly advise against greedy decoding, as it may lead to performance degradation and infinite repetition loops.
vllm serve kakaocorp/kanana-2-30b-a3b-instruct-2601 --enable-auto-tool-choice --tool-call-parser hermesvllm serve kakaocorp/kanana-2-30b-a3b-thinking-2601 --reasoning-parser deepseek_r1 --enable-auto-tool-choice --tool-call-parser hermespython3 -m sglang.launch_server --model-path kakaocorp/kanana-2-30b-a3b-instruct-2601 --tool-call-parser qwenpython3 -m sglang.launch_server --model-path kakaocorp/kanana-2-30b-a3b-thinking-2601 --reasoning-parser deepseek-r1 --tool-call-parser qwenconfig.json uploaded to HuggingFace is configured for token lengths of 32,768 or less. To process tokens beyond this length, YaRN must be applied. By updating the config.json with the following parameters, you can apply YaRN to handle token sequences up to 128K in length:1"rope_scaling": {
2 "beta_fast": 32,
3 "beta_slow": 1,
4 "factor": 4.0,
5 "mscale": 1.0,
6 "mscale_all_dim": 1.0,
7 "original_max_position_embeddings": 32768,
8 "type": "yarn",
9},vllmvllm serve ... --hf-overrides '{"max_position_embeddings": 131072, "rope_scaling": {"rope_type":"deepseek_yarn","factor":4.0,"beta_fast":32,"beta_slow":1,"mscale":1.0,"mscale_all_dim":1.0,"original_max_position_embeddings":32768}}'sglangpython3 -m sglang.launch_server ... --json-model-override-args '{"max_position_embeddings":131072, "rope_scaling":{"rope_type":"deepseek_yarn","factor":4.0,"beta_fast":32,"beta_slow":1,"mscale":1.0,"mscale_all_dim":1.0,"original_max_position_embeddings":32768}}'[!NOTE] Most leading open-source implementations of static YaRN apply a constant scaling factor, which can negatively impact performance on shorter texts. To ensure optimal performance:
- Enable
rope_scalingonly when necessary for processing long contexts.- Adjust the
factorbased on your specific needs (e.g., setfactorto 2.0 for a 65,536-token context)."
@article{,
title={Kanana-2 LLM},
author={Kanana LLM},
year={2025},
url={https://huggingface.co/collections/kakaocorp/kanana-2}
}