Views
No views yet



| Benchmark Metric: Pass@1 | Qwen3-8B | Nemotron-Nano-9B-v2 | DeepSeek-R1-0528 671B | Gemini-2.5-Flash-Thinking | Nemotron- Cascade-8B- Thinking | Nemotron- Cascade-8B |
|---|---|---|---|---|---|---|
| Knowledge Reasoning | ||||||
| MMLU | 83.0 | 82.6 | 89.9 | - | 84.0 | 83.7 |
| MMLU Pro | 75.1 | 73.3 | 85.0 | 81.9 | 75.5 | 75.7 |
| GPQA-Diamond | 62.0 | 64.0 | 81.0 | 82.8 | 66.7 | 66.5 |
| Alignment | ||||||
| ArenaHard | 85.8 | 74.6 | 95.1 | 95.7 | 85.8 | 87.9 |
| IFEval (Strict Prompt) | 85.0 | 86.1 | 84.1 | 89.8 | 83.7 | 90.2 |
| IFBench | 34.4 | 37.4 | 38.0 | 36.1 | 41.4 | 40.8 |
| Math | ||||||
| AIME 2024 | 76.0 | 81.9 | 91.4 | 82.3 | 88.8 | 89.5 |
| AIME 2025 | 67.3 | 72.0 | 87.5 | 72.0 | 81.4 | 80.1 |
| Code | ||||||
| LCB v5 (08/24-02/25) | 61.2 | 68.2 | 74.8 | 63.4 | 74.5 | 74.3 |
| LCB v6 (08/24-05/25) | 58.3 | 65.3 | 73.3 | 61.9 | 71.4 | 71.1 |
| LCB Pro 25Q2 (Easy) | 46.1 | 59.3 | 63.9 | 47.4 | 64.8 | 65.7 |
| LCB Pro 25Q2 (Med) | 2.2 | 4.8 | 7.0 | 1.8 | 6.1 | 6.4 |
| SWE Verified (Agentless) | 20.5 | - | 57.6 | 48.9 | 38.5 | 37.2 |
| Tool Calling | ||||||
| BFCL V3 | 68.1 | 66.9 | 67.9 | 68.6 | 67.0 | 64.4 |
" /think" tag should be appended to the end of the user input. Note that a leading space is included in this tag to ensure correct tokenization." /think" tag to " /no_think".1from transformers import AutoTokenizer
2
3model_name = 'nvidia/Nemotron-Cascade-8B-Thinking'
4tokenizer = AutoTokenizer.from_pretrained(model_name)
5
6'''
7single-turn example
8'''
9messages = [
10 {"role": "user", "content": "calculate 1+1?"}
11]
12
13# only thinking mode is supported (enable_thinking=True)
14prompt_thinking = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=True)
15# prompt_thinking = '<|im_start|>system\nYou are a helpful and harmless assistant.<|im_end|>\n<|im_start|>user\ncalculate 1+1? /think<|im_end|>\n<|im_start|>assistant\n'
16
17
18'''
19multi-turn example
20'''
21messages = [
22 {"role": "user", "content": "calculate 1+1?"},
23 {"role": "assistant", "content": "<think>THINKING_CONTENT</think>\nTo calculate \\(1 + 1\\):\n\n1. **Identify the operation**: This is a basic addition problem involving two integers.\n2. **Perform the addition**: \n \\(1 + 1 = 2\\).\n\n**Result**: \\(\\boxed{2}\\)",},
24 {"role": "user", "content": "what about 2+2"}
25]
26
27# only thinking mode is supported (enable_thinking=True)
28prompt_thinking = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=True)
29# prompt_thinking = '<|im_start|>system\nYou are a helpful and harmless assistant.<|im_end|>\n<|im_start|>user\ncalculate 1+1? /no_think<|im_end|>\n<|im_start|>assistant\nTo calculate \\(1 + 1\\):\n\n1. **Identify the operation**: This is a basic addition problem involving two integers.\n2. **Perform the addition**: \n \\(1 + 1 = 2\\).\n\n**Result**: \\(\\boxed\{2\}\\)<|im_end|>\n<|im_start|>user\nwhat about 2+2 /think<|im_end|>\n<|im_start|>assistant\n'@article{Nemotron_Cascade_Scaling_Cascaded_Reinforcement_Learning,
title={Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models},
author={Wang, Boxin and Lee, Chankyu and Lee, Nayeon and Lin, Sheng-Chieh and Dai, Wenliang and Chen, Yang and Chen, Yangyi and Yang, Zhuolin and Liu, Zihan and Shoeybi, Mohammad and Catanzaro, Bryan and Ping, Wei},
year={2025}
}