Views
No views yet

| Total Parameters | 550B (55B active) |
| Architecture | LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) |
| Context Length | Up to 1M tokens |
| Minimum GPU Requirement | 8x GB200/B200/GB300/B300, 16x H100, 8x H200 |
| Supported Languages | English, French, Spanish, Italian, German, Japanese, Hindi, Korean, Brazilian Portuguese, and Chinese |
| Best For | Frontier reasoning, complex agentic workflows, long-context analysis, tool use, multilingual reasoning, high-stakes RAG |
| Reasoning Mode | Configurable on/off via chat template (enable_thinking=True/False) |
| License | OpenMDW License Agreement, version 1.1 |
| Release Date | June 4, 2026 |
For running Nemotron 3 Ultra on a smaller footprint, please see: NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
| Benchmark | N-3-Ultra 550B-A55B | MiniMax-2.7 230B-A10B | GLM-5.1 744B-A40B | Kimi-K2.6 1T-A32B | Qwen-3.5 397B-17B | DS-v4-Pro 1.6T-A49B | DS-v4-Flash 284B-A13B |
|---|---|---|---|---|---|---|---|
| Agentic | |||||||
| Terminal Bench 2.1 | 56.4 | 55.5 | 59.3 | 67.2 | 49.9 | 49.2 | 54.2 |
| GDPVal | 46.7 | 47.6 | 54.7 | 50.4 | 34.6 | 54.6 | 50.2 |
| SWE-Bench Verified | 70.7 | 75.3 | 76.2 | 75.7 | 73.6 | 74.5 | 73.5 |
| SWE-Bench Multilingual | 67.7 | 71.8 | 74.8 | 77.1 | 70.9 | 76.5 | 75.0 |
| ProfBench (Search) | 56.0 | 52.0 | 46.0 | 56.0 | 53.0 | 59.9 | 57.0 |
| PinchBench | 90.0 | 77.6 | 81.2 | 90.2 | 86.6 | 88.6 | 91.3 |
| TauBench V3 | |||||||
| Airline | 81.5 | 75.3 | 85.0 | 85.8 | 76.5 | 80.8 | 80.8 |
| Retail | 86.4 | 84.9 | 84.1 | 82.9 | 88.5 | 88.9 | 89.1 |
| Telecom | 92.9 | 89.6 | 96.9 | 97.8 | 98.0 | 96.3 | 98.3 |
| Banking | 22.6 | 14.6 | 12.8 | 23.1 | 20.9 | 25.9 | 26.7 |
| Average | 70.9 | 66.1 | 69.7 | 72.4 | 71.0 | 73.2 | 73.7 |
| BrowseComp | 44.4 | 54.1 | 59.4 | 61.3 | 40.5 | 59.4 | 46.9 |
| Vals.ai Financial Agent 1.1 | |||||||
| without web search | 60.1 | 51.3 | 60.2 | 54.0 | 61.3 | 58.9 | 58.4 |
| with web search | 53.7 | 50.5 | 60.7 | 58.8 | 59.0 | 62.3 | 60.1 |
| Reasoning and Knowledge | |||||||
| IOI 2025 | 570.0 | -- | 456.5 | 585.0 | 441.3 | 580.1 | -- |
| LiveCodeBench (v6) | 89.0 | 77.2 | 85.7 | 90.2 | 79.3 | 92.5 | 90.9 |
| IMOAnswerBench (no tools) | 88.6 | 68.3 | 86.8 | 91.1 | 83.1 | 93.0 | 91.1 |
| IMOAnswerBench (with tools) | 92.3 | 75.1 | 91.1 | 93.71 | 84.51 | 85.4 | 89.6 |
| Apex-Shortlist (no tools) | 74.9 | 28.9 | 71.1 | 77.4 | 61.4 | 85.8 | 82.4 |
| Apex-Shortlist (with tools) | 84.8 | 51.9 | 79.0 | 73.2 | 60.4 | 86.5 | 82.0 |
| GPQA (no tools) | 87.0 | 86.6 | 86.1 | 91.0 | 87.1 | 87.8 | 88.5 |
| SciCode (subtask) | 44.6 | 38.3 | 47.7 | 52.0 | 48.0 | 50.5 | 48.2 |
| HLE (no tools) | 26.7 | 23.1 | 27.2 | 34.8 | 28.5 | 37.7 | 32.2 |
| HLE (with tools) | 37.4 | -- | 50.4 | 54.0 | 48.3 | 48.2 | 45.1 |
| CritPt (no tools) | 3.1 | 0.6 | 3.7 | 9.1 | 2.4 | 14.0 | 10.6 |
| MMLU-Pro | 86.8 | 81.9 | 85.9 | 88.1 | 88.3 | 87.5 | 86.4 |
| OmniScience Accuracy | 24.1 | 20.5 | 31.3 | 35.5 | 35.9 | 46.8 | 39.9 |
| OmniScience Non-Hallucination | 78.7 | 74.4 | 66.8 | 67.1 | 7.4 | 5.7 | 2.8 |
| Chat & Instruction Following | |||||||
| IFBench (prompt loose) | 81.7 | 74.6 | 76.6 | 73.7 | 78.2 | 79.1 | 82.0 |
| Multi-Challenge | 63.8 | 42.5 | 63.0 | 63.1 | 63.9 | 64.1 | 63.5 |
| Long Context | |||||||
| AA-LCR | 65.4 | 69.8 | 66.9 | 70.2 | 68.3 | 67.3 | 62.7 |
| RULER (1M) | 94.7 | -- | -- | -- | 90.1 | 94.2 | 87.7 |
| Longbench v2 (≤ 1M) | 61.9 | -- | -- | -- | 68.9 | 62.1 | 57.0 |
| Multilingual | |||||||
| MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) | 83.0 | 78.4 | 85.8 | 85.0 | 86.4 | 85.6 | 84.3 |
| WMT24++ (en→xx) | 83.7 | 82.8 | 84.4 | 84.5 | 86.8 | 85.9 | 85.9 |
1# Set the IP for the head node in RAY_HEAD_IP
2export RAY_HEAD_IP=<head_node_ip>
3export RAY_PORT=6379
4export RAY_ADDRESS=${RAY_HEAD_IP}:${RAY_PORT}
5
6# Start Ray head node (vLLM/SGLang will run on this node)
7ray start --head --node-ip-address=${RAY_HEAD_IP} --port=${RAY_PORT}
8
9# Start Ray worker node(s)
10ray start --address=${RAY_HEAD_IP}:${RAY_PORT} --block
11
12# Verify Ray cluster is ready
13ray status --address=${RAY_HEAD_IP}:${RAY_PORT}ray[cgraph] is required: uv pip install "ray[cgraph]"vllm/vllm-openai:v0.22.0.export MODEL_CKPT=PATH/TO/MODEL/CHECKPOINT1docker run -d --name nemotron-ultra-vllm \
2 --gpus all \
3 --ipc=host \
4 --network=host \
5 --shm-size=16g \
6 --ulimit memlock=-1 \
7 --ulimit stack=67108864 \
8 -v $MODEL_CKPT:/model:ro \
9 -e VLLM_WORKER_MULTIPROC_METHOD=spawn \
10 -e SAFETENSORS_FAST_GPU=1 \
11 -e NVIDIA_TF32_OVERRIDE=1 \
12 -e VLLM_LOGGING_LEVEL=INFO \
13 vllm/vllm-openai:v0.22.0 \
14 /model \
15 --host 0.0.0.0 \
16 --port 8000 \
17 --served-model-name nvidia/nemotron-3-ultra \
18 --trust-remote-code \
19 --tensor-parallel-size 8 \
20 --enable-expert-parallel \
21 --dtype bfloat16 \
22 --max-model-len 262144 \
23 --gpu-memory-utilization 0.90 \
24 --max-num-seqs 16 \
25 --max-num-batched-tokens 32768 \
26 --enable-chunked-prefill \
27 --enable-prefix-caching \
28 --reasoning-parser nemotron_v3 \
29 --enable-auto-tool-choice \
30 --tool-call-parser qwen3_coder \
31 --mamba-ssm-cache-dtype float16 \
32 --mamba-backend flashinfer \
33 --enable-mamba-cache-stochastic-rounding \
34 --mamba-cache-philox-rounds 5 \
35 --speculative-config '{"method": "nemotron_h_mtp", "num_speculative_tokens": 5}' \
36 --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 96}'1# Run on Ray head node
2vllm serve $MODEL_CKPT \
3 --host 0.0.0.0 \
4 --port 8000 \
5 --served-model-name nvidia/nemotron-3-ultra \
6 --tensor-parallel-size 8 \
7 --distributed-executor-backend ray \
8 --trust-remote-code \
9 --dtype bfloat16 \
10 --gpu-memory-utilization 0.90 \
11 --max-model-len 262144 \
12 --max-num-seqs 256 \
13 --max-num-batched-tokens 32768 \
14 --enable-chunked-prefill \
15 --enable-prefix-caching \
16 --reasoning-parser nemotron_v3 \
17 --mamba-ssm-cache-dtype float16 \
18 --mamba-backend flashinfer \
19 --enable-mamba-cache-stochastic-rounding \
20 --mamba-cache-philox-rounds 5 \
21 --enable-auto-tool-choice \
22 --tool-call-parser qwen3_coder \
23 --speculative-config '{"method": "nemotron_h_mtp", "num_speculative_tokens": 5}' \
24 --kv-cache-dtype fp8 \
25 --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 96}' \
26 --compilation-config '{"pass_config": {"fuse_allreduce_rms": false}}' \
27 --distributed-timeout-seconds 3600VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 and --max-model-len 1048576.VLLM_FLASHINFER_ALLREDUCE_BACKEND=trtllm, VLLM_FLASHINFER_MOE_BACKEND=latency (TRTLLM-Gen) or VLLM_FLASHINFER_MOE_BACKEND=throughput (CUTLASS).docker pull lmsysorg/sglang:v0.5.12.post11docker run -d --name nemotron-ultra-sglang \
2 --gpus all \
3 --cap-add SYS_NICE \
4 --ipc=host \
5 --network=host \
6 --shm-size=16g \
7 --ulimit memlock=-1 \
8 --ulimit stack=67108864 \
9 -v $MODEL_CKPT:/model:ro \
10 -e SAFETENSORS_FAST_GPU=1 \
11 -e NVIDIA_TF32_OVERRIDE=1 \
12 -e SGLANG_DISABLE_DEEP_GEMM=1 \
13 lmsysorg/sglang:v0.5.12.post1 \
14 python3 -m sglang.launch_server \
15 --model-path /model \
16 --host 0.0.0.0 \
17 --port 8000 \
18 --served-model-name nvidia/nemotron-3-ultra \
19 --tp-size 8 \
20 --ep-size 8 \
21 --context-length 262144 \
22 --mem-fraction-static 0.85 \
23 --chunked-prefill-size 32768 \
24 --fp8-gemm-backend triton \
25 --moe-runner-backend triton \
26 --mamba-scheduler-strategy no_buffer \
27 --disable-piecewise-cuda-graph \
28 --reasoning-parser nemotron_v3 \
29 --tool-call-parser qwen3_coder \
30 --speculative-algorithm EAGLE \
31 --speculative-num-steps 5 \
32 --speculative-eagle-topk 1 \
33 --speculative-num-draft-tokens 5 \
34 --trust-remote-code \
35 --log-level infoSGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 and --context-length 1048576.tools, you must set "chat_template_kwargs": {"enable_thinking": true, "force_nonempty_content": true} in the request body to parse both reasoning and tool calls correctly.docker pull nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc171cat > ./extra-llm-api-config.yml << EOF
2backend: pytorch
3trust_remote_code: true
4tensor_parallel_size: 8
5pipeline_parallel_size: 1
6context_parallel_size: 1
7gpus_per_node: 8
8moe_expert_parallel_size: 1
9disable_overlap_scheduler: false
10
11cuda_graph_config:
12 enable_padding: true
13 max_batch_size: 256
14
15enable_chunked_prefill: true
16enable_attention_dp: false
17max_batch_size: 256
18max_seq_len: null
19max_num_tokens: 32768
20num_postprocess_workers: 4
21
22kv_cache_config:
23 enable_block_reuse: false
24 max_tokens: null
25 max_attention_window: null
26 sink_token_length: null
27 free_gpu_memory_fraction: 0.75
28 host_cache_size: null
29 cross_kv_cache_fraction: null
30 secondary_offload_min_priority: null
31 event_buffer_max_size: 0
32 attention_dp_events_gather_period_ms: 5
33 enable_partial_reuse: true
34 copy_on_partial_reuse: true
35 use_uvm: false
36 max_gpu_total_bytes: 0
37 iteration_stats_interval: 1
38 dtype: fp8
39 tokens_per_block: 32
40 mamba_state_cache_interval: 256
41 use_kv_cache_manager_v2: false
42 max_util_for_resume: 0.95
43
44moe_config:
45 backend: TRTLLM
46 max_num_tokens: null
47 load_balancer: null
48 disable_finalize_fusion: false
49 use_low_precision_moe_combine: false
50EOF
51
52
53TLLM_ALLOW_LONG_MAX_MODEL_LEN=1 trtllm-serve \
54<bf16_ckpt> \
55--max_batch_size 256 \
56--tp_size 8 --ep_size 1 \
57--max_num_tokens 32768 \
58--trust_remote_code \
59--reasoning_parser nano-v3 \
60--tool_parser qwen3_coder \
61--chat_template $MODEL_DIR/chat_template.jinja \
62--extra_llm_api_options extra-llm-api-config.ymlTLLM_ALLOW_LONG_MAX_MODEL_LEN=1 as an environment variable and add --max_seq_len <seq_len> as the desired maximum context length. The MTP speculative_config block above carries over unchanged — on rc16, max_draft_len is the authoritative field and num_nextn_predict_layers is treated as deprecated.extra_body={"chat_template_kwargs": {"force_nonempty_content": True}}1from openai import OpenAI
2client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
3MODEL = "nvidia/nemotron-3-ultra"1response = client.chat.completions.create(
2 model=MODEL,
3 messages=[{"role": "user", "content": "Write a haiku about GPUs"}],
4 max_tokens=16000,
5 temperature=1.0,
6 top_p=0.95,
7 extra_body={"chat_template_kwargs": {"enable_thinking": True}}
8)
9print(response.choices[0].message.content)1response = client.chat.completions.create(
2 model=MODEL,
3 messages=[{"role": "user", "content": "What is the capital of Japan?"}],
4 max_tokens=16000,
5 temperature=1.0,
6 top_p=0.95,
7 extra_body={"chat_template_kwargs": {"enable_thinking": False}}
8)
9print(response.choices[0].message.content)1response = client.chat.completions.create(
2 model=MODEL,
3 messages=[{"role": "user", "content": "What is the capital of Japan?"}],
4 max_tokens=16000,
5 temperature=1.0,
6 top_p=0.95,
7 extra_body={"chat_template_kwargs": {"enable_thinking": True, "medium_effort": True}}
8)
9print(response.choices[0].message.content)1response = client.chat.completions.create(
2 model=MODEL,
3 messages=[{"role": "user", "content": "What's the weather in New York?"}],
4 tools=[{
5 "type": "function",
6 "function": {
7 "name": "get_weather",
8 "description": "Get the current weather for a city.",
9 "parameters": {
10 "type": "object",
11 "properties": {
12 "city": {"type": "string"},
13 "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
14 },
15 "required": ["city"]
16 }
17 }
18 }],
19 tool_choice="required",
20 max_tokens=256,
21 temperature=1.0,
22 top_p=0.95,
23 extra_body={"chat_template_kwargs": {"enable_thinking": True, "force_nonempty_content": True}}
24)~/.config/opencode/opencode.json:1{
2 "$schema": "https://opencode.ai/config.json",
3 "model": "local/nvidia-nemotron-3-ultra",
4 "provider": {
5 "local": {
6 "npm": "@ai-sdk/openai-compatible",
7 "name": "local_backend",
8 "options": {
9 "baseURL": "http://localhost:8000/v1",
10 "apiKey": "EMPTY"
11 },
12 "models": {
13 "nvidia-nemotron-3-ultra": {
14 "name": "nvidia/nemotron-3-ultra",
15 "limit": {
16 "context": 1000000,
17 "output": 32768
18 }
19 }
20 }
21 }
22 },
23 "agent": {
24 "build": {
25 "temperature": 1.0,
26 "top_p": 0.95,
27 "max_tokens": 32000
28 },
29 "plan": {
30 "temperature": 1.0,
31 "top_p": 0.95,
32 "max_tokens": 32000
33 }
34 }
35}8000, so the baseURL works as-is for vLLM, SGLang, and TensorRT-LLM.reasoning_budget. The model will attempt to close the trace at the next newline before the budget is hit; if none is found within 500 tokens it closes abruptly at reasoning_budget + 500.1from typing import Any, Dict, List
2import openai
3from transformers import AutoTokenizer
4
5
6class ThinkingBudgetClient:
7 def __init__(self, base_url: str, api_key: str, tokenizer_name_or_path: str):
8 self.tokenizer = AutoTokenizer.from_pretrained(tokenizer_name_or_path)
9 self.client = openai.OpenAI(base_url=base_url, api_key=api_key)
10
11 def chat_completion(
12 self,
13 model: str,
14 messages: List[Dict[str, Any]],
15 reasoning_budget: int = 512,
16 max_tokens: int = 1024,
17 **kwargs,
18 ) -> Dict[str, Any]:
19 assert max_tokens > reasoning_budget, (
20 f"reasoning_budget must be less than max_tokens. "
21 f"Got {max_tokens=} and {reasoning_budget=}"
22 )
23
24 # Step 1: generate the reasoning trace up to the budget
25 response = self.client.chat.completions.create(
26 model=model, messages=messages, max_tokens=reasoning_budget, **kwargs
27 )
28 reasoning_content = response.choices[0].message.content
29 if "</think>" not in reasoning_content:
30 reasoning_content = f"{reasoning_content}.\n\n</think>\n"
31
32 reasoning_tokens_len = len(
33 self.tokenizer.encode(reasoning_content, add_special_tokens=False)
34 )
35 remaining_tokens = max_tokens - reasoning_tokens_len
36 assert remaining_tokens > 0, (
37 f"No tokens remaining for response ({remaining_tokens=}). "
38 "Increase max_tokens or lower reasoning_budget."
39 )
40
41 # Step 2: continue from the closed reasoning trace
42 messages.append({"role": "assistant", "content": reasoning_content})
43 prompt = self.tokenizer.apply_chat_template(
44 messages, tokenize=False, continue_final_message=True
45 )
46 response = self.client.completions.create(
47 model=model, prompt=prompt, max_tokens=remaining_tokens, **kwargs
48 )
49
50 return {
51 "reasoning_content": reasoning_content.strip().strip("</think>").strip(),
52 "content": response.choices[0].text,
53 "finish_reason": response.choices[0].finish_reason,
54 }1client = ThinkingBudgetClient(
2 base_url="http://localhost:8000/v1",
3 api_key="EMPTY",
4 tokenizer_name_or_path="nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16",
5)
6
7result = client.chat_completion(
8 model="nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16",
9 messages=[
10 {"role": "system", "content": "You are a helpful assistant. /think"},
11 {"role": "user", "content": "What is 2+2?"},
12 ],
13 reasoning_budget=32,
14 max_tokens=512,
15 temperature=1.0,
16 top_p=0.95,
17)
18print(result)| Dataset Collection | Token Counts | Description |
|---|---|---|
| Nemotron-CC-v2 & v2.1 | 9.1T | A massive collection of English web data filtered from Common Crawl, including 2.5T+ tokens of new organic, translated, and synthetically rephrased content. |
| Nemotron-CC-Code-v1 | 427.9B | High-quality code tokens extracted from Common Crawl using the Lynx + LLM pipeline to preserve structure and equations. |
| Nemotron-Pretraining-Code-v1 & v2 & v3 | 1.7T | Curated GitHub code references with multi-stage filtering, deduplication, and large-scale synthetic code data. |
| Nemotron-CC-Math-v1 | 133.3B | High-quality math pre-training dataset preserving LaTeX formatting and mathematical structures. |
| Nemotron-Pretraining-Specialized-v1 & v1.1 & v1.2 & Nemotron-Pretraining-SFT-v1 | 660.0B | Synthetic datasets targeting specialized domains such as STEM reasoning and scientific coding. |
| Nemotron-Pretraining-Legal-v1 | 4.3B | Synthetic datasets targeting the legal domain. |
| Dataset | Modality | Dataset Size | Collection Period | Collecting Organisation |
|---|---|---|---|---|
| English Common Crawl | Text | 3.36T | 4/8/2025 | NVIDIA Advanced Deep Learning Research |
| English Common Crawl 1.1 | Text | Not disclosed | 10/2/2025 | NVIDIA Advanced Deep Learning Research |
| Multilingual Common Crawl | Text | 812.7B | 5/1/2025 | NVIDIA Advanced Deep Learning Research |
| GitHub Crawl | Text | 747.4B | 4/29/2025 | NVIDIA Advanced Deep Learning Research |
| GitHub Crawl 1.1 | Text | 172.7B | 9/30/2025 | NVIDIA Advanced Deep Learning Research |
| Dataset | Model(s) used |
|---|---|
| Global Regulation | Unknown |
| TAUS Translation Memory | Unknown |
| Scale HLE | Unknown |
| HackerRank Coding | Unknown |
| RL data for Search | Gemini 3; GPT-5 |
| Dataset | Model(s) used |
|---|---|
| Simple Minesweeper | Undisclosed |
| Simple Sudoku | Undisclosed |
| Multitool Typewriter Hard | Undisclosed |
| Machine Translation of News Commentary and TAUS Translation Memory | Undisclosed |
| Machine Translation of STEM - | Qwen2.5-14B-Instruct |
| Competitive Coding RL data from Nemotron Cascade | Undisclosed |
| Long context RL | Undisclosed |
| Single-step SWE RL for patch generation | Undisclosed |
| OpenHands SWE | Undisclosed |
| Dataset | Modality | Dataset Size | Seed Dataset | Model(s) used for generation |
|---|---|---|---|---|
| Nemotron-Pretraining-Fact-Seeking | Text | 35.0B | FineWiki | Qwen3-30B-A3B-Instruct-2507 |
| Nemotron-Pretraining-Legal | Text | 4.3B | CommonPile (caselaw_access_project_filtered); California Code of Regulations; Judicial Ethics Opinions; GLOBALCIT; CUAD; Nemotron Personas; ToSDR Terms of Service Corpus; CodeHima/TOS_Dataset; ContractNLI; CaseHOLD; Code of Federal Regulations; Canadian Case Law (subsets that allow commercial use) | Qwen3-235B-A22B-Thinking-2507 |
| Nemotron-Pretraining-Formal-Logic | Text | 128M | Nemotron Personas | Qwen3-235B-A22B-Thinking-2507 |
| Nemotron-Pretraining-Economics | Text | 73.4M | - | Qwen3-235B-A22B-Thinking-2507 |
| Nemotron-Pretraining-Multiple-Choice | Text | 1.6B | MMLU Auxiliary Train | DeepSeek-V3; Qwen3-235B-A22B |
| Nemotron-Pretraining-Code-Concepts | Text | 7.3B | - | gpt-oss-20b; gpt-oss-120b |
| Nemotron-Pretraining-Unconditional-Algorithmic | Text | 196.5M | - | gpt-oss-120b; Qwen3-235B-A22B |
| More Synthetic Tasks from DeepSeek-V3 and Qwen3-235B-A22B | Text | 1.1B | train splits of acp_bench; ai2_arc; babi; gsm8k; hendrycks_math; IFEval; MedText; mediqa_qa; mlqa; MMLU-Pro; mmlu-pro-plus; MMLU-ProX; nq_open; tinyGSM8k; truthful_qa; truthfulqa-multi; MATH-lighteval; mmlu; awesome-chatgpt-prompts; super_glue | DeepSeek v3; Qwen3-235B-A22B |
| Synthetic Tasks from DeepSeek-V3 and Qwen3-235B-A22B | Text | 6.7B | train splits of Into the Unknown; AI2 ARC (AI2 Reasoning Challenge); BLiMP (Benchmark of Linguistic Minimal Pairs); CommonSenseQA; GLUE; HeadQA; Hendrycks Ethics; Memo Trap; modus-tollens; NeQA; pattern-matching-suppression; mastermind_24_mcq_random; mastermind_24_mcq_close; quote-repetition; redefine-math; Repetitive Algebra; sig-figs; MMLU-Pro; MC-TACO; MedConceptsQA; MMLU_dataset; OpenbooksQA; PIQA (Physical Interaction Question Answering); SocialIQA; SuperGLUE; tinyAI2_arc; tinyMMLU; tinyWinogrande; TruthfulQA; WebQuestions; Winogrande; GPQA; MBPP | DeepSeek v3; Qwen3-235B-A22B |
| Synthetic Art of Problem Solving from DeepSeek-R1 | Text | 40B | Art of Problem Solving; American Mathematics Competitions 8; American Mathematics Competitions 10; | DeepSeek-R1 |
| Synthetic Moral Stories and Social Chemistry from Qwen3-235B-A22B-Thinking-2507 and Mixtral-8x22B-v0.1 | Text | 15.2M | social-chemestry-101; Moral Stories | Qwen3-235B-A22B-Thinking-2507; Mixtral-8x22B-v0.1 |
| Synthetic Moral Stories and Social Chemistry from Mixtral-8x22B-v0.1 | Text | 327M | social-chemestry-101; Moral Stories | Mixtral-8x22B-v0.1 |
| Synthetic Social Sciences seeded with OpenStax from DeepSeek-V3, Mixtral-8x22B-v0.1, and Qwen2.5-72B | Text | 83.6M | OpenStax - CC BY-SA subset | DeepSeek-V3; Mixtral-8x22B-v0.1; Qwen2.5-72B |
| Synthetic Health Sciences seeded with OpenStax from DeepSeek-V3, Mixtral-8x22B-v0.1, and Qwen2.5-72B | Text | 9.7M | OpenStax - CC BY-SA subset | DeepSeek-V3; Mixtral-8x22B-v0.1; Qwen2.5-72B |
| Synthetic STEM seeded with OpenStax, Open Textbook Library, and GSM8K from DeepSeek-R1, DeepSeek-V3, DeepSeek-V3-0324, and Qwen2.5-72B | Text | 175M | OpenStax - CC BY-SA subset; GSM8K; Open Textbook Library - CC BY-SA & GNU subset | DeepSeek-R1, DeepSeek-V3; DeepSeek-V3-0324; Qwen2.5-72B |
| Nemotron-PrismMath | Text | 4.6B | Big-Math-RL-Verified; OpenR1-Math-220k | Qwen2.5-0.5B-instruct, Qwen2.5-72B-Instruct; DeepSeek-R1-Distill-Qwen-32B |
| Synthetic Question Answering Data from Papers and Permissible Books from Qwen2.5-72B-Instruct | Text | 350M | arXiv; National Institutes of Health ExPorter; BioRxiv; PMC Article; USPTO Backgrounds; peS2o; Global Regulation; CORE; PG-19; DOAB CC BY & CC BY-SA subset; NDLTD | Qwen2.5-72B-Instruct |
| Synthetic Rephrased Math Data from Common Crawl from phi-4 | Text | 73B | Common Crawl | phi-4 |
| Synthetic Math Data from Common Crawl 4plus | Text | 52.3B | Common Crawl | phi-4 |
| Synthetic Math Data from Common Crawl 3 | Text | 80.9B | Common Crawl | phi-4 |
| Synthetic AGIEval seeded with AQUA-RAT, LogiQA, and AR-LSAT from DeepSeek-V3 and DeepSeek-V3-0324 | Text | 4.0B | AQUA-RAT; LogiQA; AR-LSAT | DeepSeek-V3; DeepSeek-V3-0324 |
| Synthetic AGIEval seeded with AQUA-RAT, LogiQA, and AR-LSAT from Qwen3-30B-A3B | Text | 4.2B | AQUA-RAT; LogiQA; AR-LSAT | Qwen3-30B-A3B |
| Synthetic Art of Problem Solving from Qwen2.5-32B-Instruct, Qwen2.5-Math-72B, Qwen2.5-Math-7B, and Qwen2.5-72B-Instruct | Text | Undisclosed | Art of Problem Solving; American Mathematics Competitions 8; American Mathematics Competitions 10; GSM8K; PRM800K | Qwen2.5-32B-Instruct; Qwen2.5-Math-72B; Qwen2.5-Math-7B; Qwen2.5-72B-Instruct |
| Synthetic MMLU Auxiliary Train from DeepSeek-R1 | Text | 0.5B | MMLU Auxiliary Train | DeepSeek-R1 |
| Synthetic Long Context Continued Post-Training Data from Papers and Permissible Books from Qwen2.5-72B-Instruct | Text | Undisclosed | arXiv; National Institutes of Health ExPorter; BioRxiv; PMC Article; USPTO Backgrounds; peS2o; Global Regulation; CORE; PG-19; DOAB CC BY & CC BY-SA subset; NDLTD | Qwen2.5-72B-Instruct |
| Synthetic Common Crawl from Qwen3-30B-A3B and Mistral-Nemo-12B-Instruct | Text | 415.8B | Common Crawl | Qwen3-30B-A3B; Mistral-NeMo-12B-Instruct |
| Synthetic Multilingual Data from Common Crawl from Qwen3-30B-A3B | Text | Undisclosed | Common Crawl | Qwen3-30B-A3B |
| Synthetic Multilingual Data from Wikimedia from Qwen3-30B-A3B | Text | Undisclosed | Wikimedia | Qwen3-30B-A3B |
| Synthetic Math Data from Wikimedia from Nemotron-4-340B-Instruct | Text | Undisclosed | - | Nemotron-4-340B-Instruct |
| Synthetic Common Crawl Code from phi-4 | Text | 427.9B | Common Crawl | phi-4 |
| Synthetic Scientific Coding from Qwen3-235B-A22B | Text | 1.2B | Wikimedia | Qwen3-235B-A22B |
| Tool Calling Data | Text | 26.2B | Qwen3-235B-A22B-2507; gpt-oss-120b | |
| Synthetic Essential-Web from QwQ-32B | Text | 28.1B | Essential-Web | QwQ-32B |
| Translated Synthetic Crawl | Text | 389.9B | Common Crawl | Qwen3-30B-A3B |
| Translated Synthetic Wikipedia | Text | 7.9B | Wikimedia | Qwen3-30B-A3B |
| Synthetic Long Context from Qwen3-235B-A22B-Instruct-2507 | Text | Undisclosed | CORE; PG-19; DOAB CC BY & CC BY-SA subset; NDLTD | Qwen3-235B-A22B-Instruct-2507 |
| Synthetic Search STEM OPENQ from DeepSeek-R1-0528 | Text | Undisclosed | - | DeepSeek-R1-0528 |
| Synthetic MCQ from Qwen2.5-32B-Instruct and DeepSeek-R1-0528 | Text | Undisclosed | - | Qwen2.5-32B-Instruct; DeepSeek-R1-0528 |
| Synthetic Offline Search MCQA HLE from DeepSeek-R1-0528 | Text | Undisclosed | - | DeepSeek-R1-0528 |
| Synthetic Offline Search MCQA GPQA from Qwen3-235B-A22B and DeepSeek-R1-0528 | Text | Undisclosed | - | Qwen3-235B-A22B; DeepSeek-R1-0528 |
| Synthetic Human Preference from QwQ-32B, Qwen3-30B-A3B, Qwen3-235B-A22B, Qwen3-235B-A22B-Instruct-2507, Mistral-Small-3.1-24B-Instruct-2503, Mistral-Small-3.2-24B-Instruct-2506, MiniMax-M1-80k, MiniMax-M1-40k, Kimi-K2-Instruct, DeepSeek-V3-0324, DeepSeek-R1-0528 | Text | Undisclosed | - | QwQ-32B; Qwen3-30B-A3B; Qwen3-235B-A22B; Qwen3-235B-A22B-Instruct-2507; Mistral-Small-3.1-24B-Instruct-2503; Mistral-Small-3.2-24B-Instruct-2506; MiniMax-M1-80k; MiniMax-M1-40k; Kimi-K2-Instruct; DeepSeek-V3-0324; DeepSeek-R1-0528 |
| Synthetic Code from Qwen3-32B | Text | Undisclosed | English Common Crawl; English Common Crawl 1.1 | Qwen3-32B |
| Synthetic OpenCodeReasoning from DeepSeek-R1 | Text | Undisclosed | OpenCodeReasoning | DeepSeek-R1 |
| Synthetic LIMO from DeepSeek-R1-0528 | Text | Undisclosed | LIMO | DeepSeek-R1-0528 |
| Synthetic SCP from DeepSeek-R1-0528 | Text | Undisclosed | SCP-116K | DeepSeek-R1-0528 |
| Synthetic Stack Exchange from DeepSeek-R1-0528 | Text | Undisclosed | Stack Exchange | DeepSeek-R1-0528 |
| Synthetic Common Crawl from Qwen3-30B-A3B | Text | Undisclosed | Common Crawl | Qwen3-30B-A3B |
| Synthetic Wikipedia from Qwen3-30B-A3B | Text | Undisclosed | Wikimedia | Qwen3-30B-A3B |
| Synthetic Essential-Web from Qwen3-30B-A3B and Qwen3-235B-A22B-Thinking-2507 | Text | Undisclosed | Essential-Web | Qwen3-30B-A3B; Qwen3-235B-A22B-Thinking-2507 |
| Synthetic Textbook Math from Qwen3-30B-A3B, Qwen3-235B-A22B, phi-4 | Text | Undisclosed | Common Crawl; FineMath | Qwen3-30B-A3B; Qwen3-235B-A22B; phi-4 |
| Synthetic Math and Code from DeepSeek-R1 and DeepSeek-R1-0528 | Text | Undisclosed | Magicoder-Evol-Instruct-110K; opc-sft-stage2; TACO; OpenCodeReasoning; OpenMathReasoning; NuminaMath CoT | DeepSeek-R1; DeepSeek-R1-0528 |
| Dataset | Modality | Dataset Size | Seed Dataset | Model(s) used for generation |
|---|---|---|---|---|
| Synthetic Competitive MATH Proofs from DeepSeek-V4-Pro | Text | Undisclosed | [AMC 8 Problems and Solutions, AMC 10 Problems and Solution, and AIME Problems and Solutions] | [deepseek-ai/DeepSeek-V4-Pro] |
| Synthetic Hermes Agent Reasoning Traces | Text | Undisclosed | [lambda/hermes-agent-reasoning-traces] | [hermes-agent-generator] |
| Synthetic Competitive Coding from DeepSeek-V4-Pro | Text | Undisclosed | [NVCompetitiveCodingV1] | [deepseek-ai/DeepSeek-V4-Pro] |
| Synthetic Competitive Science Reasoning from DeepSeek-V4-Pro | Text | Undisclosed | [AMC 8 Problems and Solutions, AMC 10 Problems and Solution, and AIME Problems and Solutions]; [EssentialAI/essential-web-v1.0]; [cdquestions.com]; [Pile-FreeLaw]; [Vedantu]; [askfilo]; [doubtnut]; [ICHO-IPH0 Dataset]; [LIMO dataset (Less is More for Reasoning)]; [AAPT]; [ChemData 700K]; [oMeBench]; [Flavor Analysis and Recognition Transformer]; [ChemCoTBench]; [Llama Nemotron Dataset] | [deepseek-ai/DeepSeek-V4-Pro] |
| Synthetic Competitive MATH CoT and TIR from Nemotron 5.5 | Text | Undisclosed | [Pile-FreeLaw]; [AMC 8 Problems and Solutions, AMC 10 Problems and Solution, and AIME Problems and Solutions] | [Nemotron 5.5] |
| Vendor Terminal Bench-like Tasks from Mercor | Text | Undisclosed | [Terminal bench like tasks curated by the vendor] | [Undisclosed - purchased dataset] |
| Turing Math Data Pack | Text | Undisclosed | [Turing Math Data Pack dataset] | [Undisclosed - purchased dataset] |
| Synthetic Holdout, Skywork, DAPO, and Turing Math from GPT-5.5 | Text | Undisclosed | [DocQA-RL-1.6K]; [DAPO-Math-17k] | [GPT-5.5] |
| Synthetic Long Context RL from QwenLong L1 and DocQA-RL-1.6K | Text | Undisclosed | [DocQA-RL-1.6K] | Undisclosed |
| Synthetic Competitive Coding Gym Tasks | Text | Undisclosed | [NVCompetitiveCodingV1.1] | Undisclosed |
| Synthetic Finance SEC Search Agent from GPT-OSS-120B and Qwen3 | Text | Undisclosed | [SEC filings from sec.gov] | [GPT-OSS-120B]; [Qwen3-235B-A22B-Instruct]; [Qwen3-4B-Instruct] |
| Synthetic Structured Outputs from Qwen3-30B-A3B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507, Qwen3-235B-A22B-Instruct-2507, and Qwen3-235B-A22B-Thinking-2507 | Text | Undisclosed | [Nemotron-RL-agent-structured-outputs-v1] | [Qwen3-30B-A3B-Instruct-2507]; [Qwen3-235B-A22B-Instruct-2507] |
| Synthetic Long Context Equivalence Rule from Qwen3-235B-A22B-Thinking-2507 and DeepSeek-R1 | Text | Undisclosed | [Long-context SFT data] | [Qwen/Qwen3-235B-A22B-Thinking-2507]; [Deepseek-ai/DeepSeek-R1] |
| Synthetic Science RL Data Blend from Qwen2.5-32B | Text | Undisclosed | [doubtnut]; [Pile-FreeLaw]; [Llama Nemotron Dataset]; [askfilo]; [EssentialAI/essential-web-v1.0]; [Vedantu]; [auxiliary_train]; [cdquestions.com]; [AMC 8 Problems and Solutions, AMC 10 Problems and Solution, and AIME Problems and Solutions]; [AAPT]; [ICHO-IPH0 Dataset]; [LIMO dataset (Less is More for Reasoning)] | [Qwen2.5-32B] |
| Synthetic Abstention Data from Nemotron Super v3 | Text | Undisclosed | [Go abstention Dataset] | [nvidia/nvidia/nemotron-3-super-v3] |
| Synthetic Chemistry Data from Nemotron Super v3 | Text | Undisclosed | [ChemData 700K] | [nvidia/nvidia/nemotron-3-super-v3] |
| Synthetic Structured Outputs from Qwen3-30B-A3B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507, Qwen3-235B-A22B-Instruct-2507, and Qwen3-235B-A22B-Thinking-2507 | Text | Undisclosed | [In-house data] | [GPT OSS 120B - Apache 2.0] |
| Synthetic Tool Call Schema for RL | Text | Undisclosed | [In-house data] | [GPT OSS 120B - Apache 2.0] |
| Synthetic Freeform Text Formatting from GPT-OSS-120B | Text | Undisclosed | [In-house data] | [GPT OSS 120B - Apache 2.0] |
| Synthetic Citation Formatting from GPT-OSS-120B | Text | Undisclosed | [In-house data] | [GPT OSS 120B - Apache 2.0] |
| Droid Harness Pivot Vendor Data | Text | Undisclosed | [Droid Harness Pivot vendor data] | Undisclosed |
| Synthetic HotpotQA Training Data from Qwen3-235B | Text | Undisclosed | [HotpotQA] | [Qwen3-235B] |
| Synthetic Natural Language Math Proofs from Nemotron 5.5 | Text | Undisclosed | [AMC8, AMC10, and AIME problem sets hosted on Art of Problem Solving]; [Pile-StackExchange] | [Nemotron 5.5] |
| Synthetic Stack Overflow OpenQ | Text | Undisclosed | [Pile-FreeLaw] | Undisclosed |
| Chemistry Ether0 Vendor Data | Text | Undisclosed | [Chemistry ether0 vendor data] | Undisclosed |
| Synthetic Litmus-Bench Chemistry from ChEMBL | Text | Undisclosed | [ChEMBL]; [Nemo Gym RL dataset generated from ChEMBL with RDKit] | Undisclosed |
| Synthetic ZINC Chemistry from Nemotron Super v3 | Text | Undisclosed | [ZINC] | [Nemotron Super v3] |
| ARC-AGI Gym Environment | Text | Undisclosed | [ARC-AGI-2] | [ARC-AGI-2] |
| Synthetic Agentic Search Tool-Use from DeepSeek-V3.2 | Text | Undisclosed | [Mercor Data] | [DeepSeek-V3.2] |
| Synthetic Text-To-SQL | Text | Undisclosed | [In-house Text-to-SQL data] | [gpt-oss-120b] |
| Dialog Memory Vendor Data | Text | Undisclosed | [Patronus external vendor agreement] | Undisclosed |
| Synthetic Indirect Prompt Injection from Nemotron Super v3 and Qwen3-Next-80B-A3B-Instruct | Text | Undisclosed | [In-house indirect prompt injection data] | [nvidia/nemotron-3-super-v3, qwen/qwen3-next-80b-a3b-instruct.] |
| Synthetic Malicious Code and Agentic Security | Text | Undisclosed | [In-house malicious-code / agentic-security data] | Undisclosed |
| Synthetic Single-Step SWE Patch Selection | Text | Undisclosed | [SWE-Gym Dataset]; [SWE Bench Verified Benchmark] | [ground truth and task checks] |
| Synthetic Natural Language Math Final Answers from Nemotron 5.5 | Text | Undisclosed | [AMC8, AMC10, and AIME problem sets hosted on Art of Problem Solving]; [Pile-StackExchange] | [nemotron 5.5] |
| Synthetic Simple Math Prompts for Token Efficiency | Text | Undisclosed | [In-house simple math prompts] | Undisclosed |
| Synthetic Abstention Data from Nemotron Super v3 | Text | Undisclosed | [CRAG] | [nvidia/nvidia/nemotron-3-super-v3] |
| Synthetic Agentless SWE | Text | 242,536 | [SWE-Rebench-V2]; [SWEbench Training Set]; [R2E-Gym/R2E-Gym-Subset]; [SWE-Gym/SWE-Gym]; [SWE-Rebench] | [openai/gpt-oss-120b] |
| Synthetic Agentic CUDA Traces from GLM-4.7 | Text | 2,276 | [Internal CUDA task data] | [GLM-4.7] |
| Synthetic Math Proofs from DeepSeek-V3.2-Speciale | Text | 820,772 | [Nemotron-Math-Proofs-v1] | [SDG: DeepSeek-V3.2-Speciale]; [Filter: proof validation] |
| Synthetic Multilingual SFT from DeepSeek-V3 | Text | 1,245,284 | [Nano v3 SFT data] | [DeepSeek-V3] |
| Synthetic Agentic Code from gpt-oss-120b | Text | 109,086 | [NVAgenticCLIPrompts-v1]; [NVAgenticSkills-v1]; [NVAgenticCLIMultiTurnPrompts-v1] | [openai/gpt-oss-120b] |
| Synthetic Agentic CLI and Web Skills from gpt-oss-120b | Text | 27,418 | [NVAgenticCLIPrompts-v1]; [NVAgenticSkills-v1]; [NVAgenticCLIMultiTurnPrompts-v1]; [NVAgenticCLIPrompts-Web-v1] | [openai/gpt-oss-120b] |
| Synthetic Agentic Coding from gpt-oss-120b | Text | 160,531 | [NVAgenticCLIPrompts-v1]; [NVAgenticSkills-v1]; [NVAgenticCLIMultiTurnPrompts-v1]; [NVAgenticCLIPrompts-Web-v1] | [openai/gpt-oss-120b] |
| Synthetic OpenCode Agentic Tasks from gpt-oss-120b | Text | 614,773 | [NVAgenticCLIPrompts-v1]; [NVAgenticSkills-v1]; [NVAgenticCLIMultiTurnPrompts-v1]; [NVAgenticCLIPrompts-Web-v1] | [openai/gpt-oss-120b] |
| Synthetic ARC-AGI Ultra Data | Text | 192,016 | [ARC-AGI-2]; [arc dataset collection] | [ARC-AGI-2] |
| Synthetic LiveCodeBench TIR from DeepSeek-R1-0528 | Text | 1,283,398 | [Nemotron-X training datasets] | [DeepSeek-R1-0528] |
| Synthetic Verilog and SystemVerilog Code from DeepSeek-R1-0528 and GPT-OSS-120B | Text | 1,233,247 | [Verilog/SystemVerilog seed code] | [SDR: DeepSeek R1 0528 and GPT-OSS-120B]; [Filtering: Claude 4 Sonnet] |
| Synthetic Aider Python Tasks from DeepSeek-R1-0528 | Text | 236,099 | [Exercism (GitHub Python)] | [Deepseek R1 0528] |
| Synthetic Chat Reasoning-Off Data from GLM-5 | Text | 646,738 | [lmarena-ai/repochat-arena-preference-4k user prompts] | [Multi-turn conversations generated by GLM-5 with best-of-4 selection via Qwen3-Nemotron-235B-A22B-GenRM:] |
| Synthetic Chat Reasoning-On Data from GLM-5 | Text | 644,286 | [lmarena-ai/repochat-arena-preference-4k user prompts]; [lmarena-ai/arena-expert-5k user prompts]; [lmarena-ai/arena-human-preference-55k user prompts]; [lmarena-ai/arena-human-preference-100k user prompts]; [lmarena-ai/arena-human-preference-140k user prompts] | [Multi-turn conversations generated by GLM-5 with best-of-4 selection via Qwen3-Nemotron-235B-A22B-GenRM:] |
| Synthetic Multilingual Safety from Riva-Translate-4B-Instruct-v1.1 | Text | 132,067 | [Safety SFT Data: Ultra] | [nvidia/Riva-Translate-4B-Instruct-v1.1] |
| Synthetic Science Reasoning Effort Medium | Text | 502,722 | [science-reasoning-effort-medium-v0] | Undisclosed |
| Synthetic Telecom Tool-Use Trajectories from gpt-oss-120b | Text | 12,455 | [Existing Tau2 telecom trajectories originally generated with DeepSeek V3.2] | [gpt-oss-120b] |
| Synthetic Terminal Bench Data from OpenReasoningv2 | Text | Undisclosed | [OpenCodeReasoningv2]; [OpenMathReasoning]; [nemo-swe-bench-repos]; [SWE-Rebench]; [SWE-Fixer-110K] | [OpenReasoningv2] |
| Synthetic Tulu Instruction Following from DeepSeek-R1-0528 | Text | 105,361 | [Nemotron-X training datasets] | [DeepSeek-R1-0528] |
| Synthetic SWE Unverified | Text | Undisclosed | [NVAgenticCLIPrompts-v1]; [NVAgenticSkills-v1]; [NVAgenticCLIMultiTurnPrompts-v1]; [NVAgenticCLIPrompts-Web-v1] | [gpt-oss-120b] |
| Synthetic Instruction Following from gpt-oss-120b | Text | 151,988 | [IFEval]; [IFEvalG] | [gpt-oss-120b] |
| Synthetic Identity Data from Qwen3-Next-80B-A3B-Instruct and Qwen3-235B-A22B-Instruct-2507 | Text | 25,992 | [Hand-written prompts] | [Qwen3-Next-80B-A3B-Instruct]; [Qwen3-235B-A22B-Instruct-2507] |
| Synthetic Terminus Ultra Agentic Reasoning Blend | Text | 96,881 | [ARC-AGI-2]; [OpenCodeReasoningv2]; [OpenMathReasoning]; [SWE-Fixer-110K]; [SWE-Rebench]; [SWE-Smith] | [DeepSeek-V3.2]; [Qwen3-235B-A22B-Thinking-2507]; [Ring-1T]; [Kimi-K2.5]; [GLM-4.7-FP8]; [Qwen3-Next-80B-A3B-Thinking]; [gpt-oss-120b]; [Ministral-3-14B-Reasoning-2512]; [LM-4.5-Air-FP8] |
| Synthetic STEM from Qwen3-235B-A22B-Thinking-2507 | Text | 1,174,694 | [IChO-IPhO-RL-v2]; [Physics-Big Dataset] | Undisclosed |
| Translation Data from TAUS | Text | 1,618,055 | [TAUS proprietary dataset] | Undisclosed |
| Synthetic Art of Problem Solving and Stack Exchange from gpt-oss-120b, Qwen2.5-32B-Instruct, and Goedel-Prover-V2-32B | Text | 860,469 | [Nemotron-Math-Proofs-v1] | [Goedel-Prover-V2-32B] |
| Synthetic Art of Problem Solving and Stack Exchange from gpt-oss-120b, Qwen2.5-32B-Instruct, and Goedel-Prover-V2-32B | Text | 1,201,815 | [Upstream released math dataset]; [AoPS]; [StackOverflow / StackExchange] | [gpt-oss-120b] |
| Synthetic Art of Problem Solving and Stack Exchange from gpt-oss-120b, Qwen2.5-32B-Instruct, and Goedel-Prover-V2-32B | Text | 1,296,676 | [Upstream released math dataset]; [AoPS]; [StackOverflow / StackExchange] | [gpt-oss-120b] |
| Synthetic Instruction Following for RL | Text | Undisclosed | [WildChat-1M]; [LMSYS-340B-Eval Dataset]; [LMSYS-Chat-1M Prompts]; [IFEval]; [IFEvalG] | [Qwen/Qwen3-235B-A22B-Thinking-2507]; [gpt-oss-120b]; [Qwen3-235B-A22B-Instruct-2507] |
| Synthetic Instruction Following for RL | Text | Undisclosed | [WildChat-1M]; [LMSYS-340B-Eval Dataset]; [LMSYS-Chat-1M Prompts]; [IFEval]; [IFEvalG] | [Qwen/Qwen3-235B-A22B-Thinking-2507]; [gpt-oss-120b]; [Qwen3-235B-A22B-Instruct-2507] |
| Synthetic Multilingual Science and Code data from DeepSeek-R1, DeepSeek-R1-0528, Qwen2.5-32B-Instruct, and Qwen3-235B-A22B, translated with Qwen2.5-32B-Instruct and Qwen2.5-14B-Instruct | Text | Undisclosed | [Nano-V3 SFT Data (without tool call)] | [Qwen/Qwen2.5-14B-Instruct]; [Qwen/Qwen3-4B-Thinking-2507] |
| Synthetic Search Graph Walk | Text | 6,977 | [Wikidata / Wikipedia KnowledgeBase] | [MiniMaxAI/MiniMax-M2] |
| Synthetic Agentic Diverse Domains | Text | 281,537 | [Handwritten prompts (synthetic; no external seed data used)] | [SDG model: deepseek-ai/DeepSeek-V3.2, deepseek-ai/DeepSeek-R1-0528, Qwen/Qwen3-235B-A22B-Thinking-2507, Qwen/Qwen3-32B]; [Filtering model: openai/gpt-oss-120b, Qwen/Qwen3-32B, Qwen/Qwen3-235B-A22B-Instruct-2507] |
| Synthetic Long Context from Qwen3-235B-A22B-Instruct-2507 | Text | 65,608 | [Long-context SFT seed blend (pre-training blend + nano-v1 post-training data)] | [Qwen/Qwen3-235B-A22B-Thinking-2507 and deepseek-ai/DeepSeek-R1] |
| Synthetic Agentless SWE | Text | 209,976 | [SWE-Bench-Train]; [SWE-Fixer-Train]; [SWE-reBench]; [SWE-Smith] | [deepseek-ai/DeepSeek-R1-0528] |
| Synthetic Nemotron Math SFT from DeepSeek-V3.2-Speciale | Text | 1,900,553 | [Nemotron-Math-v2 (AOPS and StackExchange-math problems)] | [DeepSeek-V3.2-Speciale] |
| Synthetic Nemotron Math TIR from DeepSeek-V3.2 | Text | 1,789,258 | [Nemotron-Math-v2 (AOPS and StackExchange-math problems)] | [DeepSeek-V3.2] |
| Synthetic SWE Unverified | Text | 27,911 | [NVAgenticCLIPrompts-v1]; [NVAgenticSkills-v1] | [gpt-oss-120b] |
| Synthetic SWE Unverified | Text | 28,116 | [NVAgenticCLIPrompts-v1]; [NVAgenticSkills-v1] | [Qwen3-Coder-480B-A35B-Instruct] |
| Synthetic NemoCascade OCR Distillation from gpt-oss-120b | Text | 682,864 | [Nemotron-X training datasets] | [gpt-oss-120b] |
| Synthetic SWE Unverified | Text | 26,865 | [NVAgenticCLIPrompts-v1]; [NVAgenticSkills-v1] | [gpt-oss-120b]; [Qwen/Qwen3-Coder-480B-A35B-Instruct]; [GLM-4.7-Flash] |
| Synthetic CUDA 100k | Text | 93,086 | [KernelBook]; [HuggingFace Transformers]; [FlashInfer] | [gpt-oss-120b]; [DeepSeek-R1-0528] |
| Synthetic Science MCQ and QA Diversity from GPT-OSS and Kimi-K2 | Text | 30,358 | [doubtnut]; [Pile-FreeLaw]; [Llama Nemotron Dataset]; [askfilo]; [EssentialAI/essential-web-v1.0]; [Vedantu]; [auxiliary_train]; [cdquestions.com]; [AMC 8 Problems and Solutions, AMC 10 Problems and Solution, and AIME Problems and Solutions]; [AAPT]; [ICHO-IPH0 Dataset]; [LIMO dataset (Less is More for Reasoning)] | [GPT-OSS]; [Kimi-K2] |
| Synthetic Science HLE with Python from GPT-OSS and Kimi-K2 | Text | 85,184 | [doubtnut]; [Pile-FreeLaw]; [Llama Nemotron Dataset]; [askfilo]; [EssentialAI/essential-web-v1.0]; [Vedantu]; [auxiliary_train]; [cdquestions.com]; [AMC 8 Problems and Solutions, AMC 10 Problems and Solution, and AIME Problems and Solutions]; [AAPT]; [ICHO-IPH0 Dataset]; [LIMO dataset (Less is More for Reasoning)] | [GPT-OSS]; [Kimi-K2] |
| Synthetic Science Search and Python from GPT-OSS and Kimi-K2 | Text | 6,179 | [doubtnut]; [Pile-FreeLaw]; [Llama Nemotron Dataset]; [askfilo]; [EssentialAI/essential-web-v1.0]; [Vedantu]; [auxiliary_train]; [cdquestions.com]; [AMC 8 Problems and Solutions, AMC 10 Problems and Solution, and AIME Problems and Solutions]; [AAPT]; [ICHO-IPH0 Dataset]; [LIMO dataset (Less is More for Reasoning)] | [GPT-OSS]; [Kimi-K2] |
| Synthetic Science Search from GPT-OSS and Kimi-K2 | Text | 32,554 | [doubtnut]; [Pile-FreeLaw]; [Llama Nemotron Dataset]; [askfilo]; [EssentialAI/essential-web-v1.0]; [Vedantu]; [auxiliary_train]; [cdquestions.com]; [AMC 8 Problems and Solutions, AMC 10 Problems and Solution, and AIME Problems and Solutions]; [AAPT]; [ICHO-IPH0 Dataset]; [LIMO dataset (Less is More for Reasoning)] | [GPT-OSS]; [Kimi-K2] |
| Synthetic Finance Reasoning from GPT-OSS-120B and Qwen3-235B-A22B-Instruct-2507 | Text | 326,700 | [_SEC filings] | [GPT-OSS-120B, Qwen3-235B-A22B-Instruct-2507] |
| Synthetic Science Diversity MCQ from GPT-OSS and Kimi-K2 | Text | 532,942 | [doubtnut]; [Pile-FreeLaw]; [Llama Nemotron Dataset]; [askfilo]; [EssentialAI/essential-web-v1.0]; [Vedantu]; [auxiliary_train]; [cdquestions.com]; [AMC 8 Problems and Solutions, AMC 10 Problems and Solution, and AIME Problems and Solutions]; [AAPT]; [ICHO-IPH0 Dataset]; [LIMO dataset (Less is More for Reasoning)] | [GPT-OSS]; [Kimi-K2] |
| Synthetic Science Diversity OpenQ from GPT-OSS and Kimi-K2 | Text | 131,045 | [doubtnut]; [Pile-FreeLaw]; [Llama Nemotron Dataset]; [askfilo]; [EssentialAI/essential-web-v1.0]; [Vedantu]; [auxiliary_train]; [cdquestions.com]; [AMC 8 Problems and Solutions, AMC 10 Problems and Solution, and AIME Problems and Solutions]; [AAPT]; [ICHO-IPH0 Dataset]; [LIMO dataset (Less is More for Reasoning)] | [GPT-OSS]; [Kimi-K2] |
| Synthetic Science Reasoning No-Tool from GPT-OSS and Kimi-K2 | Text | 2,085,600 | [doubtnut]; [Pile-FreeLaw]; [Llama Nemotron Dataset]; [askfilo]; [EssentialAI/essential-web-v1.0]; [Vedantu]; [auxiliary_train]; [cdquestions.com]; [AMC 8 Problems and Solutions, AMC 10 Problems and Solution, and AIME Problems and Solutions]; [AAPT]; [ICHO-IPH0 Dataset]; [LIMO dataset (Less is More for Reasoning)] | [GPT-OSS]; [Kimi-K2] |
| Synthetic Long Context from Qwen3-235B-A22B-Instruct-2507 | Text | 62,333 | [Long-context SFT data: lc_nothink 256k] | [Qwen/Qwen3-235B-A22B-Thinking-2507 and deepseek-ai/DeepSeek-R1] |
| Synthetic Long Context from Qwen3-235B-A22B-Instruct-2507 | Text | 49,698 | [Long-context SFT data: MRCR 200k] | [Qwen/Qwen3-235B-A22B-Thinking-2507 and deepseek-ai/DeepSeek-R1] |
| Synthetic Text-To-SQL | Text | 96,564 | [Undisclosed - no seed data listed] | [gpt-oss-120b] |
| Synthetic Long Context from Qwen3-235B-A22B-Instruct-2507 | Text | 397,538 | [Long-context SFT data: RULER 256k] | [Qwen/Qwen3-235B-A22B-Thinking-2507 and deepseek-ai/DeepSeek-R1] |
| Synthetic SWE Unverified | Text | 27,960 | [NVAgenticCLIPrompts-v1]; [NVAgenticSkills-v1] | [gpt-oss-120b]; [Qwen/Qwen3-Coder-480B-A35B-Instruct]; [GLM-4.7-Flash] |
| Synthetic SWE Unverified | Text | 24,632 | [NVAgenticCLIPrompts-v1]; [NVAgenticSkills-v1] | [gpt-oss-120b]; [Qwen/Qwen3-Coder-480B-A35B-Instruct]; [GLM-4.7-Flash] |
| Synthetic Long Context from Qwen3-235B-A22B-Instruct-2507 | Text | 49,902 | [Long-context SFT data] | [Qwen/Qwen3-235B-A22B-Thinking-2507 and deepseek-ai/DeepSeek-R1] |
| Synthetic Tool Call Schema for RL | Text | 469,983 | [UltraTool]; [ToolEyes]; [AutoTools]; [API-Bank]; [Nemotron-Personas-USA]; [Salesforce xLAM function-calling]; [Glaive function-calling-v2]; [Agent-Ark/Toucan-1.5M] | [DeepSeek-V3.2]; [GLM-4.6]; [gpt-oss-120b]; [Kimi-K2-Instruct] |
| Synthetic Tool Call Schema for RL | Text | 707,967 | [UltraTool]; [ToolEyes]; [AutoTools]; [API-Bank]; [Nemotron-Personas-USA]; [Salesforce xLAM function-calling]; [Glaive function-calling-v2]; [Agent-Ark/Toucan-1.5M] | [DeepSeek-V3.2]; [GLM-4.6]; [gpt-oss-120b]; [Kimi-K2-Instruct] |
| Synthetic Long Context from Qwen3-235B-A22B-Instruct-2507 | Text | 52,630 | [AALCR seed blend: SEC Filings]; [CC]; [Wikipedia]; [FinePDFs]; [ArXiv]; [Pile-NIH ExPorter]; [BioRxiv]; [PMC Article]; [USPTO Backgrounds]; [peS20]; [Global Regulations]; [CORE]; [Gutenberg (PG-19)]; [DOAB CC-BY]; [NDLTD]; [Amps]; [StackExchange]; [MathPile]; [Numinas] | [Qwen3-30B-A3B] |
| Synthetic Safety from gemma-3-4b-it, Nemotron-Nano-9B-v2, and gpt-oss-120b | Text | 44,091 | [Safety SFT Data] | [google/gemma-3-4b-it]; [Nemotron-Nano-9B-v2]; [gpt-oss-120b] |
| Language | Size |
|---|---|
| English | 8.6M |
| Italian | 138k |
| German | 138k |
| Spanish | 138k |
| French | 138k |
| Japanese | 138k |
| Chinese | 138k |
| Hindi | 138k |
| Korean | 138k |
| Brazilian Portuguese | 138k |
1@misc{nvidia_nemotron_3_ultra_2026,
2 title = {Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning},
3 author = {{NVIDIA}},
4 year = {2026},
5 url = {https://research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Ultra-Technical-Report.pdf},
6 note = {White Paper}
7}