Views
No views yet
train_sft split of HuggingFaceH4/ultrachat_200k. Responses were regenerated by NVIDIA-Nemotron-3-Super-120B-A12B.1# In virtual environment with speculators installed
2python scripts/prepare_data.py \
3 --model nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8 \
4 --data ./regenerated_data.jsonl \
5 --output ./output \
6 --seq-length 81921# In (separate) virtual environment with vllm installed
2CUDA_VISIBLE_DEVICES=0,1,2,3 vllm_venv/bin/python scripts/launch_vllm.py \
3 nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8 \
4 --target-layer-ids 7 16 25 47 69 \
5 -- --port 8000 \
6 --gpu-memory-utilization 0.9 \
7 --disable-uvicorn-access-log \
8 --tensor-parallel-size 41# In virtual environment with speculators installed
2CUDA_VISIBLE_DEVICES=4,5,6,7 torchrun \
3 --standalone \
4 --nproc_per_node 4 \
5 scripts/train.py \
6 --verifier-name-or-path nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8 \
7 --speculator-type dflash \
8 --num-layers 5 \
9 --data-path ./output \
10 --vllm-endpoint http://localhost:8000/v1 \
11 --save-path ./output/checkpoints \
12 --epochs 3 \
13 --lr 0.0006 \
14 --total-seq-len 8192 \
15 --on-missing generate \
16 --on-generate delete \
17 --seed 42 \
18 --log-freq 100 \
19 --draft-vocab-size 32000 \
20 --draft-arch llama \
21 --target-layer-ids 7 16 25 47 69 \
22 --draft-hidden-act silu \
23 --scheduler-type cosine \
24 --max-anchors 3072 \
25 --prefetch-factor 2 \
26 --num-workers 8| Base Model | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8 |
| Chat Template | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8 (use /chat/completions endpoint) |
| Format | Safetensors |
| License | Apache 2.0 |
| Validation Hardware | Nvidia B200 |
1# Install vLLM from the required PR
2pip install git+https://github.com/vllm-project/vllm.git@refs/pull/41880/head
3
4# Deploy with speculative decoding
5vllm serve nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8 \
6 --tensor-parallel-size 4 \
7 --max-model-len 16384 \
8 --speculative-config '{
9 "model": "RedHatAI/NVIDIA-Nemotron-3-Super-120B-A12B-speculator.dflash",
10 "num_speculative_tokens": 7,
11 "method": "dflash"
12 }'| Dataset | Pos 0 | Pos 1 | Pos 2 | Pos 3 | Pos 4 | Pos 5 | Pos 6 | Avg. Length |
|---|---|---|---|---|---|---|---|---|
| HumanEval | 74.0% | 51.7% | 36.1% | 25.6% | 18.6% | 13.5% | 9.7% | 3.29 |
| math_reasoning | 81.1% | 63.3% | 49.2% | 37.5% | 28.0% | 20.2% | 14.3% | 3.93 |
| qa | 67.6% | 41.0% | 24.4% | 14.6% | 8.5% | 4.7% | 2.6% | 2.63 |
| question | 68.0% | 41.3% | 24.8% | 15.4% | 9.8% | 6.3% | 4.0% | 2.70 |
| rag | 71.6% | 46.8% | 30.0% | 19.2% | 12.1% | 7.6% | 4.5% | 2.92 |
| summarization | 69.4% | 42.5% | 24.1% | 13.6% | 7.2% | 3.7% | 1.7% | 2.62 |
| tool_call | 70.8% | 46.5% | 29.6% | 19.8% | 13.7% | 9.5% | 6.2% | 2.96 |
| translation | 67.2% | 42.3% | 24.8% | 13.7% | 7.4% | 4.1% | 2.3% | 2.62 |
| writing | 67.9% | 40.8% | 24.4% | 15.4% | 9.6% | 6.0% | 3.6% | 2.68 |
| academic (long context) | 67.8% | 41.8% | 25.4% | 15.9% | 9.8% | 6.1% | 3.5% | 2.70 |
| financial (long context) | 68.1% | 41.8% | 25.4% | 15.7% | 9.9% | 6.1% | 3.7% | 2.71 |
| legal (long context) | 67.8% | 41.9% | 26.0% | 16.9% | 11.2% | 7.4% | 4.6% | 2.76 |
| agent_history_qa (long context) | 78.0% | 57.9% | 44.8% | 35.1% | 28.2% | 22.1% | 17.3% | 3.83 |
| code_repo_qa (long context) | 65.8% | 39.3% | 22.9% | 13.4% | 7.9% | 4.4% | 2.4% | 2.56 |
| detective (long context) | 62.0% | 33.5% | 16.9% | 8.3% | 4.2% | 2.2% | 1.1% | 2.28 |
| dialogue_history_qa (long context) | 63.4% | 37.8% | 22.9% | 15.4% | 5.5% | 3.4% | 1.9% | 2.50 |
| event_ordering (long context) | 64.8% | 36.9% | 19.9% | 10.6% | 6.0% | 3.4% | 2.0% | 2.44 |
| governmental (long context) | 69.2% | 44.0% | 27.8% | 18.4% | 12.2% | 8.1% | 4.9% | 2.85 |