Views
No views yet
nvidia/Nemotron-Labs-TwoTower-30B-A3B-Base-BF16.NemotronHTwoTowerForCausalLM and repaired for Atlas causal inference. The repaired payload is intended for the OpenAI-compatible Atlas inference API using the context tower.context_towerup_proj and down_proj1ATLAS_TARGET_MODEL=nemotron-3-nano-30b-a3b \
2ATLAS_TARGET_QUANT=nvfp4 \
3CUDARC_CUDA_VERSION=12000 \
4./target/debug/spark serve \
5 --model-from-path /path/to/nemotron-twotower-nvfp4 \
6 --port 8891 \
7 --max-seq-len 4096 \
8 --max-num-seqs 1 \
9 --max-batch-size 1 \
10 --gpu-memory-utilization 0.70 \
11 --kv-cache-dtype bf16 \
12 --lm-head-dtype bf16The capital of France is -> coherent answer mentioning Paris.Question: What is 2 + 2? Answer: -> 4.Write one concise sentence about the Moon: -> coherent factual sentence.