Views
No views yet
| file | use |
|---|---|
*.safetensors (13 shards, BF16) | transformers / vLLM / Space inference (trust_remote_code) |
warden-nemotron-3-nano-30b-Q8_0.gguf … Q3_K_S.gguf | llama.cpp; the game picks a tier by system RAM |
linear_qkv, linear_proj, in_proj, out_proj
(Mamba + attention; the fused grouped-MoE experts are not targeted)nvcr.io/nvidia/nemo:25.11.nemotron_3_nano) on 2× DGX Spark (GB10)| metric | this model | gate |
|---|---|---|
| JSON tool-call validity | 100% | ≥90% |
| persona-clean dialogue | 100% | ≥90% |
| persona breaks | 0 | 0 |
| injection canary leaks | 0 | 0 |
chat_template_kwargs: {"enable_thinking": false} (the game keeps it off
for latency). Recommended sampling: temperature 0.6, top_p 0.95.