This repository contains an oMLX oQ8 affine conversion of NVIDIA Nemotron 3.5 Lightning. The source is a 30B-total / 3B-active hybrid mixture-of-experts model that interleaves Mamba-2, sparse MoE, and attention layers. The upstream tokenizer, chat template, and generation configuration are preserved.
[!NOTE] oQ8 uses uniform 8-bit affine weights with a group size of 64.
Apple-silicon performance
This checkpoint was load-tested and generation-tested on the following machine:
Hardware
Configuration
Host
Mac Studio
Chip
Apple M3 Ultra
CPU
32 cores (24 performance + 8 efficiency)
Unified memory
256 GB
Runtime
MLX-LM 0.31.3 / MLX 0.32.0
A warmed local test produced:
Measurement
Result
Decode (median)
123.68 tokens/s
Reported peak memory
33.73 GB
Timed runs
3 × 256 generated tokens
Warm-up
32 generated tokens
Prompt
36 tokens after chat templating
The decode figure is the median of three greedy 256-token runs after a 32-token Metal-kernel warm-up. It is a practical local reference, not a controlled cross-platform benchmark. Prompt length, context growth, sampler settings, memory pressure, thermal state, and MLX/oMLX versions can materially change performance.
Reasoning mode is enabled by the upstream chat template by default. To disable it:
bash
1mlx_lm.generate \2 --model Vontra/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-oQ8 \3 --chat-template-config '{"enable_thinking": false}'\4 --prompt "Write a short hello-world program in Swift."\5 --max-tokens 256
For long prompts, begin with a conservative context limit and increase it while watching memory pressure. The configured 256K context is a model capability, not a guarantee that every host can prefill that context within its available unified memory.
Architecture
Nemotron 3.5 Lightning is a hybrid sparse model designed for efficient agentic and reasoning workloads.
Architecture detail
Upstream value
Total / active parameters
30B / 3B
Layers
52
Routed / shared experts
128 / 1
Active routed experts
6
Attention heads / KV heads
32 / 2
Hidden size
2,688
Expert intermediate size
1,856
Vocabulary size
131,072
Configured context
262,144 tokens
The upstream release is intended for coding, tool use, reasoning, research, and customization. For NVIDIA's evaluations, deployment guidance, intended use, limitations, safety information, and full architecture discussion, see the original model card.
Conversion and validation notes
Source weights: NVIDIA's BF16 checkpoint.
Quantization group size: 64.
Quantization mode: affine.
The upstream chat_template.jinja is preserved.
All 729 converted tensors and every indexed shard were checked locally.
The model was loaded and exercised through end-to-end generation on Apple silicon.
Quantization can reduce output quality relative to BF16; use a higher-precision variant when quality matters more than memory use.
This is a community conversion, not an official NVIDIA release. Validate quality and numerical behavior on your own representative workload before production use.
License and attribution
The upstream model is released under the OpenMDW License Agreement, version 1.1. A copy is included in this repository; review it before use or redistribution.
All model design, training, benchmark, and upstream documentation credit belongs to NVIDIA and the original contributors. The MLX conversion, Apple-silicon validation, and packaging are provided by Vontra.
Published peak memory: 33.73 GB; estimated starting tier: 64GB, leaving about 30 GB nominal headroom. The collections use published M3 Studio peaks with at least 25% nominal headroom; fit on other Macs is an estimate, and full context is not guaranteed. Start with short context and one request.
Runtime and evidence
The exact tested oMLX application version is not recorded here; a library version is not an app version. The original performance tables retain their benchmark conditions and speed figures; this documentation update adds no new test results.
Add the downloaded folder to oMLX model directories, refresh the list, and follow this card's architecture and MTP compatibility requirements before loading.
Try this in a new chat with a 128-token output limit:
Explain why the sky looks blue in three short sentences.
This is a demo prompt to try, not a recorded successful run; a captured demonstration for this documentation update is not yet available.