Views
No views yet
Brooooooklyn/Qwen3.6-27B-NVFP4-mlx
is an MLX mixed-weight-format quantization of
Qwen/Qwen3.6-27B, prepared for
experimental NVIDIA CUDA inference on Linux aarch64. The model has a 64-layer
dense Qwen3.5-family language backbone with 48 linear-attention and 16
full-attention layers, a vision tower, and one MTP layer.1b559cf7215ebe67ff10758e14f6293ba883223b.amax / 6 in E4M3, and on
real weights those land in E4M3's subnormal band, carrying three mantissa bits
instead of four -- or rounding to the zero code, which decodes all 16 weights in
the block to zero. The fold is exact, so the layer's function is unchanged.fp8_e4m3 form is
Uint8 [N, K] weight plus BF16 [N, 1] scale; it is not MLX mxfp8 and it is
not native W8A8 execution.| Tensor class | Stored format |
|---|---|
Dense FFN {gate,up,down}_proj, layers 0–55 | NVFP4 4/16 |
Dense FFN {gate,up,down}_proj, layers 56–63 | E4M3 FP8 weight + per-output BF16 scale |
Full-attention {q,k,v,o}_proj | E4M3 FP8 weight + per-output BF16 scale |
Linear-attention in_proj_qkv, in_proj_z, out_proj | E4M3 FP8 weight + per-output BF16 scale |
lm_head | E4M3 FP8 weight + per-output BF16 scale |
Embeddings; in_proj_a/b; GDN state, convolution, and norm tensors; all other norms | BF16 |
Entire 15-tensor mtp.* subtree | BF16 |
| Vision tower and merger tensors | BF16 |
nvfp4, 4-bit, group size 16, so the 168 low-class
modules inherit that default. The config carries 233 explicit
fp8_e4m3 overrides with bits: 8 and group_size: null. The final eight
dense FFNs intentionally use the higher class.aarch64-unknown-linux-gnu: Linux aarch64
with glibc and NVIDIA CUDA 13.0. mlx-node currently validates this experimental,
inference-only path on NVIDIA GB10 / DGX Spark (sm_121). It is not a generic
CUDA or x86_64 artifact.1git clone --branch v0.0.8 https://github.com/mlx-node/mlx-node.git
2cd mlx-node
3git submodule update --init --recursive
4yarn install
5yarn build1MLX_QWEN35_FORCE_EAGER=1 \
2MLX_QWEN35_PAGED_OVERRIDE=0 \
3 yarn oxnode your-script.tsyour-script.ts can load a locally downloaded copy:1import { loadSession } from '@mlx-node/lm';
2
3const session = await loadSession('./Qwen3.6-27B-NVFP4-mlx');
4const result = await session.send('Explain the purpose of a unit test in one sentence.');
5console.log(result.text);@mlx-node/lm and @mlx-node/core 0.0.8
source tree or a newer release that explicitly supports the same Linux target
and serialized modes. The MTP weights are preserved for checkpoint fidelity,
but CUDA speculative decoding is unsupported in this preview; their presence
does not establish DGX MTP support or validation.mlx-node at or after PR #131, which made the tuned MX
weight encoders and the NVFP4 power-of-two lift unconditional. v0.0.8 reproduces
the earlier revision of this repository, not the current weights.
The reproducible invocation from the mlx-node repository root was:1mlx convert \
2 --input .cache/models/qwen3.6-27b \
3 --output .cache/models/qwen3.6-27b-unsloth-nvfp4-fp8-dgx-mlx-fresh \
4 --model-type qwen3_5 \
5 --dtype bfloat16 \
6 --quantize \
7 --q-recipe unsloth \
8 --q-mode nvfp4-fresh suffix is only the local conversion directory; the canonical Hub
repository is the ID shown at the top of this card.metadata.total_size = 23,417,338,336 bytes. Its tensor dtypes are
1,031 BF16, 401 U8, and 168 U32 entries, with 401 scale sidecars and no
quantization bias sidecars.quantization and
quantization_config blocks, exact index-to-shard closure, 168 inherited
NVFP4 groups, 233 complete fp8_e4m3 groups, the expected storage dtypes and
shapes, and BF16 preservation for protected tensors. All 333 vision tensor
entries and all 15 MTP tensor entries remain BF16 without quantization
sidecars.finishReason = "length", numTokens = 1, text = "OK", and
rawText = "OK". This one-token text smoke does not validate model quality,
long-context behavior, tool use, the vision path, or speculative decoding.macOS A16 fallback only — these are not DGX/CUDA throughput results.
@mlx-node/lm, @mlx-node/core, and @mlx-node/core-darwin-arm64 0.0.8. The
run used zero warmups, a 60-second cooldown, temperature 0, reasoning effort
none, and the same 106-token prompt. Every sample generated all 512 tokens
and ended with finishReason = "length".1MLX_QWEN35_FORCE_EAGER=1 \
2MLX_QWEN35_PAGED_OVERRIDE=0 \
3 oxnode scripts/benchmark-model.ts \
4 .cache/models/qwen3.6-27b-unsloth-nvfp4-fp8-dgx-mlx-fresh \
5 --output .cache/benchmarks/qwen3.6-27b-nvfp4-macos-fallback.json| Metric | macOS A16 fallback median (not DGX/CUDA) |
|---|---|
| Load time | 435,381.661 ms |
| Time to first token | 6,211.594 ms |
| Prefill throughput | 17.065 tokens/s |
| Decode throughput | 15.382 tokens/s |
| Generation wall time | 39,775.678 ms |
| Total wall time | 473,529.119 ms |
benchmark.json. Load time varied strongly because the
weights were read from external storage and OS page-cache state differed
between fresh processes; treat that median as specific to this run. This
fallback benchmark did not exercise DGX, CUDA, native W4A4/W8A8 execution, the
vision path, or speculative decoding, and it must not be used to infer model
quality, memory requirements, or parity with upstream execution.