Views
No views yet
Brooooooklyn/Agents-A1-nvfp4-mlx
is an MLX mixed-weight-format quantization of
InternScience/Agents-A1,
prepared for experimental NVIDIA CUDA inference on Linux aarch64. Agents-A1 is
a 35B-A3B Qwen3.5-family mixture-of-experts agent model with 40 language-model
layers, 256 experts with top-8 routing, hybrid linear/full attention, and an
MTP-capable config. The published source checkpoint does not include an
mtp.* tensor subtree.55100f11160f545dc45545c677699ace74f6bd10.fp8_e4m3 form is Uint8 [..., N, K] weight plus BF16
[..., N, 1] scale; it is not MLX mxfp8 and it is not native W8A8
execution.| Tensor class | Stored format |
|---|---|
Routed expert switch_mlp.{gate,up,down}_proj, layers 0–31 | NVFP4 4/16 |
Shared expert {gate,up,down}_proj, layers 0–31 | NVFP4 4/16 |
| The same routed/shared FFN projections, layers 32–39 | E4M3 FP8 weight + per-output BF16 scale |
Full-attention {q,k,v,o}_proj | E4M3 FP8 weight + per-output BF16 scale |
Linear-attention in_proj_qkv, in_proj_z, out_proj | E4M3 FP8 weight + per-output BF16 scale |
lm_head | E4M3 FP8 weight + per-output BF16 scale |
Embeddings; router mlp.gate and shared_expert_gate; in_proj_a/b; GDN state, convolution, and norm tensors; all other norms | BF16 |
| MTP tensors | Not present in the published source checkpoint |
| Vision tower and merger tensors | BF16 |
nvfp4, 4-bit, group size 16, so the 192 low-class
modules inherit that default. The config carries 179 explicit
fp8_e4m3 overrides with bits: 8 and group_size: null. The final eight
FFN layers intentionally use the higher class; this is the regular,
accuracy-oriented 35B recipe rather than the all-FFN-FP4 "Fast" variant.aarch64-unknown-linux-gnu: Linux aarch64
with glibc and NVIDIA CUDA 13.0. mlx-node currently validates this experimental,
inference-only path on NVIDIA GB10 / DGX Spark (sm_121). It is not a generic
CUDA or x86_64 artifact.1git clone --branch v0.0.8 https://github.com/mlx-node/mlx-node.git
2cd mlx-node
3git submodule update --init --recursive
4yarn install
5yarn build1MLX_QWEN35_FORCE_EAGER=1 \
2MLX_QWEN35_PAGED_OVERRIDE=0 \
3 yarn oxnode your-script.tsyour-script.ts can load a locally downloaded copy:1import { loadSession } from '@mlx-node/lm';
2
3const session = await loadSession('./Agents-A1-nvfp4-mlx');
4const result = await session.send('Reply briefly: what can you help with?');
5console.log(result.text);@mlx-node/lm and @mlx-node/core 0.0.8
source tree or a newer release that explicitly supports the same Linux target
and serialized modes.mlx-node v0.0.8.
The reproducible invocation from the mlx-node repository root was:1mlx convert \
2 --input .cache/models/agents-a1 \
3 --output .cache/models/agents-a1-unsloth-nvfp4-fp8-dgx-mlx-fresh \
4 --model-type qwen3_5_moe \
5 --dtype bfloat16 \
6 --quantize \
7 --q-recipe unsloth \
8 --q-mode nvfp4metadata.total_size = 24,784,342,752 bytes. Its tensor dtypes are 874
BF16, 371 U8, and 192 U32 entries, with 371 scale sidecars and no quantization
bias sidecars.quantization and
quantization_config blocks, exact index-to-shard closure, 192 inherited
NVFP4 groups, 179 complete fp8_e4m3 groups, the expected storage dtypes and
shapes, and BF16 preservation for protected tensors. All 333 vision tensor
entries remain BF16, and there is no MTP tensor subtree.finishReason = "length", numTokens = 1, text = "OK", and
rawText = "OK". This one-token text smoke does not validate model quality,
long-context behavior, tool use, or the vision path.macOS A16 fallback only — these are not DGX/CUDA throughput results.
@mlx-node/lm, @mlx-node/core, and @mlx-node/core-darwin-arm64 0.0.8. The
run used zero warmups, a 60-second cooldown, temperature 0, reasoning effort
none, and the same 106-token prompt. Every sample generated all 512 tokens
and ended with finishReason = "length".| Metric | macOS A16 fallback median (not DGX/CUDA) |
|---|---|
| Load time | 53,908.621 ms |
| Time to first token | 1,813.556 ms |
| Prefill throughput | 58.449 tokens/s |
| Decode throughput | 59.637 tokens/s |
| Generation wall time | 10,610.061 ms |
| Total wall time | 64,268.905 ms |
benchmark.json. Load time varied strongly because the
weights were read from external storage and OS page-cache state differed
between fresh processes; treat that median as specific to this run. This
fallback benchmark did not exercise DGX, CUDA, native W4A4/W8A8 execution, or
the vision path, and it must not be used to infer model quality, memory
requirements, or parity with upstream execution.