There is no universal current default launcher across the historical
profiles retained in this card. For a new deployment, do not infer a default
from an older section:
Profile
Status
Scope
Portable r17 graph/NVFP4 package
Release candidate; not yet the default
Reproducible build and focused source/ABI tests; end-to-end GPU replay is still required
Dedicated-R7 r17 candidate used for the results below
Locally qualified evidence source
TP4/DCP4/MTP3, full graphs, NVFP4 DS-MLA, tasks and 260K retrieval
r17 exact-G64-Q FP8 reader
Provisional quality experiment
Eager KLD result; not the graph/speed-qualified serving profile
Gilded Gnosis r34
Historical qualified reference (2026-08-10)
TP4/DCP1, 65,536-token limit; not the present DCP4 long-context candidate
Earlier max-KV launchers
Historical measurements
Preserve provenance only; use their exact image and configuration when reproducing them
The portable r17 package becomes the recommended downloadable path only after
it repeats the end-to-end graph, task, 260K retrieval, and numerical gates on
the corrected dual-ABI build.
Sealed r17 FP8-reader KLD comparison (2026-08-22)
An immutable five-fresh-boot acceptance panel plus a factorial G64-Q-only
follow-up compared the r17 FP8 MLA reader paths. Lower KLD is better.
reader regime
five-run mean KLD
median
sample SD
result vs stock
stock reader flags
0.06033384
0.06080258
0.00122809
control
exact G64 Q only
0.05973021
0.05894077
0.00233983
-1.00% KLD; <0.060 gate passed
BF16 P.V
0.06155303
0.06090119
0.00197113
+2.02% KLD
exact G64 Q + BF16 P.V
0.06105802
0.06036178
0.00137015
+1.20% KLD
The five G64-Q-only values were 0.05855139, 0.06383442, 0.05894077,
0.05927899, and 0.05804549. The arithmetic mean passed the predeclared
<0.060 gate, but one high boot increased sample SD and the five-boot Welch
contrast with stock remains inconclusive (p=0.6276).
Provisional quality launcher:server.sh launches the digest-pinned r17
runtime with exact G64 Q-only enabled, MTP-3, eager execution, and a fixed 4 GiB KV
reservation per rank. The validated startup initialized 313,856 FP8 MLA KV
tokens at a 262,144-token model limit (1.20x maximum-request concurrency).
This is not the speed-qualified release default. One attempted routing entered
the ordinary per-expert branch and hit nonzero during capture; that does not
show that r17 or its EXL3 fallback cannot run in CUDA graphs. Full r17 graph
replay with nvfp4_ds_mla and measured speed are mandatory before the quality
battery is rerun. The three required source overlays and their
checksums ship in runtime/r17-g64-q-only.
Neither BF16 P.V combination advances. G64 uses the existing Q-row padding for
eight exact G64 scales, leaves BF16 P.V and BF16 QK disabled, and adds no
KV-record bytes, ABI change, or dynamic-shared-memory growth.
The panel used one pinned 2,048-token WikiText window, 2,047 full-vocabulary
teacher-forced positions, the official BF16-reference logits, FP8 MLA KV with
BF16 RoPE, TP4/DCP4, and 300 W per GPU (1,200 W aggregate). Boot repeats on one
window are not independent prompts, so the report does not claim general task
quality or production qualification.
This checkpoint has now been qualified on an r17-derived, dedicated R7 runtime
with full CUDA graphs. This section is separate from the FP8-reader KLD panel
above: the KLD values, KV formats, and runtime paths are not interchangeable.
ONLINE_QUANT=none for the tested checkpoint layout
enforce_eager=False
CUDAGraphMode.FULL_AND_PIECEWISE; FULL graphs captured on all ranks
4 GiB KV reservation per rank: 463,872 physical KV tokens
262,144-token model limit
Use the files in runtime/r17-graph-nvfp4 for
this profile. The self-contained Dockerfile starts from the pinned public r17
digest and applies the reviewed vLLM and ExLlamaV3 patch series. The wrapper
checks the calibrated-scale hash, selected launch flags, graph log markers, KV
allocation, and graph replay before declaring the server qualified; it does
not independently attest every source or checkpoint hash.
Publication gate: the portable image passes its build, ABI assertions, and 68
focused vLLM tests, but has not yet repeated the end-to-end GPU battery. The
measurements below belong to the earlier validated candidate whose effective
overlays this portable build reconstructs; do not promote the portable image
to the default until replayed.
Measured performance
Four nominal 300 W GPU caps, 1,200 W aggregate:
Profile
Result
Prefill 8K
1,203 tok/s
Prefill 64K
1,161 tok/s
Prefill 128K
1,107 tok/s
Decode C1
74.0 tok/s
Decode C2
118.7 tok/s
Decode C4
192.5 tok/s
Five-run task validation
Post-reseat quality used nominal caps of 300/300/300/275 W:
Profile
Verdict
Mean output tokens
Aggregate generation
LAVD
5/5 correct (2 exact, 3 near)
8,728.8
59.27 tok/s
Estonia
5/5 pass
1,744.2
55.65 tok/s
Hotel
5/5 exact
20,143.8
55.53 tok/s
Estonia-long
5/5 pass
1,951.4
54.82 tok/s
No successful run hit its output-token limit. The small TTFT figures in the
task receipt are prefix-cache affected and are not uncached-prefill claims.
260K retrieval
Five 260,000-token prompts retrieved their planted record exactly at 5%, 25%,
50%, 75%, and 95% depth. The later probes benefited from an increasing shared
prefix, so their elapsed times are not an uncached-prefill comparison.
BF16-reference KLD
Four comparable eager-harness 2,047-position observations exist for the exact
dedicated-R7/FP32-route/static-scale/NVFP4 configuration: 0.0639764262,
0.0746190367, 0.0704009716, and 0.0676657456. Their descriptive mean is
0.0691655450. These comprise one sentinel plus three completed runs from a
planned five-run panel. Run 4's log ends during checkpoint loading without a
recorded cause, and run 5 never began, so there is no five-run acceptance
mean. The numerical harness used eager execution with CUDA graphs disabled;
graph qualification comes from the separate serving/task/needle receipts.
This numerical result should not be conflated with the five-run FP8 MLA reader
result above (0.05973021 for exact G64-Q-only). The FP8 reader study used a
different KV format and was explicitly not the graph/speed-qualified serving
default.
Hardware and claim boundaries
The successful post-reseat batteries recorded no Xid, PCIe/AER, CUDA,
traceback, engine-fatal, hardware-slowdown, thermal-slowdown, or power-brake
event. Maximum observed temperature was 89 C; GPU 3/C1 reached 76 C. Telemetry
can report short samples above a configured power limit, so the stated caps
are nominal configuration limits rather than hard instantaneous clamps.
This qualifies the exact local checkpoint file set identified by the published
configuration, quantization-config, safetensors-index, and scale hashes; the
tested image; TP4/DCP4/MTP3 topology; NVFP4 DS-MLA KV contract; maximum four
sequences; and contexts through 260,000 tokens on the reference host. It does
not bind that tested file set to a particular Hub commit or establish multi-prompt KLD
generalization, concurrency above four, FP8 KV graph qualification, or other
GPU topologies.
The shared-expert MLP tensors are now stored as the original BF16 weights
instead of pre-encoded EXL3 payloads, and the serving stack encodes them as
one merged K6 payload at load. Routed R7 expert payloads and rotations are
byte-identical to the previous revision; nothing else about the quantization
changed.
If you pulled this repo before 2026-08-10, re-download
model.safetensors.index.json, config.json, quantization_config.json,
the model-layer-003 .. model-layer-078 carrier shards, and the new
model-sharedbf16.safetensors. hf download picks up the delta on its own.
What changed and why
Festr isolated a decode-throughput loss to the previous shared-expert
layout. mlp.shared_experts.{gate_proj,up_proj,down_proj} were stored as
three separately encoded K6 EXL3 trellis payloads per MoE layer (layers 3-77;
the layer-78 MTP draft shared experts were already BF16). Shared gate and up
as two separate payloads force two small-M GEMM launches per routed layer
per decode step where a merged payload needs one -- 75 extra kernel launches
per decode step at MTP0. His diagnostic, TP4 / DCP1 / MTP0, split vs merged
shared gate+up:
decode tok/s
C1
C4
C8
split gate/up (previous layout)
50.62
157.34
258.81
merged gate+up
53.86
169.07
281.72
This revision therefore stores the shared experts unencoded:
A new shard model-sharedbf16.safetensors (~5.7 GB) carries the 228
shared-expert BF16 tensors: gate/up/down for layers 3-78, including the
MTP-78 draft layer.
The 76 carrier shards model-layer-003 .. model-layer-078 are rewritten
without their shared-expert entries. All other shards are unchanged.
tensor_storage in quantization_config.json (and the copy embedded in
config.json) drops its 225 shared-expert module entries, so a loader sees
the shared experts as ordinary BF16 modules.
Routed experts (r7-experts-layer-*.safetensors) and all rotations:
byte-identical, untouched.
Serving
Set ONLINE_QUANT=exl3-b6. At load the runtime concatenates gate and up
while still BF16 and encodes one merged K6 payload per layer; the encoded
result lands in the JIT/weight cache, so the cost is paid once, on first
load. One payload, one launch.
Without ONLINE_QUANT the checkpoint still serves, with the shared experts
running in plain BF16: correct output, roughly 3.5 GB more weight memory
model-wide than the online-encoded path (5.74 GB BF16 versus 2.20 GB
encoded), and none of the merged-launch decode gain.
Quality: single pass from source
The BF16 tensors are the original shared-expert weights -- verified
bit-identical across the
willfalco 3.42-bpw checkpoint
and willfalco 3.25-bpw checkpoint
lineage checkpoints by independent ranged-read sha256 sampling. The online
merged K6 encode is one quantization pass from that source, exactly as the
previous layout's offline K6 encode was. Nothing is re-quantized from an
already-quantized representation.
KLD, 5-run gate against the same BF16 reference logits:
layout
mean
sd
previous (offline split K6 shared)
0.062450
0.001533
this (BF16 shared, ONLINE_QUANT=exl3-b6)
0.064250
0.000383
Verdict: PASS with documented delta -- 0.064250 +/- 0.000383 (5 runs, full exl3-b6 serving policy) vs 0.062450 +/- 0.001533 for the previous layout (5-run gate). Delta +0.0018; the new mean sits below the previous layout's own worst run (0.064910). Traded for +9.5% mean decode throughput and +20.5% KV-cache capacity.
Decode at MTP-3 on the reference rig (4x RTX PRO 6000, TP4 + DCP4):
On 2026-08-10 — the same day this layout shipped — the local-inference-lab
release pipeline published Gilded Gnosis r34 with this checkpoint, at this
exact revision, as its qualified reference:
The r34 runtime keeps the routed experts in their serialized K3/K4/K5 Trellis
formats and encodes the BF16 shared experts into cached merged K6 projections
(this repo's layout, consumed as intended). Loader contract: InstantTensor
BUFFERED with borrowed-buffer consumption. Qualified profile: TP4/DCP1,
B12X A16, B12X sparse MLA, NVFP4 DS-MLA KV, MTP-3, 8 sequences, graph cap 32,
model limit 65,536, GMU 0.98.
Release-gate measurements (their receipt, not mine):
profile
C1
C4
C8
prefill 8K
KV tokens
MTP0 / GMU .97
53.80
171.30
283.05
3,253
82,816
MTP3 / GMU .98
121.25
297.69
436.23
3,239
75,072
MTP-3 strict acceptance 65.44%. FULL decode graphs covered every configured
size; target verification, all three MTP forwards, and prefill remained
graph-captured. Focused vLLM, B12X host/GPU, runtime-contract, startup,
deterministic-output, and checksum gates passed.
Scope note, in the release's own words: the r34 receipt does not qualify
R7 at DCP>1, standard NVFP4, or NF3 performance. The long-context DCP4 profile
documented above (262,144-token context, 520,960 KV tokens) is the author's
own measured configuration, validated by the KLD and throughput gates in this
card, not by the r34 receipt.
Credit where it belongs: the split-payload decode loss was isolated by
Festr, whose analysis produced this layout, and the qualification is the
work of the local-inference-lab Discord community and its release
engineering. Same-day pipeline from proposal to shipped checkpoint to
qualified release — that is what a receipts-first community looks like.
Previous layout
The pre-update revision remains available at
c55c1cd4ca42
if you need the old offline-encoded shared payloads.
The exact source lineage, independent TP4 validation, conversion tool, and
focused tests are recorded in
BF16_SHARED_ONLINE_K6_VALIDATION.md.
Historical and provisional launchers
Provisional eager FP8/G64 research image (digest-pinned r17 base):
The image is built for sm_120a (Blackwell). It will not run elsewhere
without a rebuild, and the checkpoint needs a mixed-bit loader -- see the
loader compatibility warning below.
Fused MoE path (long context)
An alternative serving path that runs the routed experts on SparkInfer's fused
mixed-Trellis MoE kernel. It is opt-in (VLLM_EXL3_R7_FUSED=1) and targets
long-context work: more KV capacity per GB than the numbers in "Run it" above,
at a slightly higher KLD because it pairs with the 4-bit nvfp4_ds_mla cache
rather than fp8.
Measured on 4x RTX PRO 6000 Blackwell (SM120a, 96 GB, PCIe Gen5, no NVLink),
TP4 + DCP4, MTP-3:
value
KV cache
355,328 tokens
Max concurrency
1.36x at 262,144 tokens/request
KLD vs BF16 reference logits
0.069527 (4-run mean; fifth interrupted)
Prefill @ 8K
1,725 tok/s
Decode
75.6 tok/s
Full configuration for every number above: max_model_len 262144,
gpu_memory_utilization 0.955, max_num_batched_tokens 2048, max_num_seqs 4,
CUDA-graph size 32, kv_cache_dtype=nvfp4_ds_mlawith outer scales, BF16
RoPE, VLLM_EXL3_PREFILL_BLOCK_M=64, VLLM_EXL3_PREFILL_CAPACITY=1024,
VLLM_DCP_INDEXER_SHARDS=4, 48 fused layers. Prefill measured with a unique
prompt prefix so the prefix cache cannot serve it; decode with ignore_eos over
256 tokens.
Non-fused reference on the identical rig and settings: 497,408 tokens KV,
1,757 tok/s prefill, 73.5 tok/s decode. The fused path trades KV capacity for
decode throughput.
The 0.061282 figure in "Run it" is fp8 KV cache; the 0.069527 here is
nvfp4 KV. They are different cache formats measured against the same
reference logits, so compare them with that in mind rather than as a
regression.
The outer scales file is required
nvfp4_mla_outer_scales.json now ships in this
repo. Mount it and point VLLM_NVFP4_MLA_SCALES_FILE at it. Measured on the
identical build, changing nothing else:
nvfp4 KV
with outer scales
without
BF16 RoPE
0.069527
0.099717
FP8 RoPE
0.075542
0.114359
Omitting it costs about 30% KLD for no memory or throughput benefit. It is a
per-layer outer-scale calibration (wikitext-2-raw-v1, 2,048 context).
The table also shows why BF16 RoPE is the default here: FP8 RoPE yields roughly
16% more KV tokens but measured +8.7% KLD with scales applied.
What had to be fixed
The fused path was previously unusable on this checkpoint and would have
produced garbage output, not a mild regression. SparkInfer bounded the FC2
(down-projection) tier-local expert index using the FC1 slot count. Because
this checkpoint chooses the trellis bit width per (expert, projection), FC2 holds
more experts than FC1 -- tier1 carries 231 down-projection experts against 77
gate/up -- so most down tiles were rejected and their output silently dropped:
6,653 of 12,288 down-projection expert slots, 54.14%, across the 48 fused
layers.
The failure was fluent rather than obviously broken. zero_fc2_output=False and
the FC2 buffer aliases rotation_gate, so a rejected tile left gate-rotated
hidden states in place, which were then Hadamard-rotated, scaled by down_svh,
router-weighted and accumulated. Measured KLD 2.36 with output that
hallucinated case law and degenerated into verbatim repetition.
The fix is patches/patch_sparkinfer_projection_tiers.py (fc1_bound_ok).
Built for sm_120a (Blackwell); it will not run on other architectures without a
rebuild. Every patch in patches/ is already applied inside it -- that directory
is only needed if you are building your own image from an r28-or-later SparkInfer
base.
An r28-or-later base is required, not merely preferred:
patch_r7_broadcast_rotations.py depends on SparkInfer ABI-v6
broadcast_suh/broadcast_svh, which earlier bases do not expose. Without it
the loader expands one shared rotation row per layer into 256 identical copies,
costing about 0.5 GiB per rank.
Caveats
KLD 0.069527 is a 4-run mean; the runner was interrupted before a fifth. The
separation from the unscaled 0.099717 is far larger than the run-to-run
spread, but treat the third decimal as provisional.
Prefill and decode figures are single probes on a rig that has shown
double-digit container-to-container variance. Treat them as indicative.
The occupancy patches (patch_moe_deadscale_2cta.py,
patch_moe_stages3_only.py) raise the fused kernel from 8 to 15.4 warps/SM,
but measured at parity end-to-end on this rig. They are included for
completeness, not as a speed claim.
What is different about this quantization
Full-width down-projection encoding
Each expert down projection is encoded jointly across its full 2048-channel
input dimension. It is not independently quantized as four serving-specific
512-channel slices. Error correction can therefore compensate across the
whole tensor before a loader slices it for tensor parallelism.
Down calibration uses reconstructed gate/up outputs
Gate and up are encoded at candidate bit widths and reconstructed through the
same quantized representation that will be stored. Their reconstructed SwiGLU
output supplies the conditional calibration input for down. Down is therefore
optimized for the quantized gate/up tensors it will actually follow, not for
unquantized BF16 gate/up outputs.
Router-mass-weighted exact bit budgeting
Each layer contains 768 routed-expert tensors. Starting all tensors at 3 bits
uses 2,304 bit units; the exact 3.5 bpw target is 2,688 units, leaving exactly
384 one-bit upgrades. Candidate 3-, 4-, and 5-bit losses are weighted by the
captured float32 routing mass. A dynamic program spends all 384 upgrades, and
a tensor may receive a 4→5 upgrade only after its 3→4 upgrade is selected.
This is why an equal-average “barbell” split of very high and very low bit
widths was rejected: quantization error falls smoothly with bit width, so the
high end wastes marginal bits while the low end crosses a steep error cliff.
Expert-private intermediate reordering
Five intermediate-channel orderings are considered for each expert. The
chosen ordering is baked consistently into gate output, up output, and down
input. SwiGLU is elementwise, so a consistent permutation is functionally
free at serving time while giving the error-correcting walk a better row order.
The 6144-dimensional residual/model-space sides are shared per layer:
gate_up_suh is shared by all gate/up inputs in the layer.
down_svh is shared by all down outputs in the layer.
The 2048-dimensional intermediate sides remain private to each expert:
gate output (gate_svh)
up output (up_svh)
down input (down_suh)
The residual space must be common because tokens enter all routed experts in
one coordinate system and expert outputs are routing-weighted and summed back
into that system. The intermediate space exists only inside one expert, so
each expert can choose the sign-vector draw that best conditions its own
weights and activations without imposing a layer-wide compromise. Twelve
candidate draws are searched.
Per-128-channel scales folded into the stored representation
Scales are searched independently on a 128-channel grid and folded into the
existing per-element sign representation. This gives finer conditioning
without a separate runtime scale tensor.
Topology-neutral schema v2
Routed-expert tensors are stored whole with a per-tensor bit map. Tensor- or
expert-parallel slicing happens at load time on 128-channel boundaries; no
four-GPU topology is baked into the files.
Authoritative corrected local build
The corrected model did not rerun the expensive bit/permutation search. It
recovered and froze all 75 layer decisions from R10, corrected the absolute
normalization/global-scaling math, and rebuilt layers 3–77 causally on four
local SM120 GPUs. Layer L+1 was calibrated from the corrected installed
output of layer L.
For every layer, the successful supervisor performed:
a four-GPU attention/router-only flat capture;
one streamed absolute-normalization/GSS fit;
four pinned GPU workers consuming a dynamic 256-expert queue;
a corrected successor forward to create the next layer's input state;
atomic promotion and a durable layer seal; and
reclamation of reproducible capture, predecessor-state, expert-mini-shard,
and no-longer-needed BF16 source-window data.
The complete code, tests, decisions, receipts, and storage runbook are in
reproducibility/local-corrected-v1.
The 75 routed shards total 317,347,848,944 bytes and their sidecar manifests
total 350,486,725 bytes. All 75 layer seals are included.
Historical rental-box R10 run
The files under reproducibility/r10 record the B300
capture/search and first encoding attempt. They remain important provenance
for the corpus, deterministic prompt plan, inventories, and frozen allocation
decisions, but the old encoder is superseded for reproducing corrected routed
payload bytes.
1. Seal inputs and runtime
The BF16 source, carrier checkpoint, numeric core, compiled extension, package
versions, and runtime Python files were inventoried before timed work. One
scalar SafeTensors serializer defect was repaired after that launch-time seal;
the published runtime_inventory.r10.json is retained as the literal launch
record, while DEPLOYED_CODE.sha256 binds the exact post-repair code that
produced the completed layer shards.
2. Flat capture on NVMe
r10_capture.py performs a BF16 source forward and writes one flat capture per
MoE layer through LayerCalibRAM/memory-mapped storage. The 75 captures total
about 970 GB, so they live on NVMe rather than /dev/shm.
The deterministic capture plan selected 1,773 prompts totaling 1,049,589
tokens. The exact selection and the complete 75-layer capture summary are in
reproducibility/r10/capture.
Capture and encoding were operationally pipelined: GPU 0 published captures
in layer order while the other GPUs consumed already-sealed captures. After
all 75 captures completed, GPU 0 joined the encoder pool.
3. Dynamic 12-worker encoding
Six B300 GPUs run two pinned workers each. Workers pull completed layer
captures from a SQLite dynamic queue rather than receiving static layer
ranges. The deployed settings are:
text
1layers 3-77
2workers 12 (two per GPU)
3GPUs 6
4rotation draws 12
5held-out rows 4096
6minimum fit rows 1024
7row chunk 4096
8factor cache 512 MiB per worker
9sigma regularizer 0.025
10CPU threads 36 per worker
Each completed layer is emitted atomically as one
r7-experts-layer-NNN.safetensors shard plus its JSON manifest. A layer does
not enter the completed queue state until both artifacts exist.
4. Assemble, upload, and verify
The final assembler hard-links complete expert shards, hard-links clean
carrier shards, rewrites only carrier shards that mix retained and replaced
tensors, and creates the final tensor index and manifest. Progressive uploads
use the same final expert-shard basenames, so the authoritative final folder
upload reuses the already-present Hub objects.
What was removed for wall-clock speed
The accuracy design above was not weakened. Work that did not change emitted
bytes was removed from the timed encoder: payload/runtime re-hashing, routing
audits, fixed-point successor passes, install-and-forward checks, pack/unpack
repeat decodes, functional oracles, and assembly/conversion. Search,
allocation, reconstruction-based down calibration, rotations, permutations,
and final encoding remain.
The result is sealed by construction and written atomically. Full checkpoint
assembly and structural verification occur after all 75 layers complete.
There is deliberately no claim that expensive functional evaluation ran
during quantization.
exllamav3 numeric extension built from v0.0.43 for sm_103
4.295 TB NVMe filesystem and a 1.509 TB memory cgroup
Uncached completed layers have taken roughly 130–138 minutes per worker,
including search, all 256 expert probes, exact allocation, final encoding, and
atomic emission. Twelve concurrent slots turn that into waves; it is not a
75× serial runtime. The uploader and memory-cache reclaimer run at reduced
priority so GPU encoding remains dominant.
Corrected KLD result
The corrected local checkpoint was measured five times against the same BF16
reference logits, with standard FP8 KV cache, BF16 RoPE, TP4/DCP4, one
2,048-token context, and 2,047 scored positions per run:
The reference-logits SHA-256 is
87f992a689c054a0548a4b3863da6c809f9239beacd5786d0401e45904fec063.
The exact result JSON, raw run logs, evaluation-code hashes, runner, and
scoring script are published with the corrected bundle.
Reproducing the corrected run
Start with the exact guide in
reproducibility/local-corrected-v1/README.md.
It explains the required checkpoint identities, decision recovery, local
inventory rebinding, layer-3 causal anchor, rolling execution, seal adoption,
NVMe free-space frontier, BF16 shard paging, safe cleanup order, assembly, and
KLD command. The old reproducibility/r10/README.md
is historical lineage, not the corrected byte-production runbook.
Tuning: prefill route block size
The single largest configuration win found on this checkpoint. The default
VLLM_EXL3_PREFILL_BLOCK_M=64 leaves substantial performance on the table.
Measured end-to-end on the reference rig (4x RTX PRO 6000, TP4 + DCP4, MTP-3,
nvfp4_ds_mla KV, max_model_len 262144, max_num_batched_tokens 2048,
util 0.95). Output verified coherent and unchanged at every setting:
VLLM_EXL3_PREFILL_BLOCK_M
64 (default)
32
16
8
Prefill 8K (tok/s)
1157
1403
1471
1674
Prefill 64K (tok/s)
1134
1320
1397
1600
Prefill 128K (tok/s)
1112
1211
1308
1514
Decode (tok/s)
57.84
59.55
63.44
70.43
KV capacity (tokens)
441,344
446,720
448,512
448,512
At block_m=8: prefill +44.7% / +41.1% / +36.2%, decode +21.8%,
KV +1.6%. Prefill, decode and KV all improve together -- there is no
trade-off to balance.
Why
Nsight Compute on the isolated MoE kernel (single GPU, synthetic weights, so
no TP collectives to deadlock the profiler) at the stock block_m=64:
DRAM Throughput 24.03% <- memory not saturated
Compute (SM) Throughput 35.96% <- compute not saturated either
Registers Per Thread 244
Dynamic Shared Memory 101.38 KB of 102.40 KB configured
Block Limit Registers 1
Block Limit Shared Mem 1
Theoretical / Achieved Occupancy 16.67%
Dropping to block_m=16 moves DRAM throughput 24.03% -> 64.22% and halves
kernel duration (1530 -> 575 us), with registers 244 -> 144 and shared memory
101.38 -> 52.22 KB.
Occupancy does not change (16.67% either way). The kernel is a cooperative
one-grid launch -- grid size equals the SM count, so exactly one block per SM
by construction. The gain is instead the zero padding a 64-row route block
carries when it holds far fewer live rows; smaller blocks waste less.
block_m below 32 additionally requires a register-count table entry for the
cta_m_blocks=1 specialization ((256, 1, 16, 4, False)), absent upstream;
without it the launch model raises
missing W4A16 register count for NVFP4 BF16 specialization.
Levers that did NOT help
Measured and rejected, so nobody repeats them:
Lever
Result
FC1/FC2 tile config
tile_n=128 already optimal (tn=64 was 18% slower at m=128; tn=256 unsupported)
_STAGES pipeline depth
4 already optimal; 3 within noise; 2 was 27% slower
VLLM_EXL3_TRELLIS_MAX_M 32 -> 64
within noise (+0.7-1.0% prefill, decode flat)
max_num_batched_tokens 2048 -> 3072
trade-off, not a win: prefill +2.7-8.2% but KV -18% (448,512 -> 366,336)
KV cache
The serving stack uses MLA (kv_lora_rank=512, qk_rope_head_dim=64), so the
KV cache is one latent per token per layer rather than per-head K and V, and
DCP4 shards it across the four GPUs (--dcp-kv-cache-interleave-size 64).
Every row below is a number printed by vLLM at startup on the reference rig
(4x RTX PRO 6000, 96 GB, TP4 + DCP4, MTP-3, CUDA graphs on, nvfp4_ds_mla KV).
A KV number is meaningless without its utilization, KV dtype, context length and
batch budget, so all of them are listed.
util
KV dtype
max_model_len
max_num_batched_tokens
MTP
KV capacity
0.95
nvfp4_ds_mla
8,192
2,048
3
457,728 tokens
0.95
nvfp4_ds_mla
not recorded
3,072
3
340,224 tokens
0.97
nvfp4_ds_mla
262,144
1,280
3
502,016 tokens (1.91x)
0.97
nvfp4_ds_mla
262,144
1,536
3
489,984 tokens
0.97
nvfp4_ds_mla
262,144
3,072
3
417,280 tokens
0.955
nvfp4_ds_mla
262,144
2,048
3
262,912 tokens
The first row is a historical dedicated-R7 profile
(exl3_moe_r7_fused), measured 2026-08-05. The second is the configuration the throughput numbers in
RESULTS.md were taken under; its max_model_len was not recorded
alongside the figure, so treat it as indicative rather than reproducible.
Capacity moves inversely with max_num_batched_tokens: dropping 3,072 -> 2,048
returned roughly 117k tokens of KV, because the prefill scratch arena and
profiling peak shrink with the batch budget.
About the 1.13M figure
server.sh carries a comment claiming ~1,132,544 tokens (2.16x at 524K context)
at GPU_MEMORY_UTILIZATION=0.96. Do not read that as usable serving
capacity. It is an auto-profile ceiling measured without the speculative
decode path engaged, and a configuration that reaches it is not one you would
serve from. It is retained in the script's comment for provenance only and is
not reproduced as a headline number here.
Pinning a smaller cache
Capacity is auto-profiled at startup. Leave NUM_GPU_BLOCKS_OVERRIDE empty to
take the maximum the utilization allows, or set a positive integer to pin a
smaller cache and leave memory for other work on the same GPUs.
Note the KLD figure in RESULTS.md was measured with fp8 KV and
BF16 RoPE, while throughput and the capacities above use nvfp4_ds_mla.
They are not interchangeable.
Max-KV serving profile (2026-08-20)
A tuned profile that reaches 502,016 KV tokens at 262,144 max-model-len — a
+41% larger cache than the previously published 355,328 — while keeping MTP-3
on DCP4 and full CUDA graphs. Full sweep, traps, and per-lever attribution are in
KV_SPEED_RESULTS.md. Turn-key launch:
serve.sh and docker-compose.yaml.
So roughly 8-10% of throughput buys 147k additional KV tokens.
The prefill figures are conservative. The reference rig is power-capped at
300 W per GPU (1200 W total against a 600 W per-card maximum) and thermally
throttles under sustained long-context prefill. An uncapped, better-cooled
4x RTX PRO 6000 box should measure higher prefill than published here. Decode
is latency-bound on PCIe collectives and is much less sensitive to the cap.
Two traps worth stating plainly
GRAPH has a hard floor of MAX_NUM_SEQS x (MTP + 1) — 16 at SEQS=4, MTP=3.
Setting GRAPH=8 buys 17,152 more KV tokens and collapses decode from
69.59 to 5.61 tok/s, because every decode step falls out of the captured full
CUDA graph. Never set it below the decode row count.
MAX_BATCHED_TOKENS moves KV inversely. Transient buffers are
max_batched-shaped and MTP-3 allocates two workspace lanes, doubling them.
Raising 1,280 -> 3,072 cost 84,736 tokens. Raise it only when you want prefill
throughput and are willing to pay KV for it.
Image requirement
This profile requires an image carrying the R7 fused MoE patches, e.g.
verdictai/glm52-exl3-sparkinfer:v39-r28-r7fused-broadcast-cu132-sm120a
(sha256:12f86065d7fe64d30dad678585e68c91f47f1f2a32bed45ccaf108382f3928ac).
The public voipmonitor/vllm:infernal-invocation-...-20260817-r17 base ships
exl3_moe_fused rather than the specialized exl3_moe_r7_fused symbol. The
portable companion package adds the dedicated projection-mixed R7 ABI while
preserving the legacy FP16 route ABI for ordinary EXL3 calls. That reconstructed
image has not yet repeated the end-to-end GPU battery, so it is not the default.
An older launcher routed around the available generic symbol and raised a
missing-symbol error; the blanket claim that r17 cannot serve the checkpoint is
superseded.
Repository map
text
1calibration/
2 reap_recall_calib.jsonl exact live corpus
3reproducibility/r10/
4 README.md superseded B300 run and provenance
5 RUN_METADATA.json human-readable run parameters
6 capture/ exact prompt plan and capture summary
7 inventories/ source, carrier, numeric, runtime seals
8 DEPLOYED_CODE.sha256 exact post-repair deployed-code hashes
9 lineage/encode_tr3_v31.py proven numeric core
10 r7_encoder/ exact deployed encoder package
11 run/ exact launch, supervisor, guard, assembly,
12 finisher, and progressive-upload scripts
13reproducibility/local-corrected-v1/
14 README.md authoritative corrected reproduction guide
15 STORAGE_BOUNDED_RUNBOOK.md NVMe paging, seals, and reclamation order
16 code/ exact executed local correction code
17 decisions/ frozen decisions for all 75 routed layers
18 receipts/ preflight, source windows, and 75 seals
19 tests/ local correction regression tests
20 results/ corrected five-run KLD record
21r7-experts-layer-*.safetensors progressively uploaded expert weights
22r7-experts-layer-*.json per-layer bit maps and provenance
Loader compatibility boundary
This schema-v2 checkpoint requires a mixed-R7-aware loader. The pinned r34 and
validated r17-derived runtimes described above implement that contract; stock
or generic loaders without mixed K3/K4/K5 R7 support do not. The portable r17
companion implements the required dedicated ABI but remains a release candidate
until its end-to-end replay is complete. The included converter refuses to
label an unsupported conversion as ready.
Josh Cartu for the MTP-78 recipe and
rank-sliced runtime work credited by the upstream MTP overlay.
Luke Alonso for the
B12X kernels and Blackwell
serving work on which these profiles depend.
Martin Vit and
yatesdr for the Infernal Invocation / RC2 image
engineering credited by the upstream MTP overlay.
Festr for isolating the shared-expert split-payload decode loss. This is a
Discord handle; no public profile is linked because one has not been verified.
The quantized weights inherit the license and use restrictions of the upstream
GLM-5.2 model. ExLlamaV3 is MIT-licensed; consult every linked upstream project
for its own license and terms. The proposed runtime bundle carries
THIRD_PARTY_NOTICES.md
and pinned Apache-2.0/MIT license copies alongside the redistributed patches.