8-bit (group 64) Qwen3.5-4B with a calibrated native-MTP draft head, built
for
MTPLX on Apple Silicon. 4.58 GB on disk, ~4.8 GiB
peak at load. The highest-fidelity 4B MTPLX ships, with the largest MTP
multiplier in the fleet.
The 8-bit trunk keeps output quality close to the BF16 reference while
the calibrated draft head converts that fidelity into a 2.2x decode
multiplier. The engine reads the tuned depth from mtplx_runtime.json.
Runs on any Apple Silicon Mac with 8 GB+ of unified memory.
New artifact (July 2026), forged with the fixed MTPLX forge after the 4B
zero-acceptance defect
(
#176) was root-caused:
draft-head RMSNorms in the original export are stored zero-centered and
must be restored at extraction. The draft head is quantized int4 (group
64) with fc and norms kept in BF16, calibrated so acceptance matches the
BF16 head.