4-bit (group 64) Qwen3.5-4B with a calibrated native-MTP draft head, built
for
MTPLX on Apple Silicon. 2.47 GB on disk, ~2.9 GiB
peak at load. The fastest model MTPLX ships.
Earlier revisions of this repo shipped a defective MTP sidecar: the
RMSNorm weights were stored in the raw zero-centered convention and never
restored, so the draft head proposed garbage and MTP made the model
slower than plain decoding (reported as
#176, thanks lBroth).
This revision is rebuilt from source with the fixed forge. MTPLX 2.2.0+
also detects and heals the old sidecar at load, so existing downloads
recover without re-downloading.