Views
No views yet
num_nextn_predict_layers=1, num_experts=192), for self-speculative decoding on MTPLX.mtplx inspect (release v2.0.1) recognizes it — "HY V3 MTP markers recognized, but MTPLX does not yet have a native MLX runtime backend for this family" — i.e. it is recognized-backend-pending. The native hy_v3 runtime backend is in flight as MTPLX PR #142, gated on hy_v3 reaching MTPLX's mlx-lm pin. Until that lands, use the AR daily-driver (lite-v1) instead.num_experts. So the base checkpoint's mtp.* sidecar grafts directly onto the fused AR trunk — no re-heal, no re-prune of the trunk.mtplx inspect receipts: https://github.com/PhilipJohnBasile/hy3-demolition-mlx (scripts/38_mtp_sidecar_graft.py, eval/receipts/mtplx_inspect_*.json).mtplx inspect), not a live run.hy_v3 must reach mainline mlx-lm (ml-explore/mlx-lm#1211). And note: even once loadable, LM Studio runs AR-only (no MTP speculative decoding) — the MTP heads only pay off on MTPLX. For LM Studio, use the AR sibling instead.src/hy3_streaming.py in the source repo, bit-identical,
zero quality loss). The MTP head is small and stays resident; only trunk experts
stream. See docs/64gb-feasibility.md for the measured 16/32/64 GB tiers.