Views
No views yet
--trust-remote-codeglm_moe_dsa.py (declared via model_file in config.json), a fixed runtime
for this architecture, and needs it:mlx_lm.generate --model pipenetwork/GLM-5.2-MLX-4bit --trust-remote-code --prompt "..." --max-tokens 300glm_moe_dsa builds a lightning indexer on all 78 layers, but GLM-5.2 ships indexer weights
on 21 (indexer_types: the other 57 "shared" layers reuse the previous full layer's top-k selection).
mlx_lm.load loads leniently and left those 57 indexers at random initialisation. Prompts up to 2048
tokens were unaffected (the indexer is bypassed below index_topk); beyond that, 57 of 78 layers attended
to keys chosen by random projections. The bundled runtime implements the schedule as the reference does
(plus fp32 indexer scores and router logits and the indexer LayerNorm epsilon); tiny-config parity against
transformers 5.16 is 4e-7 with the sparse path live, and a strict load of this checkpoint reports zero
missing and zero unexpected tensors. Details, tests and the GLM-5.3 builds made with it:
github.com/PipeNetwork/glm53-mlx. The weights are unchanged.glm_moe_dsa MoE (256 experts, DeepSeek-V3.2-style sparse attention) — quantized to 4-bit.| Variant | Notes |
|---|---|
| 8-bit | 8-bit · ~800GB · needs ~1TB RAM · integrity-checked |
| 6-bit | 6-bit · ~625GB · needs ~768GB RAM · integrity-checked |
| 5-bit | 5-bit · ~530GB · needs ~640GB RAM · integrity-checked |
| 4-bit (this repo) | 4-bit · ~430GB · tight on 512GB · smoke-tested |
| mixed | mixed · experts@3-bit / non-expert@6-bit · ~360GB · 512GB-fit · smoke-tested |
1pip install mlx-lm
2python -m mlx_lm generate --model pipenetwork/GLM-5.2-MLX-4bit --prompt "Hello" -m 256{"group_size": 64, "bits": 4, "mode": "affine"}.