Views
No views yet
| Size | |
|---|---|
| FP8/FP4 source | 167 GB |
| BF16 equivalent | ~570 GB |
| Config-I (3.05 bpw) | 108 GB (101 GiB) |
| Tensor group | Precision |
|---|---|
| Expert gate/up (routed + shared) | 2-bit |
| Expert down | 3-bit |
| MLA attention + indexer + compressor | 4-bit |
| Boundary layers (first 2 + last 2): attention | 8-bit |
| Boundary layers: experts | 4-bit |
| Embeddings + head | 8-bit |
| Router gates, norms, mHC params | f16 |
Note: this MLX build has had lighter testing than the GGUF sibling. If you want the more thoroughly measured artifact (perplexity, backend traps, behavioral notes), start there.
max_tokens truncates mid-thought and looks like a failure.deepseek_v4 is a registered model type — loads directly.deepseek_v4 model class (upstream PR pending). Until it lands, drop the included deepseek_v4.py into mlx_lm/models/.iogpu.wired_limit_mb helps sustained throughput.