Views
No views yet
Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon (no PyTorch, no cloud). Try the Lab · All OptiQ quants · Docs
static build — per-layer bit-widths assigned to a 2.5
target bits-per-weight.optiq >= 0.4.4, which ships the vendored llada2_moe decoder (the
256-expert diffusion MoE) and the block-diffusion decode loop. Stock mlx-lm
has no llada2_moe arch and cannot load or generate from this repo.pip install -U optiqoptiq serve detects the arch and routes it through OptiQ's
vendored decoder and the block-diffusion decode loop automatically:optiq serve --model mlx-community/LLaDA2.2-flash-OptiQ-2bithttp://localhost:8000/v1, or use
it from the OptiQ Lab and optiq code.