Views
No views yet
google/gemma-4-26B-A4B-it, for Apple Silicon (~26 GB).
Sparse-MoE vision-language model — run it with mlx-vlm, not mlx-lm. The expert router (router.proj) is kept at 8-bit precision.pip install -U mlx-vlm1python -m mlx_vlm.generate \
2 --model TyKaoz/gemma-4-26B-A4B-it-8bit \
3 --prompt "Explique la quantization en une phrase." \
4 --max-tokens 200| Base | Tool | Precision | Size |
|---|---|---|---|
google/gemma-4-26B-A4B-it | mlx-vlm | 8-bit · group 64 (router 8-bit) | ~26 GB |