MLX
static mixed-precision 4-bit quant of
unsloth/Qwen3.8-27B.
Measured weights footprint:
13.33 GB.
Recommended sampling from a per-model temperature ladder: temperature
0.4 (this recipe
failed convergence screens at t0.6; a capped scan located t0.4), top_p 0.95, top_k 20,
min_p 0.0, presence_penalty 0.0, thinking ON (budget 81920). Campaign methodology and
results:
https://github.com/ivan-avramov/mlx_local_stack.
The original conversion was language-model-only. This revision grafts the vision tower back
from the upstream base (unsloth repackaging of the family release): 333 vision_tower.*
tensors kept bf16 (exactly what the vision-retaining mlx_vlm convert produces for this
family), +0.92 GB.