Views
No views yet
JetBrains/Mellum2-12B-A2.5B-Instruct,
the instruction-tuned Mixture-of-Experts coding assistant from JetBrains. It is derived from the
full-precision jedisct1/Mellum2-12B-A2.5B-Instruct-mlx
conversion.<think>
reasoning block. Mellum 2 uses 64 experts with 8 active per token (about 2.5B active parameters
out of 12B), a mix of sliding-window and full-attention layers, and a 131,072-token context
window.q_proj, k_proj, v_proj, o_proj) — 8-bitlm_head — 8-bitmlp.gate) — 8-bitswitch_mlp) — 4-bitmlx_lm.server driven by the swival agent
harness, run side by side with the full-precision model. Across repeated multi-step coding tasks
the 4-bit weights matched the full model's behavior: well-formed read_file, edit_file,
write_file, list_files, and shell-command calls, and no malformed tool calls. Generation
stops cleanly on <|im_end|> (the eos_token_id is set to [0, 28], which is what lets agent
harnesses see a proper tool_calls finish reason).mellum architecture is not supported by the stock mlx-lm code yet.mlx-lm from source:pip install git+https://github.com/jedisct1/mlx-lmuv:uvx --from git+https://github.com/jedisct1/mlx-lm mlx_lm.server1uvx --from git+https://github.com/jedisct1/mlx-lm \
2 mlx_lm.generate --model jedisct1/Mellum2-12B-A2.5B-Instruct-mlx-4bit \
3 --prompt "Write a Python function that reverses a linked list." \
4 --max-tokens 16384 \
5 --temp 0.6 --top-p 0.95 --top-k 201uvx --from git+https://github.com/jedisct1/mlx-lm \
2 mlx_lm.server --model jedisct1/Mellum2-12B-A2.5B-Instruct-mlx-4bit \
3 --max-tokens 16384 \
4 --temp 0.6 --top-p 0.95 --top-k 20temperature=0.6, top_p=0.95, top_k=20.uv tool install swivalswival --provider llamacpp --model jedisct1/Mellum2-12B-A2.5B-Instruct-mlx-4bit