Views
No views yet
JetBrains/Mellum2-12B-A2.5B-Thinking,
the reasoning-augmented Mixture-of-Experts coding assistant from JetBrains. It is derived from the
full-precision jedisct1/Mellum2-12B-A2.5B-Thinking-mlx
conversion.<think>...</think> blocks before the final answer.edit-style tool calls whose JSON arguments contained mis-escaped quotes, so the server
dropped them as malformed and agent runs stalled. The culprit was 4-bit precision on the
attention and output layers, which is where the model copies literal source text verbatim into a
tool argument.q_proj, k_proj, v_proj, o_proj) — 8-bitlm_head — 8-bitmlp.gate) — 8-bitswitch_mlp) — 4-bitmlx_lm.server driven by the swival agent
harness: across repeated runs the model issued well-formed read_file, edit_file,
write_file, list_files, and shell-command calls, completed multi-step coding tasks, and never
produced a malformed tool call. Generation stops cleanly on <|im_end|> (the eos_token_id is
set to [0, 28], which is what lets agent harnesses see a proper tool_calls finish reason).mellum architecture is not supported by the stock mlx-lm code yet.mlx-lm from source:pip install git+https://github.com/jedisct1/mlx-lmuv:uvx --from git+https://github.com/jedisct1/mlx-lm mlx_lm.server1uvx --from git+https://github.com/jedisct1/mlx-lm \
2 mlx_lm.generate --model jedisct1/Mellum2-12B-A2.5B-Thinking-mlx-4bit \
3 --prompt "How are you doing today?" \
4 --max-tokens 16384 \
5 --temp 0.6 --top-p 0.95 --top-k 201uvx --from git+https://github.com/jedisct1/mlx-lm \
2 mlx_lm.server --model jedisct1/Mellum2-12B-A2.5B-Thinking-mlx-4bit \
3 --max-tokens 16384 \
4 --temp 0.6 --top-p 0.95 --top-k 20temperature=0.6, top_p=0.95, top_k=20.uv tool install swivalswival --provider llamacpp --model jedisct1/Mellum2-12B-A2.5B-Thinking-mlx-4bit