Views
No views yet
thinkingmachines/Inkling-Small
for Apple Silicon (mlx-vlm). 276B total / 12B active sparse MoE (42 layers, 256 routed experts top-6 + 2 shared),
text + image + audio in, text out.ToPo-ToPo/Inkling-Small-mlx-4bit.models/inkling can load an official Inkling checkpoint through the
public loader, and the first that implements the MoE global_scale / gate.bias tensors. On 0.6.7 / 0.6.8
this repo will not load.1from mlx_vlm import load, generate
2model, processor = load("ToPo-ToPo/Inkling-Small-mlx-2bit")thinkingmachines/Inkling-Small (license: apache-2.0, bf16, 531.9 GB)mlx-vlm 0.6.9 — mlx_vlm.convert --hf-path thinkingmachines/Inkling-Small --mlx-path . -q --q-bits 2 --q-group-size 64pad_token / eos_token added to tokenizer_config.json
(the official TokenizersBackend config sets neither, so transformers raises on any padded call).
Both point at existing ids — the vocabulary is unchanged.Thinking effort level: system message (default 0.9). Control it
with the OpenAI-compatible reasoning_effort — "none" / "minimal" / "low" / "medium" / "high" /
"max", or a float in [0.0, 0.99]. "none" disables thinking entirely.mlx_vlm.server, note that Inkling wraps its answer in structural tokens
(<|message_model|>, <|content_text|>, <|end_message|>) which the server's fixed
_CONTENT_MARKERS list does not strip, and that its reasoning channel is
<|content_thinking|> … <|end_message|><|message_model|> rather than one of the built-in marker pairs.
Set MLX_VLM_THINKING_START_TOKEN / MLX_VLM_THINKING_END_TOKEN accordingly and strip the structural
tokens, or the reasoning and those markers end up in content.models/inkling did not implement the MoE mlp.global_scale (50 keys) and mlp.gate.bias
(40 keys) present in the official checkpoint, so those tensors were silently dropped. It also
shipped a translated config (renamed intermediate_size / dense_intermediate_size, etc.) that 0.6.9
rejects. If you pulled this repo before this date, re-download it.