Views
No views yet
Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs
| Property | Value |
|---|---|
| Base | LiquidAI/LFM2.5-1.2B-Thinking |
| Architecture | lfm2 — hybrid conv + full attention |
| Method | OptiQ mixed-precision, sensitivity-driven (bf16 reference) |
| On disk | 820 MB (bf16: 2.34 GB) |
| Context | 128k |
| KV cache | mixed-precision kv_config.json bundled (6 attention layers, 4.67 avg bits) |
| Metric | Score |
|---|---|
| MMLU (5-shot, 969 samples) | 64.7% |
| GSM8K (1000 samples) | 82.8% |
| IFEval (full set, strict) | 84.3% |
| BFCL-V3 simple (200 calls) | 45.5% |
| HumanEval (164 problems, pass@1) | 47.6% |
| HashHop (long-context retrieval) | 0.0% |
| Capability Score (mean of 6) | 54.14 |
pip install mlx-optiq1from mlx_lm import load, generate
2
3model, tok = load("mlx-community/LFM2.5-1.2B-Thinking-OptiQ-4bit")
4prompt = tok.apply_chat_template(
5 [{"role": "user", "content": "What is 17 * 23? Think briefly then answer."}],
6 tokenize=False, add_generation_prompt=True,
7)
8print(generate(model, tok, prompt=prompt, max_tokens=2048))max_tokens. Cap it too low and the whole allowance is consumed inside the reasoning block, so you get an empty answer back rather than a short one. 2048 is a sensible floor for anything non-trivial.optiq serve --model mlx-community/LFM2.5-1.2B-Thinking-OptiQ-4bit --kv-config kv_config.json<|tool_call_start|>[get_weather(city="Paris")]<|tool_call_end|>tools= to apply_chat_template and they are rendered into the system prompt.