Views
No views yet
Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs · LFM2.5 family
| Property | Value |
|---|---|
| Base | LiquidAI/LFM2.5-2.6B |
| Architecture | lfm2 — hybrid conv + full attention, 30 layers |
| Method | OptiQ mixed-precision, sensitivity-driven (bf16 reference) |
| On disk | 1.93 GB (bf16: 5.2 GB) |
| Context | 128k |
| Metric | Score |
|---|---|
| MMLU (5-shot, 969 samples) | 59.5% |
| GSM8K (1000 samples) | 38.1% |
| IFEval (full set, strict) | 33.1% |
| BFCL-V3 simple (200 calls) | 48.5% |
| HumanEval (164 problems, pass@1) | 18.9% |
| HashHop (long-context retrieval) | 13.0% |
| Capability Score (mean of 6) | 35.19 |
1pip install mlx-optiq
2optiq serve --model mlx-community/LFM2.5-2.6B-OptiQ-4bittemperature 0.1, top_k 50, repetition_penalty 1.1. Those ship in generation_config.json in this repo and optiq serve applies them automatically, so you get the publisher's settings without passing any flags.1from mlx_lm import load, generate
2
3model, tokenizer = load("mlx-community/LFM2.5-2.6B-OptiQ-4bit")
4prompt = tokenizer.apply_chat_template(
5 [{"role": "user", "content": "Extract the dates from this invoice."}],
6 add_generation_prompt=True, tokenize=False,
7)
8print(generate(model, tokenizer, prompt=prompt, max_tokens=512))