Views
No views yet
Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs · LFM2.5 family
| Property | Value |
|---|---|
| Base | LiquidAI/LFM2.5-8B-A1B |
| Architecture | lfm2_moe — hybrid conv + attention, 24 layers, 32 experts, 4 active |
| Method | OptiQ mixed-precision, sensitivity-driven (uniform-4-bit reference) |
| On disk | 5.46 GB (bf16: 15.8 GB) |
| Context | 128k |
| Metric | Score |
|---|---|
| MMLU (5-shot, 969 samples) | 46.1% |
| GSM8K (1000 samples) | 54.6% |
| IFEval (full set, strict) | 32.9% |
| BFCL-V3 simple (200 calls) | 24.5% |
| HumanEval (164 problems, pass@1) | 14.0% |
| HashHop (long-context retrieval) | 0.0% |
| Capability Score (mean of 6) | 28.69 |
1pip install mlx-optiq
2optiq serve --model mlx-community/LFM2.5-8B-A1B-OptiQ-4bittemperature 0.1, top_k 50, repetition_penalty 1.1. Those ship in generation_config.json in this repo and optiq serve applies them automatically, so you get the publisher's settings without passing any flags.1from mlx_lm import load, generate
2
3model, tokenizer = load("mlx-community/LFM2.5-8B-A1B-OptiQ-4bit")
4prompt = tokenizer.apply_chat_template(
5 [{"role": "user", "content": "Extract every date from this email."}],
6 add_generation_prompt=True, tokenize=False,
7)
8print(generate(model, tokenizer, prompt=prompt, max_tokens=512))