Views
No views yet
| This model | FP16 baseline | |
|---|---|---|
| Decode tok/s (avg, long traces) | 633.26 | 203.66 |
| Peak memory (GB) | 1.126 | 2.732 |
| Disk size (MB) | 598 | 2237 |
| Benchmark | This model | FP16 baseline | n |
|---|---|---|---|
| JMMLU (Japanese MMLU, accuracy) | 33.3% | 38.3% | 60 |
| MGSM-ja (Japanese math, accuracy) | 20.0% | 33.3% | 30 |
| JCommonsenseQA (Japanese commonsense, accuracy) | 75.0% | 77.5% | 40 |
| Long-context JA (length-met rate) | 100.0% | 100.0% | 8 |
| Context length | Decode tok/s |
|---|---|
| ~128 tokens | 371.9 |
| ~256 tokens | 364.4 |
| ~512 tokens | 364.6 |
| ~1024 tokens | 361.0 |
pip install mlx-lm1from mlx_lm import load, generate
2
3model, tokenizer = load("sahilchachra/lfm2.5-1.2b-jp-mxfp4-mlx")
4response = generate(model, tokenizer, prompt="Your prompt here", max_tokens=256, verbose=True)| Model | Variant |
|---|---|
| sahilchachra/lfm2.5-1.2b-jp-mxfp4-mlx | Block float MX FP4 ← this model |
| sahilchachra/lfm2.5-1.2b-jp-optiq-5bpw-mlx | OptiQ mixed-precision (target 5.0 bpw) |