Views
No views yet
| This model | FP16 baseline | |
|---|---|---|
| Decode tok/s (avg, long traces) | 99.85 | 31.96 |
| Peak memory (GB) | 2.872 | 8.82 |
| Disk size (MB) | 2151 | 8907 |
| Benchmark | This model | FP16 baseline | n |
|---|---|---|---|
| MATH-500 (math reasoning) | 73.3% (answered 27/30) | 86.7% (answered 28/30) | 30 |
| IFEval (instruction following) | 50.0% | 30.0% | 30 |
| GSM8K (math, accuracy) | 100.0% | 93.3% | 30 |
| HumanEval (code, pass@1) | 63.3% | 86.7% | 30 |
| MMLU (knowledge, accuracy) | 63.3% | 56.7% | 30 |
| Level | This model | FP16 baseline |
|---|---|---|
| level 1 | 83.3% | 83.3% |
| level 2 | 83.3% | 100.0% |
| level 3 | 66.7% | 66.7% |
| level 4 | 66.7% | 83.3% |
| level 5 | 66.7% | 100.0% |
| Context length | Decode tok/s |
|---|---|
| ~128 tokens | 100.2 |
| ~256 tokens | 101.6 |
| ~512 tokens | 102.8 |
| ~1024 tokens | 102.7 |
pip install mlx-lm1from mlx_lm import load, generate
2
3model, tokenizer = load("sahilchachra/tessera-4b-mxfp4-mlx")
4response = generate(model, tokenizer, prompt="Your prompt here", max_tokens=256, verbose=True)| Model | Variant |
|---|---|
| sahilchachra/tessera-4b-4bit-mlx | Affine int4 |
| sahilchachra/tessera-4b-8bit-mlx | Affine int8 |
| sahilchachra/tessera-4b-mxfp4-mlx | Block float MX FP4 ← this model |
| sahilchachra/tessera-4b-mxfp8-mlx | Block float MX FP8 |
| sahilchachra/tessera-4b-optiq-5bpw-mlx | OptiQ mixed-precision (target 5.0 bpw) |