Views
No views yet
| This model | FP16 baseline | |
|---|---|---|
| Decode tok/s (avg, long traces) | 57.57 | 31.96 |
| Peak memory (GB) | 4.885 | 8.82 |
| Disk size (MB) | 4156 | 8907 |
| Benchmark | This model | FP16 baseline | n |
|---|---|---|---|
| MATH-500 (math reasoning) | 83.3% (answered 27/30) | 86.7% (answered 28/30) | 30 |
| IFEval (instruction following) | 40.0% | 30.0% | 30 |
| GSM8K (math, accuracy) | 96.7% | 93.3% | 30 |
| HumanEval (code, pass@1) | 93.3% | 86.7% | 30 |
| MMLU (knowledge, accuracy) | 56.7% | 56.7% | 30 |
| Level | This model | FP16 baseline |
|---|---|---|
| level 1 | 83.3% | 83.3% |
| level 2 | 100.0% | 100.0% |
| level 3 | 83.3% | 66.7% |
| level 4 | 66.7% | 83.3% |
| level 5 | 83.3% | 100.0% |
| Context length | Decode tok/s |
|---|---|
| ~128 tokens | 58.6 |
| ~256 tokens | 58.5 |
| ~512 tokens | 58.5 |
| ~1024 tokens | 58.4 |
pip install mlx-lm1from mlx_lm import load, generate
2
3model, tokenizer = load("sahilchachra/tessera-4b-mxfp8-mlx")
4response = generate(model, tokenizer, prompt="Your prompt here", max_tokens=256, verbose=True)| Model | Variant |
|---|---|
| sahilchachra/tessera-4b-4bit-mlx | Affine int4 |
| sahilchachra/tessera-4b-8bit-mlx | Affine int8 |
| sahilchachra/tessera-4b-mxfp4-mlx | Block float MX FP4 |
| sahilchachra/tessera-4b-mxfp8-mlx | Block float MX FP8 ← this model |
| sahilchachra/tessera-4b-optiq-5bpw-mlx | OptiQ mixed-precision (target 5.0 bpw) |