Views
No views yet
1pip install mlx-lm
2
3mlx_lm.generate \
4 --model mlx-community/Youtu-LLM-2B-4bit \
5 --prompt "Hello, what can you do?" \
6 --max-tokens 100| Quant | Prompt | Generation | Memory |
|---|---|---|---|
| bf16 | 118 tok/s | 112 tok/s | 4.7GB |
| 4-bit | 202 tok/s | 205 tok/s | 1.3GB |
<think> tags for Chain of Thought| Benchmark | Youtu-LLM-2B | Qwen3-4B |
|---|---|---|
| HumanEval | 95.9% | 95.4% |
| SWE-Bench | 17.7% | 5.7% |
| GAIA | 33.9% | 25.5% |