Views
No views yet
| Metric | Score |
|---|---|
| Real-World Tasks | 16.7% (2/12) |
| Training Tasks | 52.3% |
| Held-Out Tasks | 60.0% |
| Quant | Quality | Size | Recommendation |
|---|---|---|---|
| f16 | Full | ~2.1GB | Best quality |
| q8_0 | Excellent | ~1.1GB | Recommended — near-identical to f16 |
| q5_k_m | Good | ~0.8GB | Reasoning OK, response may degrade on longer outputs |
| q4_k_m | Fair | ~0.7GB | Reasoning OK, response degrades into repetition |
| q3_k_m | Poor | ~0.6GB | Not recommended |
| q2_k | Poor | ~0.5GB | Not recommended |
enable_thinking=true, so the model will always produce reasoning followed by response. If your inference engine supports enable_thinking=false, you can skip reasoning for faster responses.