Views
No views yet
cached-experts-v2; it includes the Mellum2 support from
add-mellum-2-support. These files are not compatible with upstream
llama.cpp and are not expected to become compatible with upstream llama.cpp
unless Mellum2 support is implemented there separately.<think>...</think> blocks before the final answer. See the original JetBrains
model card for full architecture details, training notes, evaluations, and
license information.| File | Description |
|---|---|
Mellum2-12B-A2.5B-Thinking.BF16.gguf | BF16 GGUF conversion made with the experimental Mellum2 branch |
Mellum2-12B-A2.5B-Thinking.Q8_0.gguf | Q8_0 quantization made with the experimental Mellum2 branch |
Mellum2-12B-A2.5B-Thinking.Q6_K.gguf | Q6_K quantization made with the experimental Mellum2 branch |
Mellum2-12B-A2.5B-Thinking.Q4_K_M.gguf | Q4_K_M quantization made with the experimental Mellum2 branch |
cached-experts-v2, which includes add-mellum-2-support. Do
Do not expect them to load with upstream llama.cpp.1llama-cli \
2 -m Mellum2-12B-A2.5B-Thinking.Q4_K_M.gguf \
3 -p "def quicksort(xs):"temperature=0). Result is Mellum2 minus Qwen2.5-Coder.
BigCodeBench was evaluated on the local Hard/Instruct subset with
--execution local --no_gt.| Benchmark | Qwen2.5-Coder-7B-Instruct Q8 | Mellum2-12B-A2.5B-Thinking Q8 | Result |
|---|---|---|---|
| HumanEval | 85.98% | 87.20% | +1.22% |
| HumanEval+ | 78.66% | 83.54% | +4.88% |
| MBPP | 83.07% | 84.39% | +1.32% |
| MBPP+ | 70.90% | 69.84% | -1.06% |
| MultiPL-E Java | 71.52% | 69.62% | -1.90% |
| MultiPL-E TypeScript | 82.39% | 77.36% | -5.03% |
| Kotlin HumanEval | 62.11% | 77.64% | +15.53% |
| SAFIM API | 4.19% | 15.48% | +11.29% |
| BigCodeBench Hard/Instruct | 19.59% | 25.00% | +5.41% |