Views
No views yet
JetBrains/Mellum2-12B-A2.5B-Instruct, ready to run with
llama.cpp, Ollama, LM Studio, and
other GGUF-compatible runtimes.| File | Size |
|---|---|
Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf | 8.1 GB |
| Quantization | Description | Size | KLD vs BF16 ↓ | Top-token match ↑ |
|---|---|---|---|---|
BF16 | 16-bit, no quantization (reference) | 24.3 GB | — | — |
Q8_0 | 8-bit, effectively lossless | 12.9 GB | 0.016 | 95.2% |
Q6_K | 6-bit k-quant, very high quality | 10.9 GB | 0.038 | 92.9% |
Q4_K_M (this repo) | 4-bit k-quant, balanced (recommended) | 8.1 GB | 0.106 | 87.2% |
MXFP4_MOE | MXFP4 4-bit on MoE experts, smallest | 7.0 GB | 0.166 | 84.2% |
n_ctx=512); lower KLD / higher agreement means closer to the
unquantized model. (Perplexity is omitted here — it is unreliable for
instruction-tuned models on Wikitext-2, which is out of distribution.)hf download JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf --local-dir .1# Pull and serve in one step (downloads the GGUF automatically)
2llama-server -hf JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M \
3 --ctx-size 131072 \
4 --temp 0.6 --top-p 0.95 --top-k 20
5
6# Or run a one-off prompt with a local file
7llama-cli -m Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf \
8 --ctx-size 131072 \
9 --temp 0.6 --top-p 0.95 --top-k 20 \
10 -p "Write a Python function to reverse a string."http://localhost:8080/v1:1from openai import OpenAI
2
3client = OpenAI(base_url="http://localhost:8080/v1", api_key="llama.cpp")
4
5chat_response = client.chat.completions.create(
6 model="JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M",
7 messages=[
8 {"role": "user", "content": "Write a Python function to reverse a string."},
9 ],
10 max_tokens=81920,
11 temperature=0.6,
12 top_p=0.95,
13 extra_body={"top_k": 20},
14)
15print(chat_response.choices[0].message.content)ollama run hf.co/JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M