Required usage: pooling + instruction
Qwen3-Embedding uses last-token pooling and is instruction-aware — both are required for correct results:
- Serve with last-token pooling:
llama-server -m <model>.gguf --embedding --pooling last
- Prepend a task instruction to queries (not documents):
Instruct: {task description}\nQuery:{your query}
- Inputs should end with the EOS token
<|endoftext|> (recent llama.cpp appends this automatically).
Without last-token pooling and the instruction prefix, retrieval quality collapses. Instructions improve results ~1-5%. The MTEB score below (SciFact 0.6982 nDCG@10) was measured with both. Native dimension 1024, Matryoshka (MRL) custom dimensions supported, multilingual (100+ languages), 8192-token context.
Quant note: Q4_K_M drift dips to ~0.945 min — prefer Q5_K_M or Q8_0 for best fidelity.
Qwen3-Embedding-0.6B — Embedding GGUF (quantization-verified)
Quantized embedding model in GGUF, served in --embedding mode via llama.cpp. This is an encoder — it outputs vectors, not text. It is validated for retrieval quality and quantization fidelity, not chat behavior.
Files
Qwen3-Embedding-0.6B-Q4_K_M.gguf (396.5 MB)
Qwen3-Embedding-0.6B-Q5_K_M.gguf (444.2 MB)
Qwen3-Embedding-0.6B-Q8_0.gguf (639.2 MB)
Quantization drift (vs f16)
Mean cosine similarity of embeddings vs the f16 baseline. 1.0 = identical.
| Quant | Mean cosine | Min cosine | Verdict |
|---|
| Q4_K_M | 0.97469 | 0.94327 | good (>0.97) |
| Q5_K_M | 0.99117 | 0.97828 | excellent (>0.99) |
| Q8_0 | 0.99934 | 0.99887 | excellent (>0.99) |
Per-domain fidelity at Q4_K_M (which content types the quant preserves best):
| Domain | Mean cosine | Min |
|---|
| science | 0.96758 | 0.94327 |
| legal | 0.96894 | 0.96257 |
| long_form | 0.97152 | 0.97059 |
| code | 0.97431 | 0.96263 |
| everyday | 0.97454 | 0.9666 |
| medical | 0.9754 | 0.97154 |
| finance | 0.97994 | 0.97184 |
| short_queries | 0.98367 | 0.98154 |
Retrieval sanity (lightweight)
Built-in 12-query retrieval check (no external corpus): top-1 accuracy 1.0, MRR 1.0. healthy (top-1 >= 0.9)
Retrieval (MTEB)
Standardized
MTEB retrieval scores (main metric, usually nDCG@10 — higher is better). These are comparable across models on the MTEB leaderboard.
Metric: main_score (retrieval tasks: nDCG@10). Measured on the Q8_0 quant served via llama.cpp.
Dense-retrieval mode. These scores are for standard single-vector dense retrieval (what llama.cpp serves). Models like BGE-M3 that also support sparse/multi-vector (ColBERT) modes score higher in hybrid setups — that capability isn't exercised here, so compare this number against other models' dense scores, not hybrid ones.
What this is NOT
This card carries no safety, red-team, or viewpoint scores: those do not apply to an embedding model. For chat-model governance cards, see the SmartTasks text-LLM line.