Views
No views yet
/api/embed /
llama-server --embedding endpoints).| File | Precision | Size | Note |
|---|---|---|---|
islamqa-retriever-quran-f16.gguf | F16 | 287 MB | recommended — max cosine deviation vs original: 0.0003 |
islamqa-retriever-quran-q8_0.gguf | Q8_0 | 154 MB | smaller; max cosine deviation: 0.005 |
modern-bert (requires a llama.cpp runtime from Dec 2025 or newer)"the current context does not logits computation"
(HTTP 500): you sent a chat message to an embedding model. This model cannot generate
text anywhere, by design. Use it as the embedding model for RAG / document search, or via
POST /v1/embeddings.