Quantized with llama.cpp using a public Wikitext-2 importance matrix.
Queries require the query: prefix and documents require passage: .
Embeddings use mean pooling and L2 normalization.
This model is not affiliated with or endorsed by nvidia, mistral, or any other company. These are simply quantized gguf files of the original model.