mdbr-leaf-mt-asym is a high-performance text embedding model designed for classification, clustering, semantic sentence similarity and summarization tasks.
If you are looking to perform semantic search / information retrieval (e.g. for RAGs), please check out our mdbr-leaf-ir model, which is specifically trained for these tasks.
[!Note] Note: this model has been developed by the ML team of MongoDB Research. At the time of writing it is not used in any of MongoDB's commercial product or service offerings.
Technical Report
A technical report detailing our proposed LEAF training procedure is available here.
Highlights
State-of-the-Art Performance: mdbr-leaf-mt-asym achieves state-of-the-art results for compact embedding models, ranking #1 on the public MTEB v2 (Eng) leaderboard for models with ≤30M parameters.
Flexible Architecture Support: mdbr-leaf-mt-asym uses an asymmetric retrieval architecture enabling even greater retrieval results.
MRL and Quantization Support: embedding vectors generated by mdbr-leaf-mt-asym compress well when truncated (MRL) and can be stored using more efficient types like int8 and binary. See below for more information.
Benchmark Comparison
The table below shows the scores for mdbr-leaf-mt on the MTEB v2 (English) benchmark, compared to other retrieval models.
mdbr-leaf-mt ranks #1 on this benchmark for models with <30M parameters.
Model
Size
MTEB v2 (Eng)
OpenAI text-embedding-3-large
Unknown
66.43
OpenAI text-embedding-3-small
Unknown
64.56
mdbr-leaf-mt
23M
63.97
gte-small
33M
63.22
snowflake-arctic-embed-s
32M
61.59
e5-small-v2
33M
61.32
granite-embedding-small-english-r2
47M
61.07
all-MiniLM-L6-v2
22M
59.03
Quickstart
Sentence Transformers
python
1from sentence_transformers import SentenceTransformer
23# Load the model 4model = SentenceTransformer("MongoDB/mdbr-leaf-mt-asym")56# Example queries and documents7queries =[8"What is machine learning?",9"How does neural network training work?",10]1112documents =[13"Machine learning is a subset of artificial intelligence that focuses on algorithms that can learn from data.",14"Neural networks are trained through backpropagation, adjusting weights to minimize prediction errors.",15]1617# Encode queries and documents18query_embeddings = model.encode_query(queries)19document_embeddings = model.encode_document(documents)2021# Compute similarity scores22scores = model.similarity(query_embeddings, document_embeddings)2324# Print results25for i, query inenumerate(queries):26print(f"Query: {query}")27for j, doc inenumerate(documents):28print(f" Similarity: {scores[i, j]:.4f} | Document {j}: {doc[:80]}...")
See example output
Query: What is machine learning?
Similarity: 0.8483 | Document 0: Machine learning is a subset of artificial intelligence that focuses on algorith...
Similarity: 0.6805 | Document 1: Neural networks are trained through backpropagation, adjusting weights to minimi...
Query: How does neural network training work?
Similarity: 0.6050 | Document 0: Machine learning is a subset of artificial intelligence that focuses on algorith...
Similarity: 0.7689 | Document 1: Neural networks are trained through backpropagation, adjusting weights to minimi...
mdbr-leaf-mt is aligned to mxbai-embed-large-v1, the model it has been distilled from.
This enables flexible architectures in which, for example, documents are encoded using the larger model,
while queries can be encoded faster and more efficiently with the compact leaf model.
This usually outperforms the symmetric setup in which both queries and documents are encoded with leaf.
To use exclusively the leaf model, use mdbr-leaf-mt.
MRL Truncation
Embeddings have been trained via MRL and can be truncated for more efficient storage:
Vector quantization, for example to int8 or binary, can be performed as follows:
Note: For vector quantization to types other than binary, we suggest performing a calibration to determine the optimal ranges, see here.
Good initial values are -1.0 and +1.0.
python
1from sentence_transformers.quantization import quantize_embeddings
2import torch
34query_embeds = model.encode_query(queries)5doc_embeds = model.encode_document(documents)67# Quantize embeddings to int8 using -1.0 and +1.08ranges = torch.tensor([[-1.0],[+1.0]]).expand(2, query_embeds.shape[1]).cpu().numpy()9query_embeds = quantize_embeddings(query_embeds,"int8", ranges=ranges)10doc_embeds = quantize_embeddings(doc_embeds,"int8", ranges=ranges)1112# Calculate similarities; cast to int64 to avoid under/overflow13similarities = query_embeds.astype(int) @ doc_embeds.astype(int).T
1415print('After quantization:')16print(f"* Embeddings type: {query_embeds.dtype}")17print(f"* Similarities:\n{similarities}")