This is a sentence-transformers model: It maps sentences & paragraphs to a 768 dimensional dense vector space and was designed for semantic search. It has been trained on 215M (question, answer) pairs from diverse sources. For an introduction to semantic search, have a look at: SBERT.net - Semantic Search
1from sentence_transformers import SentenceTransformer, util
23query ="How many people live in London?"4docs =["Around 9 Million people live in London","London is known for its financial district"]56#Load the model7model = SentenceTransformer('sentence-transformers/multi-qa-distilbert-cos-v1')89#Encode query and documents10query_emb = model.encode(query)11doc_emb = model.encode(docs)1213#Compute dot score between query and all document embeddings14scores = util.dot_score(query_emb, doc_emb)[0].cpu().tolist()1516#Combine docs & scores17doc_score_pairs =list(zip(docs, scores))1819#Sort by decreasing score20doc_score_pairs =sorted(doc_score_pairs, key=lambda x: x[1], reverse=True)2122#Output passages & scores23for doc, score in doc_score_pairs:24print(score, doc)
Usage (HuggingFace Transformers)
Without sentence-transformers, you can use the model like this: First, you pass your input through the transformer model, then you have to apply the correct pooling-operation on-top of the contextualized word embeddings.
python
1from transformers import AutoTokenizer, AutoModel
2import torch
3import torch.nn.functional as F
45#Mean Pooling - Take average of all tokens6defmean_pooling(model_output, attention_mask):7 token_embeddings = model_output.last_hidden_state #First element of model_output contains all token embeddings8 input_mask_expanded = attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()9return torch.sum(token_embeddings * input_mask_expanded,1)/ torch.clamp(input_mask_expanded.sum(1),min=1e-9)101112#Encode text13defencode(texts):14# Tokenize sentences15 encoded_input = tokenizer(texts, padding=True, truncation=True, return_tensors='pt')1617# Compute token embeddings18with torch.no_grad():19 model_output = model(**encoded_input, return_dict=True)2021# Perform pooling22 embeddings = mean_pooling(model_output, encoded_input['attention_mask'])2324# Normalize embeddings25 embeddings = F.normalize(embeddings, p=2, dim=1)2627return embeddings
282930# Sentences we want sentence embeddings for31query ="How many people live in London?"32docs =["Around 9 Million people live in London","London is known for its financial district"]3334# Load model from HuggingFace Hub35tokenizer = AutoTokenizer.from_pretrained("sentence-transformers/multi-qa-distilbert-cos-v1")36model = AutoModel.from_pretrained("sentence-transformers/multi-qa-distilbert-cos-v1")3738#Encode query and docs39query_emb = encode(query)40doc_emb = encode(docs)4142#Compute dot score between query and all document embeddings43scores = torch.mm(query_emb, doc_emb.transpose(0,1))[0].cpu().tolist()4445#Combine docs & scores46doc_score_pairs =list(zip(docs, scores))4748#Sort by decreasing score49doc_score_pairs =sorted(doc_score_pairs, key=lambda x: x[1], reverse=True)5051#Output passages & scores52for doc, score in doc_score_pairs:53print(score, doc)
Technical Details
In the following some technical details how this model must be used:
Setting
Value
Dimensions
768
Produces normalized embeddings
Yes
Pooling-Method
Mean pooling
Suitable score functions
dot-product (util.dot_score), cosine-similarity (util.cos_sim), or euclidean distance
Note: When loaded with sentence-transformers, this model produces normalized embeddings with length 1. In that case, dot-product and cosine-similarity are equivalent. dot-product is preferred as it is faster. Euclidean distance is proportional to dot-product and can also be used.
Background
The project aims to train sentence embedding models on very large sentence level datasets using a self-supervised
contrastive learning objective. We use a contrastive learning objective: given a sentence from the pair, the model should predict which out of a set of randomly sampled other sentences, was actually paired with it in our dataset.
Our model is intented to be used for semantic search: It encodes queries / questions and text paragraphs in a dense vector space. It finds relevant documents for the given passages.
Note that there is a limit of 512 word pieces: Text longer than that will be truncated. Further note that the model was just trained on input text up to 250 word pieces. It might not work well for longer text.
Training procedure
The full training script is accessible in this current repository: train_script.py.
Pre-training
We use the pretrained distilbert-base-uncased model. Please refer to the model card for more detailed information about the pre-training procedure.
Training
We use the concatenation from multiple datasets to fine-tune our model. In total we have about 215M (question, answer) pairs.
We sampled each dataset given a weighted probability which configuration is detailed in the data_config.json file.
The model was trained with MultipleNegativesRankingLoss using Mean-pooling, cosine-similarity as similarity function, and a scale of 20.
Dataset
Number of training tuples
WikiAnswers Duplicate question pairs from WikiAnswers
77,427,422
PAQ Automatically generated (Question, Paragraph) pairs for each paragraph in Wikipedia
64,371,441
Stack Exchange (Title, Body) pairs from all StackExchanges
25,316,456
Stack Exchange (Title, Answer) pairs from all StackExchanges
21,396,559
MS MARCO Triplets (query, answer, hard_negative) for 500k queries from Bing search engine