This model was trained on the
MS Marco Passage Ranking task.
The model can be used for Information Retrieval: Given a query, encode the query will all possible passages (e.g. retrieved with ElasticSearch). Then sort the passages in a decreasing order. See
SBERT.net Retrieve & Re-rank for more details. The training code is available here:
SBERT.net Training MS Marco
The usage is easy when you have
SentenceTransformers installed. Then you can use the pre-trained models like this:
1from sentence_transformers import CrossEncoder
2
3model = CrossEncoder('cross-encoder/ms-marco-MiniLM-L6-v2')
4scores = model.predict([
5 ("How many people live in Berlin?", "Berlin had a population of 3,520,031 registered inhabitants in an area of 891.82 square kilometers."),
6 ("How many people live in Berlin?", "Berlin is well known for its museums."),
7])
8print(scores)
9# [ 8.607138 -4.320078]
1from transformers import AutoTokenizer, AutoModelForSequenceClassification
2import torch
3
4model = AutoModelForSequenceClassification.from_pretrained('cross-encoder/ms-marco-MiniLM-L6-v2')
5tokenizer = AutoTokenizer.from_pretrained('cross-encoder/ms-marco-MiniLM-L6-v2')
6
7features = tokenizer(['How many people live in Berlin?', 'How many people live in Berlin?'], ['Berlin has a population of 3,520,031 registered inhabitants in an area of 891.82 square kilometers.', 'New York City is famous for the Metropolitan Museum of Art.'], padding=True, truncation=True, return_tensors="pt")
8
9model.eval()
10with torch.no_grad():
11 scores = model(**features).logits
12 print(scores)
In the following table, we provide various pre-trained Cross-Encoders together with their performance on the
TREC Deep Learning 2019 and the
MS Marco Passage Reranking dataset.
Note: Runtime was computed on a V100 GPU.