Views
No views yet
[!NOTE] 🚀 Try our demo: https://huggingface.co/spaces/Tarka-AIR/Tarka-Embedding
| MTEB English / Models | Param. | Mean(Task) | Mean(Type) | Class. | Clust. | Pair Class. | Rerank. | Retri. | STS | Summ. |
|---|---|---|---|---|---|---|---|---|---|---|
| multilingual-e5-large-instruct | 0.6B | 65.53 | 61.21 | 75.54 | 49.89 | 86.24 | 48.74 | 53.47 | 84.72 | 29.89 |
| NV-Embed-v2 | 7.8B | 69.81 | 65.00 | 87.19 | 47.66 | 88.69 | 49.61 | 62.84 | 83.82 | 35.21 |
| GritLM-7B | 7.2B | 67.07 | 63.22 | 81.25 | 50.82 | 87.29 | 49.59 | 54.95 | 83.03 | 35.65 |
| gte-Qwen2-1.5B-instruct | 1.5B | 67.20 | 63.26 | 85.84 | 53.54 | 87.52 | 49.25 | 50.25 | 82.51 | 33.94 |
| stella_en_1.5B_v5 | 1.5B | 69.43 | 65.32 | 89.38 | 57.06 | 88.02 | 50.19 | 52.42 | 83.27 | 36.91 |
| gte-Qwen2-7B-instruct | 7.6B | 70.72 | 65.77 | 88.52 | 58.97 | 85.9 | 50.47 | 58.09 | 82.69 | 35.74 |
| gemini-embedding-exp-03-07 | - | 73.3 | 67.67 | 90.05 | 59.39 | 87.7 | 48.59 | 64.35 | 85.29 | 38.28 |
| Tarka-Embedding-350M-V1 | 350M | 69.29 | 63.29 | 88.43 | 55.73 | 83.96 | 47.77 | 55.14 | 84.59 | 27.43 |
[!IMPORTANT]
For the best performance use Flash attention with bfloat16
1from sentence_transformers import SentenceTransformer
2
3# We recommend enabling flash_attention_2 for better acceleration and memory saving,
4model = SentenceTransformer(
5 "Tarka-AIR/Tarka-Embedding-350M-V1",
6 trust_remote_code=True,
7 model_kwargs={
8 "attn_implementation": "flash_attention_2",
9 "device_map": "cuda",
10 "torch_dtype": "bfloat16",
11 },
12 tokenizer_kwargs={"padding_side": "left"},
13)
14
15# The queries and documents to embed
16queries = [
17 "What is the capital of China?",
18 "Explain gravity",
19]
20documents = [
21 "The capital of China is Beijing.",
22 "Gravity is a force that attracts two bodies towards each other. It gives weight to physical objects and is responsible for the movement of planets around the sun.",
23]
24
25# Encode the queries and documents. Note that queries benefit from using a prompt
26# Here we use the prompt called "query" stored under `model.prompts`, but you can
27# also pass your own prompt via the `prompt` argument
28query_embeddings = model.encode(queries, prompt_name="query")
29document_embeddings = model.encode(documents)
30
31# Compute the (cosine) similarity between the query and document embeddings
32similarity = model.similarity(query_embeddings, document_embeddings)
33print(similarity)
34
35# tensor([[0.9177, 0.3923],
36# [0.2975, 0.7631]])Disclaimer The creator of this Model is not responsible for any misuse, damages, or legal issues arising from the use of this model.
1@misc{tarka_ai_research_2025,
2 author = { Tarka AI Research },
3 title = { Tarka-Embedding-350M-V1 (Revision f4b5de8) },
4 year = 2025,
5 url = { https://huggingface.co/Tarka-AIR/Tarka-Embedding-350M-V1 },
6 doi = { 10.57967/hf/6979 },
7 publisher = { Hugging Face }
8}