Views
No views yet
| Metric | Original | Trimmed | Reduction |
|---|---|---|---|
| Vocabulary size | 262,144 tokens | 32,768 tokens | 87.50% |
| Model size | 307,581,696 params | 131,420,928 params | 57.27% |

| File | Type | Size |
|---|---|---|
embeddinggemma-rus-32768-Q8_0.gguf | Q8_0 (8-bit) | 136 MB |
embeddinggemma-rus-32768-F32.gguf | F32 (lossless reference) | 503 MB |
SHA256SUMS. Q8_0 is the recommended default; F32 is provided as a lossless reference equivalent to the source safetensors.tokenizer.json (Gemma SPM-style BPE) and no tokenizer.model (SentencePiece). The Gemma3 HF→GGUF converter only takes the SentencePiece path when tokenizer.model exists, so the model's tokenizer chkhsh (b847c511…) was registered as the gemma4 pre-type (get_vocab_base_pre()) in conversion/base.py. This maps the SPM-style BPE (normalizer →▁, BPE over whole text, ByteFallback on raw UTF-8) correctly.--sentence-transformers-dense-modules. EmbeddingGemma has 2_Dense/3_Dense projection layers; without this flag they are silently dropped and embeddings drift from the SentenceTransformers baseline.llama-embedding, L2-normalized, OpenAI-style JSON output) is included as smoke-embedding.json.llama-server (OpenAI-compatible embeddings endpoint):1llama-server \
2 -m embeddinggemma-rus-32768-Q8_0.gguf \
3 --embeddings --host 0.0.0.0 --port 80801curl http://localhost:8080/v1/embeddings \
2 -H "Content-Type: application/json" \
3 -d '{"input": "task: search result | query: тестовый русский запрос", "model": "embeddinggemma-rus-32768"}'llama-embedding:1llama-embedding \
2 -m embeddinggemma-rus-32768-Q8_0.gguf \
3 --embd-output-format json --embd-normalize 2 \
4 -p "task: search result | query: тестовый русский запрос"EmbeddingGemma is instruction-tuned; prefix inputs withtask: <task> | query: <text>(e.g.task: search result,task: title,task: web query).
@misc{vera2025embeddinggemmapowerfullightweighttext,
title={EmbeddingGemma: Powerful and Lightweight Text Representations},
author={Henrique Schechter Vera and Sahil Dua and Biao Zhang and Daniel Salz and Ryan Mullins and Sindhu Raghuram Panyam and Sara Smoot and Iftekhar Naim and Joe Zou and Feiyang Chen and Daniel Cer and Alice Lisak and Min Choi and Lucas Gonzalez and Omar Sanseviero and Glenn Cameron and Ian Ballantyne and Kat Black and Kaifeng Chen and Weiyi Wang and Zhe Li and Gus Martins and Jinhyuk Lee and Mark Sherwood and Juyeong Ji and Renjie Wu and Jingxiao Zheng and Jyotinder Singh and Abheesht Sharma and Divyashree Sreepathihalli and Aashi Jain and Adham Elarabawy and AJ Co and Andreas Doumanoglou and Babak Samari and Ben Hora and Brian Potetz and Dahun Kim and Enrique Alfonseca and Fedor Moiseev and Feng Han and Frank Palma Gomez and Gustavo Hernández Ábrego and Hesen Zhang and Hui Hui and Jay Han and Karan Gill and Ke Chen and Koert Chen and Madhuri Shanbhogue and Michael Boratko and Paul Suganthan and Sai Meher Karthik Duddu and Sandeep Mariserla and Setareh Ariafar and Shanfeng Zhang and Shijie Zhang and Simon Baumgartner and Sonam Goenka and Steve Qiu and Tanmaya Dabral and Trevor Walker and Vikram Rao and Waleed Khawaja and Wenlei Zhou and Xiaoqi Ren and Ye Xia and Yichang Chen and Yi-Ting Chen and Zhe Dong and Zhongli Ding and Francesco Visin and Gaël Liu and Jiageng Zhang and Kathleen Kenealy and Michelle Casbon and Ravin Kumar and Thomas Mesnard and Zach Gleicher and Cormac Brick and Olivier Lacombe and Adam Roberts and Qin Yin and Yunhsuan Sung and Raphael Hoffmann and Tris Warkentin and Armand Joulin and Tom Duerig and Mojtaba Seyedhosseini},
year={2025},
eprint={2509.20354},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2509.20354},
}@misc{hf_blogpost_trimming,
title={Introduction to Trimming},
author={Loïck BOURDOIS and Tom AARSEN and Bram VANROY and Christopher AKIKI and Woojun JUNG and Manuel ROMERO and Prithiv SAKTHI},
year={2026},
url={https://huggingface.co/blog/lbourdois/introduction-to-trimming},
}