Views
No views yet

Qwen/Qwen3-VL-Embedding-2B prepared for direct vLLM deployment through the modelopt_fp4 backend.NVFP4_DEFAULT_CFGhf_quant_config.jsonmodelopt_fp4NVFP416model.visual* and lm_headmteb/MSMARCO-PL, mteb/NQ-PL, mteb/FiQA-PLen, de, es, fr, javidore/colpali_train_set and lmms-lab/flickr30kmteb/MSMARCO-PLmteb/NQ-PLen, de, es, fr, javidore/vidore_v3_industrialvidore/vidore_v3_computer_sciencenDCG@10Recall@10MRR@10Qwen/Qwen3-VL-Embedding-2B checkpoint on the local full benchmark:| Metric | Stock FP16 | NVFP4 | Delta |
|---|---|---|---|
nDCG@10 | 0.56222 | 0.55008 | -0.01214 |
Recall@10 | 0.64934 | 0.63794 | -0.01141 |
MRR@10 | 0.78883 | 0.77870 | -0.01013 |
| Benchmark wall time | 434.853 s | 377.707 s | 13.14% faster |
| Average request latency | 0.332726 s | 0.277620 s | -0.055106 s |
| Throughput | 18.4338 rps | 21.2228 rps | +2.7890 rps |
0.9172 against the FP16 checkpoint.1HF_TOKEN=hf_xxx \
2vllm serve LifetimeMistake/Qwen3-VL-Embedding-2B-NVFP4 \
3 --runner pooling \
4 --convert embed \
5 --trust-remote-code \
6 --quantization modelopt_fp4 \
7 --limit-mm-per-prompt '{"image":1}'chat_template.jinja, download the repo locally and pass --chat-template /path/to/chat_template.jinja.Qwen/Qwen3-VL-Embedding-2B, the 2B member of Qwen’s multimodal embedding series.2048, with support for smaller embedding dimensions