Source checkpoint is jinaai/jina-embeddings-v4-vllm-text-matching — a stock Qwen2.5-VL-3B
with the text-matching task LoRA merged into the base weights (no custom adapter code, single
full checkpoint). It is one of three per-task variants of jina-embeddings-v4; this repo hosts the
ONNX conversion of the text-matching one.
text-matching is the symmetric similarity task: both sides of a pair use the same Query:
prefix (unlike retrieval/code, which are asymmetric Query: / Passage:). Use it for
sentence-similarity, STS, deduplication, and clustering.
This is a single-vector model: the embedding is a masked mean-pool over the last hidden state,
L2-normalized, 2048-d, with Matryoshka truncation to 128/256/512/1024/2048.
Decomposition
The model is split into three ONNX sub-parts (same pattern as the other recipes in this repo), so
the heavy backbone is stored once and reused by both the text and image paths:
Sub-part
Input → Output
Notes
vision.onnx
pixel_values [N,1176] → image_features [N,2048]
task-agnostic (vision tower has no LoRA); grid baked at build resolution
Pooling and Matryoshka truncation happen in the driver (nothing baked), so one build serves every
output dimension.
Why host-computed position_ids?
The model uses MROPE (mrope_section [16,24,24]), which onnxruntime-genai's ModelBuilder cannot
emit. position_ids [3,B,S] are therefore computed on the host and fed in: cumulative positions for
text, and get_rope_index(...) over the image grid for image inputs. The graph stays clean.
Prompts
Both sides of a comparison use the sameQuery: prefix (symmetric). The image prompt is the
fixed template used across all tasks. manifest.json records this per build under prompts
({"query": "Query:", "document": "Query:", "symmetric": true}).
text : "Query: <your text>"
image: "<|im_start|>user\n<|vision_start|><|image_pad|><|vision_end|>Describe the image.<|im_end|>\n"
Files
Each build directory is self-contained (sub-parts + image_meta.npz + tokenizer/processor assets +
manifest.json, which records the source hf_model id, precision, and any quantized sub-parts):
Dir
Precision
total
fp16
fp16
7.0 GB
fp32
fp32
14 GB
int8
fp16, backbone int8
~4.6 GB
int4
fp16, backbone int4
3.3 GB
int8 (fp16 graph + int8 backbone) is the recommended quantized option — smallest build that still
clears the 0.999 fidelity bar.
cuda_* directories, if present and empty, are placeholders. This environment's PyTorch/ORT are
CPU builds, so GPU builds produce nothing there. The exported ONNX is execution-provider agnostic —
the same files run on CUDAExecutionProvider via onnxruntime-gpu with no rebuild and no
device flag.
Fidelity vs full PyTorch
Composed ONNX chain vs the full Qwen2_5_VLForConditionalGeneration (pooled-embedding cosine,
worst of 3 text samples + 1 image):
Build
worst cosine
verdict
fp32
0.999984
✅
fp16
0.999987
✅
int8 (backbone)
0.999436
✅ (≥0.999)
int4 (backbone)
0.912048
❌ not for production
(Reference model loaded in fp16 for the eval; eval.py dedupes repeated paths, so listing a dir
twice runs it once.)
int8 is the quantization sweet spot — ~35 % smaller than fp16 with negligible cosine drift. int4
is too coarse for an embedding model (the pooled/normalized vector amplifies 4-bit weight error into
~8 % drift, which wrecks similarity ranking) — build.py emits it with a warning, not a failure.
Reproducing / using
CPU only; runs in the repo's uv project env (transformers 5.x, torchvision for the image
processor). The pipeline is three task-agnostic scripts sharing common.py — point --model at the
text-matching source:
bash
1# build sub-parts — one --precision flag: fp16 (default) | fp32 | int8 | int42uv run build.py --model vllm-text-matching --output onnx/fp16 # fp163uv run build.py --model vllm-text-matching --output onnx/fp32 --precision fp32
4uv run build.py --model vllm-text-matching --output onnx/int8 --precision int8 # fp16 graph + int8 backbone5uv run build.py --model vllm-text-matching --output onnx/int4 --precision int4 # lossy (see above)67# eval — accepts multiple build dirs (positional), dedupes repeats, auto-detects each one's precision8uv run eval.py --model vllm-text-matching onnx/fp16 onnx/fp32 onnx/int8 onnx/int4
910# inference (no PyTorch load) — text-matching uses the Query: prefix for BOTH texts11uv run inference.py --onnx-dir onnx/fp16 --text "The impacts of climate change on coastal cities"12uv run inference.py --onnx-dir onnx/fp16 --text "..." --truncate-dim 25613uv run inference.py --onnx-dir onnx/fp16 --image doc.png
int8/int4 build the fp16 graph then weight-quantize the backbone in place (block-wise
MatMulNBits; vision/embeddings stay fp16). build.py runs a composed self-sanity check: fp16 /
fp32 / int8 must hit cosine ≥ 0.999 or the build fails, while int4 only warns. The vision sub-part
is identical across all tasks (no LoRA), so a single vision.onnx can be shared to save disk.
Same scripts serve the other tasks — --model vllm-retrieval or --model vllm-text-code — the
only difference being the prompt convention (those are asymmetric Query: / Passage:).
Minimal ONNX Runtime example (text)
python
1import json, numpy as np, onnxruntime as ort
2from pathlib import Path
3from transformers import AutoTokenizer
45d = Path("onnx/fp16")6man = json.loads((d /"manifest.json").read_text())7npdt = np.float16 if man["precision"]=="fp16"else np.float32
8tok = AutoTokenizer.from_pretrained(str(d))910defsess(name):# log level raised to silence the harmless constant-fold notice11 so = ort.SessionOptions(); so.log_severity_level =312return ort.InferenceSession(str(d / name), so, providers=["CPUExecutionProvider"])1314emb_s, back_s = sess("embeddings.onnx"), sess("backbone.onnx")1516# text-matching: both texts use the SAME "Query:" prefix17enc = tok(["Query: The impacts of climate change on coastal cities"], return_tensors="np", padding="longest")18ids, am = enc["input_ids"], enc["attention_mask"]19pos = np.clip(np.cumsum(am,-1)-1,0,None)[None].repeat(3,0)# MROPE (text)20e = emb_s.run(None,{"input_ids": ids,"image_features": np.zeros((0,2048), npdt)})[0]21h = back_s.run(None,{"inputs_embeds": e,"attention_mask": am,"position_ids": pos})[0]2223pooled =(h * am[...,None]).sum(1)/ am.sum(1, keepdims=True)# masked mean-pool24emb = pooled / np.linalg.norm(pooled, axis=-1, keepdims=True)# L2-norm → [1, 2048]