A banking intent classification system built on top of jinaai/jina-embeddings-v5-text-nano-classification. The model uses Jina's nano embedding model to produce text representations, which are then fed into a Support Vector Classifier (SVC) with probability estimates to predict the intent category of incoming banking-related messages.
All credit for the underlying embedding model goes to Jina AI. This repo simply packages their nano classification model alongside a trained SVC head (classifier.pkl) for use in a banking message routing pipeline.
How it works
Input text is encoded using jina-embeddings-v5-text-nano-classification via the ONNX backend
The resulting embedding is passed to an SVC classifier trained on labeled banking messages
The SVC outputs a predicted intent label along with a confidence score via predict_proba
Contents
File
Description
onnx/model.onnx + onnx/model.onnx_data
Jina nano ONNX weights
classifier.pkl
Trained SVC classifier head
config.json, tokenizer.json, etc.
Jina model config and tokenizer files
Usage
python
1from huggingface_hub import snapshot_download
2from sentence_transformers import SentenceTransformer
3import pickle, numpy as np
45local_path = snapshot_download("TonitoMC/jina-analitica-bam")6embedder = SentenceTransformer(local_path, trust_remote_code=True, backend="onnx")78withopen(f"{local_path}/classifier.pkl","rb")as f:9 clf = pickle.load(f)1011text ="I want to transfer money to another account"12embedding = embedder.encode([text])13label = clf.predict(embedding)[0]14confidence = clf.predict_proba(embedding).max()
`jina-embeddings-v5-text-nano-classification` is a compact, high-performance text embedding model designed for classification.
It is part of the jina-embeddings-v5-text model family, which also includes jina-embeddings-v5-text-small, for better performance at a bigger size.
Trained using a novel approach that combines distillation with task-specific contrastive losses, jina-embeddings-v5-text-nano-classification outperforms existing state-of-the-art models of similar size across diverse embedding benchmarks.
Feature
Value
Parameters
239M
Supported Tasks
classification
Max Sequence Length
8192
Embedding Dimension
768
Matryoshka Dimensions
32, 64, 128, 256, 512, 768
Pooling Strategy
Last-token pooling
Base Model
jinaai/jina-embeddings-v5-text-nano
image
Training and Evaluation
For training details and evaluation results, see our technical report.
Usage
Requirements
The following Python packages are required:
transformers>=5.1.0
torch>=2.8.0
peft>=0.15.2
vllm==0.15.1
Optional / Recommended
flash-attention: Installing flash-attention is recommended for improved inference speed and efficiency, but not mandatory.
sentence-transformers: If you want to use the model via the sentence-transformers interface, install this package as well.
The fastest way to use v5-text in production. Elastic Inference Service (EIS) provides managed embedding inference with built-in scaling, so you can generate embeddings directly within your Elastic deployment.
1from sentence_transformers import SentenceTransformer
2import torch
34model = SentenceTransformer(5"jinaai/jina-embeddings-v5-text-nano-classification",6 trust_remote_code=True,7 model_kwargs={"dtype": torch.bfloat16},# Recommended for GPUs8 config_kwargs={"_attn_implementation":"flash_attention_2"},# Recommended but optional9)10# Optional: set truncate_dim in encode() to control embedding size1112texts =[13"My order hasn't arrived yet and it's been two weeks.",14"How do I reset my password?",15"I'd like a refund for my recent purchase.",16"Your product exceeded my expectations. Great job!",17]1819# Encode texts20embeddings = model.encode(texts)21print(embeddings.shape)22# (4, 768)2324similarity = model.similarity(embeddings, embeddings)25print(similarity)26# tensor([[1.0000, 0.7152, 0.8378, 0.8101],27# [0.7152, 1.0000, 0.7512, 0.6940],28# [0.8378, 0.7512, 1.0000, 0.7741],29# [0.8101, 0.6940, 0.7741, 1.0000]])
Since our nano model is based on jinaai/jina-embeddings-v5-text-nano, which is not yet supported by llama.cpp, we provide our own branch of llama.cpp, which implements the necessary changes to support it for now.
To start the OpenAI API compatible HTTP server, run with the respective model version:
curl -X POST "http://127.0.0.1:8080/v1/embeddings" \
-H "Content-Type: application/json" \
-d '{
"input": [
"Document: A beautiful sunset over the beach",
"Document: Un beau coucher de soleil sur la plage",
"Document: 海滩上美丽的日落",
"Document: 浜辺に沈む美しい夕日",
"Document: Golden sunlight melts into the horizon, painting waves in warm amber and rose, while the sky whispers goodnight to the quiet, endless sea."
]
}'
Note: For the classification variant, always add Document: prefix in front of your input as shown above.
You can run the ONNX-optimized version of the model locally using Hugging Face's optimum library. Make sure you have the required dependencies installed (e.g., pip install optimum[onnxruntime] transformers torch):
python
1from optimum.onnxruntime import ORTModelForFeatureExtraction
2from transformers import AutoTokenizer
3import torch
45model_id ="jinaai/jina-embeddings-v5-text-nano-classification"67# 1. Load tokenizer and ONNX model8# We specify the subfolder 'onnx' where the weights are located9tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)10model = ORTModelForFeatureExtraction.from_pretrained(11 model_id,12 subfolder="onnx",13 file_name="model.onnx",14 provider="CPUExecutionProvider",# Or "CUDAExecutionProvider" for GPU15 trust_remote_code=True,16)1718# 2. Prepare input19texts =["Document: How do I use Jina ONNX models?","Document: Information about semantic matching."]20inputs = tokenizer(texts, padding=True, truncation=True, return_tensors="pt")212223# 4. Inference24with torch.no_grad():25 outputs = model(**inputs)2627# 5. Pooling (Crucial for Jina-v5)28# Jina-v5 uses LAST-TOKEN pooling.29# We take the hidden state of the last non-padding token.30last_hidden_state = outputs.last_hidden_state
31# Find the indices of the last token (usually the end of the sequence)32sequence_lengths = inputs.attention_mask.sum(dim=1)-133embeddings = last_hidden_state[torch.arange(last_hidden_state.size(0)), sequence_lengths]3435print('embeddings shape:', embeddings.shape)36print('embeddings:', embeddings)
License
The model is licensed under CC BY-NC 4.0. For commercial use, please contact us.
Citation
If you find jina-embeddings-v5-text-nano-classification useful in your research, please cite the following paper:
@misc{akram2026jinaembeddingsv5texttasktargetedembeddingdistillation,
title={jina-embeddings-v5-text: Task-Targeted Embedding Distillation},
author={Mohammad Kalim Akram and Saba Sturua and Nastia Havriushenko and Quentin Herreros and Michael Günther and Maximilian Werk and Han Xiao},
year={2026},
eprint={2602.15547},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2602.15547},
}