Views
No views yet
uform3-image-text-english-base UForm model is a tiny vision and English language encoder, mapping them into a shared vector space.
This model produces up to 256-dimensional embeddings and is made of:| Dataset | Recall@1 | Recall@5 | Recall@10 |
|---|---|---|---|
| Zero-Shot Flickr | 0.727 | 0.915 | 0.949 |
| MS-COCO ¹ | 0.510 | 0.761 | 0.838 |
¹ It's important to note, that the MS-COCO train split was present in the training data.
pip install "uform[torch,onnx]"1from uform import get_model, Modality
2
3import requests
4from io import BytesIO
5from PIL import Image
6
7model_name = 'unum-cloud/uform3-image-text-english-base'
8modalities = [Modality.TEXT_ENCODER, Modality.IMAGE_ENCODER]
9processors, models = get_model(model_name, modalities=modalities)
10
11model_text = models[Modality.TEXT_ENCODER]
12model_image = models[Modality.IMAGE_ENCODER]
13processor_text = processors[Modality.TEXT_ENCODER]
14processor_image = processors[Modality.IMAGE_ENCODER]1text = 'a cityscape bathed in the warm glow of the sun, with varied architecture and a towering, snow-capped mountain rising majestically in the background'
2image_url = 'https://media-cdn.tripadvisor.com/media/photo-s/1b/28/6b/53/lovely-armenia.jpg'
3image_url = Image.open(BytesIO(requests.get(image_url).content))
4
5image_data = processor_image(image)
6text_data = processor_text(text)
7image_features, image_embedding = model_image.encode(image_data, return_features=True)
8text_features, text_embedding = model_text.encode(text_data, return_features=True)