Views
No views yet

| Scale | Tower | Params | Width | Depth | MLP | Heads | CLIP Dim | Resolution / Context Len |
|---|---|---|---|---|---|---|---|---|
| B/16 | Vision | 0.09B | 768 | 12 | 3072 | 12 | 1024 | 224px |
| Text | 0.31B | 1024 | 24 | 4096 | 16 | 1024 | 32 tokens | |
| L/14 | Vision | 0.32B | 1024 | 24 | 4096 | 16 | 1024 | 336px |
| Text | 0.31B | 1024 | 24 | 4096 | 16 | 1024 | 32 tokens | |
| G/14 | Vision | 1.88B | 1536 | 50 | 8960 | 16 | 1280 | 448px |
| Text | 0.47B | 1280 | 24 | 5120 | 20 | 1280 | 72 tokens |
| Model | Checkpoint | IN-1k | IN-v2 | IN-A | ObjectNet | COCO-T2I | Kinetics-400 | VTT-T2I |
|---|---|---|---|---|---|---|---|---|
| B/16 224px | PE-Core-B16-224 | 78.4 | 71.7 | 62.4 | 71.9 | 50.9 | 65.6 | 47.6 |
| L/14 336px | PE-Core-L14-336 | 83.5 | 77.9 | 89.0 | 84.7 | 57.1 | 73.4 | 50.3 |
| G/14 448px | PE-Core-G14-448 | 85.4 | 80.2 | 92.6 | 88.2 | 58.1 | 76.9 | 51.2 |
1git clone https://github.com/facebookresearch/perception_models.git
2cd perception_models
3conda create --name perception_models python=3.12
4conda activate perception_models
5# Install PyTorch
6pip install torch==2.5.1 torchvision==0.20.1 torchaudio==2.5.1 xformers --index-url https://download.pytorch.org/whl/cu124
7# We use torchcodec for decoding videos into PyTorch tensors
8conda install ffmpeg -c conda-forge
9pip install torchcodec==0.1 --index-url=https://download.pytorch.org/whl/cu124
10pip install -e .1import torch
2from PIL import Image
3import core.vision_encoder.pe as pe
4import core.vision_encoder.transforms as transforms
5
6print("CLIP configs:", pe.CLIP.available_configs())
7# CLIP configs: ['PE-Core-G14-448', 'PE-Core-L14-336', 'PE-Core-B16-224']
8
9model = pe.CLIP.from_config("PE-Core-L14-336", pretrained=True) # Downloads from HF
10model = model.cuda()
11
12preprocess = transforms.get_image_transform(model.image_size)
13tokenizer = transforms.get_text_tokenizer(model.context_length)
14
15image = preprocess(Image.open("docs/assets/cat.png")).unsqueeze(0).cuda()
16text = tokenizer(["a diagram", "a dog", "a cat"]).cuda()
17
18with torch.no_grad(), torch.autocast("cuda"):
19 image_features, text_features, logit_scale = model(image, text)
20 text_probs = (logit_scale * image_features @ text_features.T).softmax(dim=-1)
21
22print("Label probs:", text_probs) # prints: [[0.0, 0.0, 1.0]]@article{bolya2025PerceptionEncoder,
title={Perception Encoder: The best visual embeddings are not at the output of the network},
author={Daniel Bolya and Po-Yao Huang and Peize Sun and Jang Hyun Cho and Andrea Madotto and Chen Wei and Tengyu Ma and Jiale Zhi and Jathushan Rajasegaran and Hanoona Rasheed and Junke Wang and Marco Monteiro and Hu Xu and Shiyu Dong and Nikhila Ravi and Daniel Li and Piotr Doll{\'a}r and Christoph Feichtenhofer},
journal={arXiv},
year={2025}
}
@article{cho2025PerceptionLM,
title={PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding},
author={Jang Hyun Cho and Andrea Madotto and Effrosyni Mavroudi and Triantafyllos Afouras and Tushar Nagarajan and Muhammad Maaz and Yale Song and Tengyu Ma and Shuming Hu and Hanoona Rasheed and Peize Sun and Po-Yao Huang and Daniel Bolya and Suyog Jain and Miguel Martin and Huiyu Wang and Nikhila Ravi and Shashank Jain and Temmy Stark and Shane Moon and Babak Damavandi and Vivian Lee and Andrew Westbury and Salman Khan and Philipp Kr\"{a}henb\"{u}hl and Piotr Doll{\'a}r and Lorenzo Torresani and Kristen Grauman and Christoph Feichtenhofer},
journal={arXiv},
year={2025}
}