Contrastive Language-Image Pretraining (CLIP) model pre-trained on 2.5 billion data points of CommonCrawl at resolution 224x224. It was introduced in the paper
Learning Transferable Visual Models From Natural Language Supervision and further reproduced in the follow-up paper
Demystifying CLIP Data.
The weights were converted from the
b32_fullcc2.5b.pt file presented in the
original repository.