CLIP Italian is a CLIP-like Model for Italian. The CLIP model (Contrastive Language–Image Pre-training) was developed by researchers at OpenAI and is able to efficiently learn visual concepts from natural language supervision.
We fine-tuned a competitive Italian CLIP model with only ~1.4 million Italian image-text pairs. This model is part of the
Flax/Jax Community Week, organized by
HuggingFace and TPU usage sponsored by Google.
Preprocessing, hardware used, hyperparameters...