This model is a fine-tuned version of
google/vit-base-patch16-224 on the pcuenq/oxford-pets dataset.
It achieves the following results on the evaluation set:
Zero-Shot Evaluation
Model used: openai/clip-vit-large-patch14
Dataset: Oxford-IIIT-Pets ( pcuenq/oxford-pets )
Accuracy: 0.8800
Precision: 0.8768
Recall: 0.8800
The zero-shot evaluation was done using Hugging Face Transformers and the CLIP model on the Oxford-Pet dataset.