This model is a fine-tuned version of
google/vit-base-patch16-224-in21k on the food101 dataset.
It achieves the following results on the evaluation set:
This is an image classification model fine tuned from the Google Vision Transformer (ViT) to classify images of food.
The training set contained 101 food classes, over a dataset of 101,000 images. The train/eval split was 80/20