ViT-Flower: Vision Transformer for Flower Classification
Model Description
ViT-Flower is a Vision Transformer (ViT-Base-Patch16-384) model fine-tuned for flower classification. The model classifies images into 152 different flower categories with both Chinese and English names.
Model Architecture: Vision Transformer (ViT-Base-Patch16-384)
See id2label_full.json for complete 152 class mappings.
Training Details
Parameter
Value
Epochs
40
Learning Rate
1.5e-4
Optimizer
AdamW
Weight Decay
8e-5
Batch Size
32
Warmup Epochs
6
Frozen Blocks
10 of 12
Loss Function
FocalCrossEntropyLoss (α=0.25, γ=2)
Label Smoothing
0.1
Limitations
Input images must be RGB format
Optimal performance on flower images similar to training distribution
Model expects 384×384 input resolution
Citation
bibtex
1@article{dosovitskiy2021vit,
2 title={An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale},
3 author={Dosovitskiy, Alexey and Beyer, Lucas and Kolesnikov, Alexander and others},
4 journal={ICLR},
5 year={2021}
6}