This model is a fine-tuned version of google/vit-base-patch16-224-in21k on a balanced subset of the ashraq/fashion-product-images-small dataset.
Model Description
Task: Multi-class Image Classification
Categories: Apparel, Accessories, Footwear, Personal Care
Base Model: Vision Transformer (ViT) Base
Training Hyperparameters
Image Size: 224
Batch Size: 8
Epochs: 1
Learning Rate: 2e-05
Seed: 5
Evaluation Results
Based on the test set (1200 samples):
Metric
Value
Accuracy (Pre-trained)
0.2308
Accuracy (Fine-tuned)
0.9850
Avg Inference Time
0.92 ms
Pipeline Integration
This model serves as Pipeline 1 in an automated ad generation system, providing the product category which is then combined with BLIP captions for LLM-based ad copy generation.