1from transformers import pipeline
2
3clf = pipeline("text-classification", model="kash620/fashion-classifier")
4
5clf("Levi's 511 Slim Fit Stretch Jeans")
6# [{'label': 'Bottoms', 'score': 0.99}]
~157,000 real product titles scraped from fashion e-commerce stores, balanced to ~15,000 samples per class. Stratified 90/10 train/test split.
Cased tokenizer was chosen to preserve brand names and acronyms. Class weights handle residual imbalance after downsampling.
Evaluated on a stratified 10% holdout (~15,700 samples).
All 12 classes exceed 0.94 F1. The model handles ambiguous titles well (e.g. "Oversized Utility Shirt Jacket" → Outerwear, not Tops).
Evaluated on a stratified 10% holdout (~15,700 samples).
1@article{sanh2019distilbert,
2 title={DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter},
3 author={Sanh, Victor and Debut, Lysandre and Chaumond, Julien and Wolf, Thomas},
4 journal={arXiv preprint arXiv:1910.01108},
5 year={2019}
6}