Views
No views yet
1torch == 1.13.1
2torchvision == 0.14.1
3timm == 0.6.12
4einops == 0.6.1| Model | Parameters (M) | FLOPs(G) | Top-1 Accuracy (%) |
|---|---|---|---|
| FDViT-Ti | 4.6 | 0.6 | 73.74 |
| FDViT-S | 21.6 | 2.8 | 81.45 |
| FDViT-B | 68.1 | 11.9 | 82.39 |
1from transformers import AutoModelForImageClassification
2import torch
3
4model = AutoModelForImageClassification.from_pretrained("FDViT_s", trust_remote_code=True)
5
6model.eval()
7
8inp = torch.ones(1,3,224,224)
9out = model(inp)@inproceedings{xu2023fdvit,
title={FDViT: Improve the Hierarchical Architecture of Vision Transformer},
author={Xu, Yixing and Li, Chao and Li, Dong and Sheng, Xiao and Jiang, Fan and Tian, Lu and Sirasao, Ashish},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
pages={5950--5960},
year={2023}
}