Views
No views yet
| Backbone | # Mixture tokens | Avg ZS Acc. | HF model id |
|---|---|---|---|
| ViT-B/16 | 32 | 69.6 | lavoies/llip-vitb-16-224 |
| ViT-G/14 | 64 | 79.3 | lavoies/llip-vitG-14-224 |
>>> from transformers import AutoModel
>>> model = AutoModel.from_pretrained('lavoies/llip-vitb-16-224')@inproceedings{lavoie2024modeling,
title={Modeling Caption Diversity in Contrastive Vision-Language Pretraining},
author={Samuel Lavoie and Polina Kirichenko and Mark Ibrahim and Mido Assran and Andrew Gordon Wilson and Aaron Courville and Nicolas Ballas},
booktitle={Forty-first International Conference on Machine Learning},
year={2024},
url={https://openreview.net/forum?id=iaV2fU6Dif}
}