Views
No views yet
| model name | #bits | CR↑ | 0-shot↑ | rFID↓ | HF Link |
|---|---|---|---|---|---|
| QLIP-B-16-256 | 28 | 219.4 | 74.3 | 3.21 | 🤗 link |
| QLIP-B-8-256 | 28 | 54.8 | 75.6 | 0.70 | 🤗 link |
| QLIP-L-14-392 | 28 | 168 | 79.1 | 1.46 | 🤗 link |
1@article{zhao2025qlip,
2 title={QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation},
3 author={Zhao, Yue and Xue, Fuzhao and Reed, Scott and Fan, Linxi and Zhu, Yuke and Kautz, Jan and Yu, Zhiding and Krähenbühl, Philipp and Huang, De-An},
4 journal={arXiv preprint arXiv:2502.05178},
5 year={2025}
6}