This repository hosts the pretrained model weights for ViQ. For the inference / training / weight-conversion code, see the main repo: https://github.com/yuxumin/ViQ.
ViQ is trained in two stages, and this repository provides weights for both stages:
The text-aligned, any-resolution ViT encoders produced after Stage 1 pre-training. Two backbone sizes are released:
Size
Backbone
File
400M
SigLIP2-SO400M
anyres_vit/so400m/siglip2_so400m_anyres_s4.pth
1B
SigLIP2-g
anyres_vit/giant1b/siglip2_g_anyres_s4.pth
🔢 ViQ/ — Stage 2 (Discrete Tokenizers)
The discretized ViQ tokenizers produced after Stage 2, released in several FSQ codebook sizes. Each converted_<size>/ folder contains the ViQ-inference-format weights:
1@article{yu2026viq,
2 title = {ViQ: Text-Aligned Visual Quantized Representations at Any Resolution},
3 author = {Yu, Xumin and Liu, Zuyan and Yang, Zhenyu and Dong, Yuhao and Qian, Shengsheng and Lu, Jiwen and Hu, Han and Rao, Yongming},
4 journal = {arXiv preprint arXiv:xxxx.xxxxx},
5 year = {2026}
6}