ObjEmbed is a multimodal embedding model that decomposes an input image into multiple regional embeddings, each corresponding to an individual object, along with global embeddings. It is designed to bridge the gap between global image-text alignment and fine-grained region-phrase alignment.
If you find our work helpful for your research, please consider citing our paper:
1@article{fu2026objembed,
2 title={ObjEmbed: Towards Universal Multimodal Object Embeddings},
3 author={Fu, Shenghao and Su, Yukun and Rao, Fengyun and LYU, Jing and Xie, Xiaohua and Zheng, Wei-Shi},
4 journal={arXiv preprint arXiv:2602.01753},
5 year={2026}
6}