Code:
https://github.com/360CVGroup/FG-CLIP
FG-CLIP 2 is the foundation model for fine-grained vision-language understanding in both English and Chinese.
Across 29 datasets and 8 diverse tasks, it consistently surpasses recent strong baselines such as SigLIP 2 and MetaCLIP 2, achieving the best reported performance to date in both languages.
FG-CLIP 2: A Bilingual Fine-grained Vision-language Alignment Model… See the full description on the dataset page:
https://huggingface.co/datasets/qihoo360/DCI-CN.