Views
No views yet
1import torch
2from PIL import Image
3from transformers import AutoModel, AutoImageProcessor
4
5model_id = "tiiuae/siglino-0.6B"
6model = AutoModel.from_pretrained(model_id, trust_remote_code=True).to("cuda", dtype=torch.bfloat16)
7processor = AutoImageProcessor.from_pretrained(model_id, trust_remote_code=True)
8
9image = Image.open("image.jpg").convert("RGB")
10inputs = processor(image, return_tensors="pt").to("cuda")
11inputs["pixel_values"] = inputs["pixel_values"].to(torch.bfloat16)
12
13with torch.no_grad():
14 outputs = model(**inputs)
15
16# Options: 'siglino' (1280d), 'siglip2' (1152d), 'dinov3' (1024d)
17patch_features = outputs["patch_features"]["siglino"] # (Batch, Tokens, 1280)
18summary_features = outputs["summary_features"]["siglip2"] # (Batch, 1152)| Property | Value |
|---|---|
| Architecture | Dense |
| Parameters | 0.6B |
| Layers | 18 |
| Hidden Dim | 1280 |
| FFN Dim | 5120 |
| Patch Size | 16x16 |
| Teachers | DINOv3, SigLIP2 |
| Task | Metric | Score |
|---|---|---|
| kNN (ImageNet) | Acc | 86.1 |
| kNN (6-dataset avg) | Acc | 90.7 |
| Zero-shot cls (ImageNet) | Acc | 80.5 |
| Flickr30K I2T | R@1 | 94.2 |
| MSCOCO I2T | R@1 | 72.9 |
| Pascal VOC (1024) | mIoU | 89.8 |
| Cityscapes (1024) | mIoU | 67.3 |
1@article{chaybouti2025amoe,
2 title={AMoE: Agglomerative Mixture-of-Experts Vision Foundation Models},
3 author={Chaybouti, Sofian and Narayan, Sanath and Dahou, Yasser and Le Khac, Phuc H. and Singh, Ankit and Huynh, Ngoc Dung and Para, Wamiq Reyaz and Kuehne, Hilde and Hacid, Hakim},
4 journal={arXiv preprint arXiv:2512.20157},
5 year={2025}
6}