Views
No views yet
sapiens2_1b_pretrain.safetensorspip install -e .).1import torch
2from huggingface_hub import hf_hub_download
3from safetensors.torch import load_file
4from sapiens.backbones.standalone.sapiens2 import Sapiens2
5
6# Build the model and load the pretrained checkpoint
7model = Sapiens2(arch="sapiens2_1b", img_size=(1024, 768), patch_size=16).eval().cuda() # img_size is (H, W)
8ckpt_path = hf_hub_download(repo_id="facebook/sapiens2-pretrain-1b", filename="sapiens2_1b_pretrain.safetensors")
9model.load_state_dict(load_file(ckpt_path))
10
11# Forward pass on a single image (RGB; ImageNet normalization recommended)
12x = torch.randn(1, 3, 1024, 768).cuda()
13with torch.no_grad():
14 features = model(x)[0] # dense backbone features: (B, num_tokens, embed_dim)| Field | Value |
|---|---|
| Architecture | Sapiens2 ViT (RoPE, GQA, SwiGLU, RMSNorm, QK-norm) |
| Parameters | 1.462 B |
| FLOPs | 4.715 T |
| Embedding dim | 1536 |
| Layers | 40 |
| Attention heads | 24 |
| Pretraining resolution | 1024 × 768 (H × W) |
| Patch size | 16 |
| Pretraining data | 1B human images |
| Model | Params | FLOPs | Embed dim | Layers | Heads |
|---|---|---|---|---|---|
| Sapiens2-0.1B | 0.114 B | 0.342 T | 768 | 12 | 12 |
| Sapiens2-0.4B | 0.398 B | 1.260 T | 1024 | 24 | 16 |
| Sapiens2-0.8B | 0.818 B | 2.592 T | 1280 | 32 | 16 |
| Sapiens2-1B (this) | 1.462 B | 4.715 T | 1536 | 40 | 24 |
| Sapiens2-1B-4K | 1.607 B | — | 1536 | 40 | 24 |
| Sapiens2-5B | 5.071 B | 15.722 T | 2432 | 56 | 32 |
1@article{khirodkarsapiens2,
2 title={Sapiens2},
3 author={Khirodkar, Rawal and Wen, He and Martinez, Julieta and Dong, Yuan and Su, Zhaoen and Saito, Shunsuke},
4 journal={arXiv preprint arXiv:2604.21681},
5 year={2026}
6}