Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Dinov2-SigLIP-Phi3-LoRA – AI Model by nms05 | AlphaNeural AI
You can deploy this model and start earning money today!
nms05
/
Dinov2-SigLIP-Phi3-LoRA
like
0
visual-question-answering
en
liuhaotian/LLaVA-Instruct-150K
liuhaotian/LLaVA-CC3M-Pretrain-595K
2401.06209
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
DinoV2-SigLIP-Phi3(LoRA) VLM
Vision Encoder
- DinoV2 + SigLIP @384px resolution.
Why 2 vision encoders?
Connector
- MLP (Dino and SigLIP features are concatenated and then projected to Phi3 representation space)
Language Model
- Phi3 + LoRA
Pre-train (Align) Dataset
- LLaVA-CC3M-Pretrain-595K
Fine-tune (Instruction) Dataset
- LLAVA-v1.5-Instruct + LRV-Instruct
Scripts to build and train the models are available at
NMS05/DinoV2-SigLIP-Phi3-LoRA-VLM
.