nanoLLaVA - Sub 1B Vision-Language Model
Description
nanoLLaVA is a "small but mighty" 1B vision-language model designed to run efficiently on edge devices.
Base LLM: Quyen-SE-v0.1 (Qwen1.5-0.5B)
Vision Encoder: google/siglip-so400m-patch14-384
Training Data will be released later as I am still writing a… See the full description on the dataset page:
https://huggingface.co/datasets/taiseimatsuoka/test-public.