This repository contains the code and data for the paper "Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos".
🏠 Project Page📜 arXiv
🧑💻 GitHub
Sa2VA is the first unified model for the dense grounded understanding of both images and videos. Unlike existing multi-modal large language models, which are often limited to specific modalities and tasks, Sa2VA supports a wide range of image and video tasks, including referring segmentation and conversation… See the full description on the dataset page:
https://huggingface.co/datasets/tlzhang96/Sa2VA-Training.