Beta
Explore
Marketplace
Neural Labs
Playground
Wallet
Docs
Video-LLaVA-Seg – AI Model by fun-research | AlphaNeural AI
You can deploy this model and start earning money today!
fun-research
/
Video-LLaVA-Seg
like
0
transformers
safetensors
llava_llama
text-generation
video-text-to-text
2412.09754
apache-2.0
autotrain_compatible
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Video-LLaVA-Seg
Project
|
Arxiv
This is the official baseline implementation for the ViCas dataset, presented in the paper
ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation
.
For details about setting up the model, refer to the
Video-LLaVA-Seg GitHub repo
For details about downloading and evaluating the dataset benchmark, refer to the
ViCaS GitHub repo