Model type:
LLaVA-Next-Video is an open-source chatbot trained by fine-tuning LLM on multimodal instruction-following data. This model is the one mentioned in:
https://llava-vl.github.io/blog/2024-04-30-llava-next-video/
Base LLM: lmsys/vicuna-7b-v1.5
Llama 2 is licensed under the LLAMA 2 Community License,
Copyright (c) Meta Platforms, Inc. All Rights Reserved.
A collection of 4 benchmarks, including 3 academic VQA benchmarks and 1 captioning benchmark.