Video-CCAM-14B is a lightweight Video-MLLM built on
Phi-3-medium-4k-instruct and
SigLIP SO400M.
Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.10:
torch==2.1.0
torchvision==0.16.0
transformers==4.40.2
peft==0.10.0
Please refer to
Video-CCAM on inference and evaluation.
The model is licensed under the MIT license.