VideoITG-8B is a multimodal video understanding model trained with instructed temporal grounding, equipped with the ability to enhance Video Large Language Models through intelligent frame selection. The model tackles the complexities of real-world video scenarios by aligning frame sampling with user instructions. Please check our paper for more details.
1@article{wang2025videoitg,
2 title = {VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding},
3 author = {Shihao Wang and Guo Chen and De-An Huang and Zhiqi Li and Minghan Li and Guilin Liu and Jose M. Alvarez and Lei Zhang and Zhiding Yu},
4 journal = {arXiv preprint arXiv:2507.13353},
5 year = {2025}
6}