The ProLongVid-v1 models are 7B parameter models trained on
ProLongVid_data, based on our extended Qwen2.5 language model with a context window of 256K tokens.
We suggest using this model with up to 256 frames.
1@inproceedings{wang2025prolongvid,
2 title={ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction Tuning},
3 author={Wang, Rui and Li, Bohao and Dai, Xiyang and Yang, Jianwei and Chen, Yi-Ling and Xing, Zhen and Yang, Yifan and Chen, Dongdong and Qiu, Xipeng and Wu, Zuxuan and others},
4 booktitle={EMNLP},
5 year={2025}
6}