This model is a fine-tuned version of
Qwen/Qwen3-VL-8B-Instruct on the all_sft_formats_unbalanced_20251122_part_1 dataset.
This model belongs to a family of video-language models (VLMs) optimized for precise video captioning using the CHAI (Critique-based Human–AI Oversight) framework. CHAI pairs trained human experts with model-generated pre-captions: experts provide correctional critiques that guide revisions into improved post-captions.
1@inproceedings{chai2026,
2 title = {Building a Precise Video Language with Human--AI Oversight},
3 author = {Zhiqiu Lin and Chancharik Mitra and Siyuan Cen and Isaac Li and Yuhan Huang and Yu Tong Tiffany Ling and Hewei Wang and Irene Pi and Shihang Zhu and Ryan Rao and George Liu and Jiaxi Li and Ruojin Li and Yili Han and Yilun Du and Deva Ramanan},
4 booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
5 year = {2026}
6}