Views
No views yet
| Model | Base Architecture | Other Remarks |
|---|---|---|
| VBVR-Wan2.1 | Wan2.1-I2V-14B-720P | Diffusers format |
| VBVR-Wan2.2 | Wan2.2-I2V-A14B | Diffusers format |
| VBVR-Wan2.1-diffsynth | Wan2.1-I2V-14B-720P | DiffSynth LoRA format |
| VBVR-Wan2.2-diffsynth | Wan2.2-I2V-A14B | DiffSynth LoRA format |
| VBVR-LTX2.3-diffsynth | LTX-Video-2.3 | DiffSynth LoRA format |
| Model | Overall | ID | ID-Abst. | ID-Know. | ID-Perc. | ID-Spat. | ID-Trans. | OOD | OOD-Abst. | OOD-Know. | OOD-Perc. | OOD-Spat. | OOD-Trans. |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Human | 0.974 | 0.960 | 0.919 | 0.956 | 1.00 | 0.95 | 1.00 | 0.988 | 1.00 | 1.00 | 0.990 | 1.00 | 0.970 |
| Open-source Models | |||||||||||||
| CogVideoX1.5-5B-I2V | 0.273 | 0.283 | 0.241 | 0.328 | 0.257 | 0.328 | 0.305 | 0.262 | 0.281 | 0.235 | 0.250 | 0.254 | 0.282 |
| HunyuanVideo-I2V | 0.273 | 0.280 | 0.207 | 0.357 | 0.293 | 0.280 | 0.316 | 0.265 | 0.175 | 0.369 | 0.290 | 0.253 | 0.250 |
| Wan2.2-I2V-A14B | 0.371 | 0.412 | 0.430 | 0.382 | 0.415 | 0.404 | 0.419 | 0.329 | 0.405 | 0.308 | 0.343 | 0.236 | 0.307 |
| LTX-2 | 0.313 | 0.329 | 0.316 | 0.362 | 0.326 | 0.340 | 0.306 | 0.297 | 0.244 | 0.337 | 0.317 | 0.231 | 0.311 |
| Proprietary Models | |||||||||||||
| Seedance 2.0 | 0.544 | 0.570 | 0.593 | 0.498 | 0.618 | 0.514 | 0.602 | 0.517 | 0.643 | 0.398 | 0.492 | 0.427 | 0.556 |
| Runway Gen-4 Turbo | 0.403 | 0.392 | 0.396 | 0.409 | 0.429 | 0.341 | 0.363 | 0.414 | 0.515 | 0.429 | 0.419 | 0.327 | 0.373 |
| Sora 2 | 0.546 | 0.569 | 0.602 | 0.477 | 0.581 | 0.572 | 0.597 | 0.523 | 0.546 | 0.472 | 0.525 | 0.462 | 0.546 |
| Kling 2.6 | 0.369 | 0.408 | 0.465 | 0.323 | 0.375 | 0.347 | 0.519 | 0.330 | 0.528 | 0.135 | 0.272 | 0.356 | 0.359 |
| Veo 3.1 | 0.480 | 0.531 | 0.611 | 0.503 | 0.520 | 0.444 | 0.510 | 0.429 | 0.577 | 0.277 | 0.420 | 0.441 | 0.404 |
| Data Scaling Strong Baseline | |||||||||||||
| VBVR-LTX2.3 | 0.516 | 0.580 | 0.608 | 0.631 | 0.529 | 0.454 | 0.680 | 0.453 | 0.608 | 0.577 | 0.409 | 0.414 | 0.388 |
| VBVR-Wan2.1 | 0.592 | 0.724 | 0.705 | 0.710 | 0.727 | 0.719 | 0.784 | 0.461 | 0.674 | 0.592 | 0.387 | 0.461 | 0.387 |
| VBVR-Wan2.2 | 0.685 | 0.760 | 0.724 | 0.750 | 0.782 | 0.745 | 0.833 | 0.610 | 0.768 | 0.572 | 0.547 | 0.618 | 0.615 |
uv installation guide: https://docs.astral.sh/uv/getting-started/installation/#installing-uv
1pip install torch>=2.4.0 torchvision>=0.19.0 transformers Pillow huggingface_hub[cli]
2uv pip install git+https://github.com/huggingface/diffusers1huggingface-cli download Video-Reason/VBVR-Wan2.2 --local-dir ./VBVR-Wan2.2
2
3python example.py \
4 --model_path ./VBVR-Wan2.21@article{vbvr2026,
2 title = {A Very Big Video Reasoning Suite},
3 author = {Wang, Maijunxian and Wang, Ruisi and Lin, Juyi and Ji, Ran and
4 Wiedemer, Thadd{\"a}us and Gao, Qingying and Luo, Dezhi and
5 Qian, Yaoyao and Huang, Lianyu and Hong, Zelong and Ge, Jiahui and
6 Ma, Qianli and He, Hang and Zhou, Yifan and Guo, Lingzi and
7 Mei, Lantao and Li, Jiachen and Xing, Hanwen and Zhao, Tianqi and
8 Yu, Fengyuan and Xiao, Weihang and Jiao, Yizheng and
9 Hou, Jianheng and Zhang, Danyang and Xu, Pengcheng and
10 Zhong, Boyang and Zhao, Zehong and Fang, Gaoyun and Kitaoka, John and
11 Xu, Yile and Xu, Hua bureau and Blacutt, Kenton and Nguyen, Tin and
12 Song, Siyuan and Sun, Haoran and Wen, Shaoyue and He, Linyang and
13 Wang, Runming and Wang, Yanzhi and Yang, Mengyue and Ma, Ziqiao and
14 Milli{\`e}re, Rapha{\"e}l and Shi, Freda and Vasconcelos, Nuno and
15 Khashabi, Daniel and Yuille, Alan and Du, Yilun and Liu, Ziming and
16 Lin, Dahua and Liu, Ziwei and Kumar, Vikash and Li, Yijiang and
17 Yang, Lei and Cai, Zhongang and Deng, Hokin},
18 journal = {arXiv preprint arXiv:2602.20159},
19 year = {2026},
20 url = {https://arxiv.org/abs/2602.20159}
21}