Views
No views yet
| Model | Base Architecture | Other Remarks |
|---|---|---|
| Image Generation Models | ||
| VBVR-Pro-BAGEL | BAGEL-7B-MoT | Complete model |
| VBVR-Pro-FLUX2-dev | FLUX.2-dev | Complete model, Diffusers format |
| VBVR-Pro-FLUX2-dev-diffsynth | FLUX.2-dev | LoRA model, DiffSynth format |
| VBVR-Pro-Qwen-Image-Edit | Qwen-Image-Edit-2511 | Complete model, Diffusers format |
| VBVR-Pro-Qwen-Image-Edit-diffsynth | Qwen-Image-Edit-2511 | LoRA model, DiffSynth format |
| Interleaved Image Generation Models | ||
| VBVR-Pro-ThinkMorph | ThinkMorph-7B | Complete model |
| VBVR-Pro-SenseNova-U1 | SenseNova-U1-8B-MoT | Complete model |
| Video Generation Models | ||
| VBVR-Pro-LTX2.3 | LTX-Video-2.3 | Complete model, Diffusers format |
| VBVR-Pro-LTX2.3-diffsynth | LTX-Video-2.3 | LoRA model, DiffSynth format |
| VBVR-Pro-Wan2.1-I2V-14B | Wan2.1-I2V-14B-720P | Complete model, Diffusers format |
| VBVR-Pro-Wan2.1-I2V-14B-diffsynth | Wan2.1-I2V-14B-720P | LoRA model, DiffSynth format |
| VBVR-Pro-Wan2.2-I2V-A14B | Wan2.2-I2V-A14B | Complete model, Diffusers format |
| VBVR-Pro-Wan2.2-I2V-A14B-diffsynth | Wan2.2-I2V-A14B | LoRA model, DiffSynth format |
| VBVR-Pro-Wan2.2-TI2V-5B | Wan2.2-TI2V-5B | Complete model, Diffusers format |
| VBVR-Pro-Wan2.2-TI2V-5B-diffsynth | Wan2.2-TI2V-5B | LoRA model, DiffSynth format |
| Models | Overall | In-Domain by Category | Out-of-Domain by Category | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Avg. | Abst. | Know. | Perc. | Spat. | Trans. | Avg. | Abst. | Know. | Perc. | Spat. | Trans. | ||
| Image Generation Models | |||||||||||||
| Proprietary Models | |||||||||||||
| Qwen-Image-2.0 | 0.313 | 0.248 | 0.269 | 0.196 | 0.225 | 0.170 | 0.132 | 0.378 | 0.341 | 0.235 | 0.391 | 0.384 | 0.080 |
| Seedream-5.0-Pro | 0.557 | 0.485 | 0.518 | 0.312 | 0.509 | 0.401 | 0.217 | 0.629 | 0.507 | 0.455 | 0.661 | 0.559 | 0.202 |
| Open-source Models | |||||||||||||
| BAGEL-7B-MoT | 0.089 | 0.066 | 0.039 | 0.085 | 0.067 | 0.046 | 0.027 | 0.111 | 0.201 | 0.031 | 0.073 | 0.028 | 0.121 |
| FLUX.2-dev | 0.157 | 0.108 | 0.088 | 0.109 | 0.072 | 0.100 | 0.066 | 0.206 | 0.197 | 0.165 | 0.184 | 0.241 | 0.077 |
| Qwen-Image-Edit | 0.134 | 0.108 | 0.092 | 0.082 | 0.100 | 0.109 | 0.056 | 0.159 | 0.176 | 0.063 | 0.141 | 0.182 | 0.082 |
| Strong Baselines | |||||||||||||
| VBVR-Pro-BAGEL | 0.172 | 0.168 | 0.199 | 0.105 | 0.110 | 0.213 | 0.055 | 0.176 | 0.254 | 0.104 | 0.148 | 0.015 | 0.145 |
| VBVR-Pro-FLUX.2 | 0.407 | 0.484 | 0.483 | 0.323 | 0.367 | 0.449 | 0.336 | 0.330 | 0.361 | 0.272 | 0.255 | 0.454 | 0.128 |
| VBVR-Pro-Qwen-Image | 0.322 | 0.332 | 0.298 | 0.217 | 0.193 | 0.431 | 0.222 | 0.311 | 0.341 | 0.239 | 0.233 | 0.413 | 0.181 |
| Interleaved Image Generation Models | |||||||||||||
| Proprietary Models | |||||||||||||
| GPT-Image-2 | 0.507 | 0.428 | 0.456 | 0.318 | 0.428 | 0.206 | 0.300 | 0.587 | 0.398 | 0.413 | 0.633 | 0.480 | 0.303 |
| Nano Banana Pro | 0.564 | 0.480 | 0.518 | 0.422 | 0.512 | 0.285 | 0.174 | 0.648 | 0.553 | 0.499 | 0.657 | 0.585 | 0.220 |
| Open-source Models | |||||||||||||
| ThinkMorph-7B | 0.154 | 0.113 | 0.100 | 0.082 | 0.101 | 0.148 | 0.031 | 0.195 | 0.176 | 0.166 | 0.163 | 0.253 | 0.103 |
| VBVR-SenseNova-U1 | 0.408 | 0.469 | 0.356 | 0.313 | 0.373 | 0.386 | 0.477 | 0.347 | 0.291 | 0.317 | 0.275 | 0.480 | 0.238 |
| SenseNova-U1-8B-MoT | 0.565 | 0.533 | 0.501 | 0.395 | 0.544 | 0.355 | 0.349 | 0.597 | 0.448 | 0.495 | 0.533 | 0.717 | 0.401 |
| Strong Baselines | |||||||||||||
| VBVR-Pro-ThinkMorph | 0.373 | 0.402 | 0.403 | 0.344 | 0.238 | 0.454 | 0.184 | 0.344 | 0.367 | 0.224 | 0.238 | 0.535 | 0.257 |
| VBVR-Pro-SenseNova-U1 | 0.638 | 0.811 | 0.648 | 0.695 | 0.621 | 0.770 | 0.541 | 0.464 | 0.480 | 0.328 | 0.344 | 0.558 | 0.408 |
| Video Generation Models | |||||||||||||
| Proprietary Models | |||||||||||||
| Veo 3.1 | 0.309 | 0.312 | 0.275 | 0.299 | 0.252 | 0.267 | 0.157 | 0.305 | 0.305 | 0.233 | 0.252 | 0.312 | 0.219 |
| Kling V3 | 0.392 | 0.356 | 0.213 | 0.326 | 0.320 | 0.355 | 0.229 | 0.427 | 0.294 | 0.564 | 0.375 | 0.242 | 0.412 |
| SeedDance 2.0 | 0.499 | 0.451 | 0.338 | 0.361 | 0.353 | 0.468 | 0.308 | 0.547 | 0.369 | 0.511 | 0.478 | 0.538 | 0.532 |
| Open-source Models | |||||||||||||
| HunyuanVideo-I2V | 0.054 | 0.054 | 0.023 | 0.064 | 0.015 | 0.084 | 0.032 | 0.053 | 0.088 | 0.014 | 0.028 | 0.062 | 0.055 |
| CogVideoX1.5-5B-I2V | 0.085 | 0.100 | 0.061 | 0.118 | 0.069 | 0.092 | 0.060 | 0.070 | 0.125 | 0.038 | 0.051 | 0.040 | 0.024 |
| Wan2.1-I2V-14B | 0.100 | 0.105 | 0.052 | 0.125 | 0.091 | 0.102 | 0.052 | 0.095 | 0.112 | 0.073 | 0.071 | 0.123 | 0.044 |
| Wan2.2-TI2V-5B | 0.094 | 0.066 | 0.029 | 0.073 | 0.050 | 0.083 | 0.031 | 0.122 | 0.156 | 0.052 | 0.106 | 0.063 | 0.099 |
| Wan2.2-I2V-14B-720P | 0.182 | 0.157 | 0.082 | 0.131 | 0.110 | 0.161 | 0.156 | 0.207 | 0.224 | 0.139 | 0.140 | 0.195 | 0.273 |
| LTX2.3-I2AV | 0.112 | 0.106 | 0.062 | 0.109 | 0.070 | 0.133 | 0.055 | 0.119 | 0.161 | 0.135 | 0.086 | 0.091 | 0.050 |
| VBVR-Wan2.2 | 0.517 | 0.548 | 0.237 | 0.499 | 0.334 | 0.566 | 0.591 | 0.486 | 0.310 | 0.343 | 0.345 | 0.732 | 0.684 |
| Strong Baselines | |||||||||||||
| VBVR-Pro-LTX2.3 | 0.425 | 0.527 | 0.409 | 0.510 | 0.346 | 0.460 | 0.390 | 0.324 | 0.381 | 0.108 | 0.201 | 0.477 | 0.386 |
| VBVR-Pro-Wan2.1-I2V-14B | 0.562 | 0.730 | 0.617 | 0.580 | 0.452 | 0.676 | 0.623 | 0.395 | 0.410 | 0.305 | 0.230 | 0.617 | 0.439 |
| VBVR-Pro-Wan2.2-TI2V-5B | 0.470 | 0.641 | 0.528 | 0.556 | 0.373 | 0.565 | 0.557 | 0.300 | 0.333 | 0.127 | 0.161 | 0.505 | 0.409 |
| VBVR-Pro-Wan2.2-I2V-14B | 0.670 | 0.808 | 0.632 | 0.685 | 0.556 | 0.751 | 0.636 | 0.532 | 0.479 | 0.418 | 0.350 | 0.679 | 0.690 |
pip install -U diffusers transformers accelerate pillow imageio imageio-ffmpegexample.pyexample.py loads the merged checkpoint directly
with Diffusers and enables model CPU offloading.1python example.py \
2 --model_path Video-Reason/VBVR-Pro-Wan2.1-I2V-14B \
3 --image input.png \
4 --prompt "The subject walks toward the doorway." \
5 --num_frames 81 --width 832 --height 480 \
6 --output output.mp41git clone https://github.com/Video-Reason/VBVR-Pro.git
2cd VBVR-Pro/
3uv sync --extra cu124 # or one of [cu118|cu121|cu124|cu126|cu128|cu129]
4source .venv/bin/activate1python example.py \
2 --model_path Video-Reason/VBVR-Pro-Wan2.1-I2V-14B \
3 --image_paths input.png \
4 --prompt "The subject walks toward the doorway." \
5 --num_frames 81 --width 832 --height 480 \
6 --output output.mp41@misc{xu2026vbvrproscalableverifiablesuite,
2 title={VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning},
3 author={Junxiang Xu and Ruisi Wang and Fanyi Pu and Maijunxian Wang and Ran Ji and Tongxi Zhou and Chenyang Gu and Jing Zuo and Hongcan Xiao and Yimeng Geng and Wanqi Yin and Wei Chen and Oscar Qian and Zhengan Yan and Ziqi Huang and Haiwen Diao and Liang Pan and Bo Li and Xiangyu Fan and Dezhi Luo and Fengyuan Yu and Zehong Zhao and Qingying Gao and Tinghui Zhu and Yilan Zhang and Jingqi Tong and Pinyuan Feng and Zhengze Jiang and Letian Wang and Ziyu Guo and Renrui Zhang and Jieneng Chen and Sonia Joseph and Constantin Venhoff and Saman Motamed and Mengyue Yang and Chandra Sripada and Alan Yuille and Philip Torr and Lvmin Zhang and Vikash Kumar and Daniel Khashabi and Nikolaus Kriegeskorte and Raphaël Millière and Vincent C. Müller and Anyi Rao and Quan Wang and Ziwei Liu and Dahua Lin and Lei Yang and Hokin Deng and Zhongang Cai},
4 year={2026},
5 eprint={2608.26105},
6 archivePrefix={arXiv},
7 primaryClass={cs.CV},
8 url={https://arxiv.org/abs/2608.26105},
9}