Views
No views yet
UnifiedReward-2.0-qwen3vl-8b is the first unified reward model based on Qwen/Qwen3-VL-8B-Instruct for multimodal understanding and generation assessment, enabling both pairwise ranking and pointwise scoring, which can be employed for vision model preference alignment.| Reward Model | Method | Image Generation | Image Understanding | Video Generation | Video Understanding |
|---|---|---|---|---|---|
| PickScore | Point | √ | |||
| HPS | Point | √ | |||
| ImageReward | Point | √ | |||
| LLaVA-Critic | Pair/Point | √ | |||
| IXC-2.5-Reward | Pair/Point | √ | √ | ||
| VideoScore | Point | √ | |||
| LiFT | Point | √ | |||
| VisionReward | Point | √ | √ | ||
| VideoReward | Point | √ | |||
| UnifiedReward (Ours) | Pair/Point | √ | √ | √ | √ |
@article{unifiedreward,
title={Unified reward model for multimodal understanding and generation},
author={Wang, Yibin and Zang, Yuhang and Li, Hao and Jin, Cheng and Wang, Jiaqi},
journal={arXiv preprint arXiv:2503.05236},
year={2025}
}