Views
No views yet
A self-supervised training framework that aligns understanding and generation in modest compute, with huge zero-shot gain on generation and editing capability.
| Model | GenEval ↑ | DPGBench ↑ | WISE ↑ |
|---|---|---|---|
| Show-o | 0.57 | 70.65 | 0.33 |
| Show-o-RecA | 0.62 | 75.70 | 0.34 |
@article{xie2025reconstruction,
title={Reconstruction Alignment Improves Unified Multimodal Models},
author={Xie, Ji and Darrell, Trevor and Zettlemoyer, Luke and Wang, XuDong},
journal={arXiv preprint arXiv:2509.07295},
year={2025}
}