Views
No views yet
A self-supervised training framework that aligns understanding and generation in modest compute, with huge zero-shot gain on generation and editing capability.
| Model | GenEval ↑ | DPGBench ↑ | WISE ↑ |
|---|---|---|---|
| BAGEL | 0.787 | 84.03 | 0.50 |
| BAGEL-RecA | 0.824 | 85.29 | 0.52 |
| Model | GEdit-Bench-EN (SC) ↑ | GEdit-Bench-EN (PQ) ↑ | GEdit-Bench-EN (O) ↑ | ImgEdit ↑ |
|---|---|---|---|---|
| BAGEL | 7.96 | 6.64 | 6.94 | 3.38 |
| BAGEL-NHR | 8.04 | 6.87 | 7.08 | 3.48 |
| BAGEL-RecA | 8.24 | 6.87 | 7.27 | 3.75 |
| FLUX Kontext | 6.95 | 7.30 | 6.27 | 3.59 |

@article{xie2025reconstruction,
title={Reconstruction Alignment Improves Unified Multimodal Models},
author={Xie, Ji and Darrell, Trevor and Zettlemoyer, Luke and Wang, XuDong},
journal={arXiv preprint arXiv:2509.07295},
year={2025}
}