Views
No views yet
cfg_text_scale: Use 4.0–8.0 for balanced prompt following.cfg_renorm_type: Use global for general Text-to-Image tasks.timestep_shift: Higher values for better layout; lower values for finer details.num_timesteps: Standard setting is 50.| Model | TIIF (Short/Long) | WISE (Overall) | OneIG-EN (Overall) | CompBench (Overall) | DPG (Score) | Geneval (Score) |
|---|---|---|---|---|---|---|
| BAGEL | 71.0 / 71.8 | 50.0 | 36.1 | 82.2 | 84.0 | 78.0 |
| UniCorn | 74.7 / 72.9 | 55.0 | 42.6 | 88.5 | 86.8 | 82.0 |
| $\Delta$(vs. BAGEL) | +3.7 / +1.1 | +5.0 | +6.5 | +6.3 | +2.8 | +4.0 |
1@article{han2026unicorn,
2 title={UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision},
3 author={Han, Ruiyan and Fang, Zhen and Sun, Xinyu and Ma, Yuchen and Wang, Ziheng and Zeng, Yu and Chen, Zehui and Chen, Lin and Huang, Wenxuan and Xu, Wei-Jie and others},
4 journal={arXiv preprint arXiv:2601.03193},
5 year={2026}
6}
7