SR-GRPO-LoRA is a LoRA adapter for preference-aligned super-resolution, obtained by fine-tuning a FLUX.1-dev-based ControlNet SR model with GRPO.
For inference, load the released
SR-GRPO-LoRA weights into the super-resolution pipeline built with
FLUX.1-dev and the
C-FLUX ControlNet checkpoint distributed by
DP2O-SR.
The data were prepared with the
SeeSR degradation pipeline to produce ground-truth images, degraded LR images, and tag prompts.
This model card does not declare a new license for the released LoRA. The relevant upstream terms are:
Users are responsible for reviewing and complying with all applicable upstream terms.
1@article{song2026refreward,
2 title={RefReward-SR: LR-Conditioned Reward Modeling for Preference-Aligned Super-Resolution},
3 author={Song, Yushuai and Quan, Weize and Wang, Weining and Sun, Jiahui and Liu, Jing and Li, Meng and Yu, Pengbin and Chen, Zhentao and Shen, Wei and Yuan, Lunxi and Yan, Dong-ming},
4 journal={arXiv preprint arXiv:2603.24198},
5 year={2026}
6}