Views
No views yet
mair-lab/earl-thinking-sft-simple.rl-simple-n-complexmair-lab/sft-think-simplesft-think-simple) and is optimized using reinforcement learning across both simple and complex edit instructions. While it incorporates chain-of-thought supervision, it still trails the non-reasoning RL model in overall benchmark performance.| Model Description | OmniEdit | EmuEdit | AURORA | MB | VisMin | I2EBench | AVG |
|---|---|---|---|---|---|---|---|
| SFT think (S) | 4.34 | 3.76 | 2.88 | 3.36 | 3.46 | 3.21 | 3.50 |
| EARL SFT think (S) → RL (S+C) | 4.65 | 3.78 | 3.23 | 3.67 | 3.39 | 3.36 | 3.68 |
📉 Note: The RL version improves modestly over the SFT think (S) baseline but does not match the performance of the non-reasoning SFT (S) RL model, indicating current limitations of reasoning-guided supervision in image editing.