Researchers have developed a new method, UMM-Reflection, that uses reinforcement learning to enable unified multimodal models to correct their own image generation flaws. These models, which can both understand and create images, are trained to identify errors in their generated images, revise them, and then evaluate the success of the revision. This process is learned jointly over the entire generation loop.
Traditional supervised fine-tuning (SFT) provides a starting point for this self-correction process but does not effectively discover successful repair strategies. Naive reinforcement learning (RL) approaches that focus on only one part of the model, such as the image renderer or a single output head, leave significant potential gains untapped. UMM-Reflection addresses this by applying RL to the complete "reflection" trajectories within a single model.
In this interleaved reinforcement learning approach, multiple generation attempts that start from the same initial image share a common advantage calculation. This allows for a comparison of different reflection strategies. Furthermore, a single trajectory-level advantage updates both the textual reflections and the image revisions, avoiding the complexity of assigning credit for success or failure at each individual step.
Unlike methods that edit images in a single round or use separate external critics, UMM-Reflection allows credit to flow across multiple rounds and updates both the model's diagnostic and revision capabilities simultaneously. This eliminates the need for a separate verifier during inference.
Testing on the BAGEL benchmark showed that UMM-Reflection improved the GenEval score by 12.05 points compared to SFT. The gains were also observed on other benchmarks, including WISE, OneIG-Bench, and T2I-CompBench++, even though these datasets were not used in the training of UMM-Reflection. The researchers note that while imitation learning can teach models to produce revisions, it does not guarantee reliability. Reinforcement learning, through UMM-Reflection, makes these revisions reliable by selecting repair paths that the model already possesses the capability to produce.
The code and models for UMM-Reflection were released on September 28, 2026, along with a project page and a video demonstration.
