Google Research has introduced Diffusion Controller, a lightweight side network that steers text-to-image models during generation. It attaches to the base model without modifying it, which matters for closed-source systems. The approach reframes the denoising process as a continuous control problem, with the side network acting like a steering damper on a motorcycle.

The framework aims to solve the trade-off between following prompts and preserving image quality. For example, a model may produce a lizard without sunglasses, or distort the lizard's face when forced to add them. Diffusion Controller adjusts the generation trajectory toward user-defined rewards while a penalty guardrail prevents distortion.

The authors describe two fine-tuning methods based on final reward scores: a policy-gradient/PPO approach with clipping for stable updates, and a reward-weighted loss that directly favors high-quality generations. Both let engineers customize even black-box or gray-box models by observing intermediate states and injecting small corrections.

In evaluations, the fully unlocked version achieved a 90% win rate over the baseline, and the lightweight version outperformed the industry standard for matching human preferences. The work unifies previously disconnected guidance and fine-tuning techniques under one mathematical framework.