Achieving 4‑Step Diffusion Generation by Replacing MSE with Perceptual Loss in Five Lines of Code
By swapping the traditional MSE loss for a perceptual loss in Flow Matching training, the authors enable high‑quality diffusion generation in only 4–8 inference steps—down from 35–50—without teacher models, distribution or trajectory distillation, and they substantiate the claim with extensive experiments and a new distribution‑distance metric.
